⚡ vLLM
← All demos
Simulated demo
Model
Llama-3-70B
Throughput
0 tok/s
Active requests
0
Queue depth
0
GPU utilization (4x A100)
KV cache usage
block pool
0%
Recent completions
Request ID
Prompt tokens
Output tokens
Latency
Status