Tokens per second is quoted in two senses, and mixing them up leads to bad capacity plans.
Per user: streaming speed
How fast one answer appears. It's the inverse of time per output token (TPOT): at 25 ms per token, a user sees 40 tokens per second. People read roughly 5–10 tokens per second, so 20+ feels instant for chat. Agents and code generation want more.
Per GPU: throughput
How many output tokens the GPU produces in total, across all requests it serves at once. A GPU decodes many requests in each step, so its throughput is far higher than any one user's speed: an H100 served Qwen3-30B-A3B at about 3,800 output tokens per second while each user's tokens arrived roughly every 33 ms.
The trade-off
Serving more requests at once raises GPU throughput but slows each user down, because every decoding step does more work. So "maximum tokens per second" alone means little. The useful number is throughput at a latency target: the most output tokens per second a GPU delivers while 99% of tokens still arrive within, say, 50 ms. We call that SLO goodput. It's what decides cost per token.
Values we measured (throughput at the SLO)
All at p99 time to first token ≤ 2 s and p99 time per output token ≤ 50 ms, with a chat-like workload:
| GPU | Model | Output tokens/s per GPU | Median time per token |
|---|---|---|---|
| H100 SXM (vLLM, full mode) | Qwen3-30B-A3B FP8 | 3,764 | 33 ms |
| H100 SXM (SGLang, full mode) | Qwen3-30B-A3B FP8 | 3,814 | 24 ms |
| A100 SXM 80 GB | Qwen3-8B BF16 | 469 | 16 ms |
| RTX 4090 | Qwen3-8B FP8 | 276 | 14 ms |
| L40S | Qwen3-8B BF16 | 272 | 29 ms |
| A40 | Qwen3-8B BF16 | 102 | 34 ms |
Sources: SGLang vs vLLM on H100, L40S vs A100 vs RTX 4090 vs A40. The 30B model produces far more tokens per second than the 8B one because it's a mixture-of-experts model that reads only about 3 billion parameters per token.
FAQ
How many tokens per second is fast enough?
Per user, 20 or more tokens per second feels fast for chat, since people read 5–10. Per GPU, what matters is throughput while still meeting your latency target, not the peak.
Why is my GPU's tokens per second much higher than what one user sees?
The GPU serves many requests at once and produces tokens for all of them in each step. Its total throughput is the sum across users; each user sees only their own stream.
Does a faster GPU always mean more tokens per second?
For token generation, memory bandwidth matters most, because each step reads the model's weights. The model matters as much as the GPU: a mixture-of-experts model with few active parameters generates far more tokens per second than a dense model of similar size.