← Methodology

Glossary

Tokens per second for LLM inference: per user, per GPU, and why both matter

\"Tokens per second\" means two different things in LLM inference: how fast one user's answer streams, and how many tokens a whole GPU produces. They pull in opposite directions. Here's how they relate and what we measured on real GPUs.

Tokens per second is quoted in two senses, and mixing them up leads to bad capacity plans.

Per user: streaming speed

How fast one answer appears. It's the inverse of time per output token (TPOT): at 25 ms per token, a user sees 40 tokens per second. People read roughly 5–10 tokens per second, so 20+ feels instant for chat. Agents and code generation want more.

Per GPU: throughput

How many output tokens the GPU produces in total, across all requests it serves at once. A GPU decodes many requests in each step, so its throughput is far higher than any one user's speed: an H100 served Qwen3-30B-A3B at about 3,800 output tokens per second while each user's tokens arrived roughly every 33 ms.

The trade-off

Serving more requests at once raises GPU throughput but slows each user down, because every decoding step does more work. So "maximum tokens per second" alone means little. The useful number is throughput at a latency target: the most output tokens per second a GPU delivers while 99% of tokens still arrive within, say, 50 ms. We call that SLO goodput. It's what decides cost per token.

Values we measured (throughput at the SLO)

All at p99 time to first token ≤ 2 s and p99 time per output token ≤ 50 ms, with a chat-like workload:

GPU Model Output tokens/s per GPU Median time per token
H100 SXM (vLLM, full mode) Qwen3-30B-A3B FP8 3,764 33 ms
H100 SXM (SGLang, full mode) Qwen3-30B-A3B FP8 3,814 24 ms
A100 SXM 80 GB Qwen3-8B BF16 469 16 ms
RTX 4090 Qwen3-8B FP8 276 14 ms
L40S Qwen3-8B BF16 272 29 ms
A40 Qwen3-8B BF16 102 34 ms

Sources: SGLang vs vLLM on H100, L40S vs A100 vs RTX 4090 vs A40. The 30B model produces far more tokens per second than the 8B one because it's a mixture-of-experts model that reads only about 3 billion parameters per token.

FAQ

How many tokens per second is fast enough?

Per user, 20 or more tokens per second feels fast for chat, since people read 5–10. Per GPU, what matters is throughput while still meeting your latency target, not the peak.

Why is my GPU's tokens per second much higher than what one user sees?

The GPU serves many requests at once and produces tokens for all of them in each step. Its total throughput is the sum across users; each user sees only their own stream.

Does a faster GPU always mean more tokens per second?

For token generation, memory bandwidth matters most, because each step reads the model's weights. The model matters as much as the GPU: a mixture-of-experts model with few active parameters generates far more tokens per second than a dense model of similar size.