FP8 vs BF16 for LLM inference: faster tokens, but not always more capacity
We served Qwen3-8B with BF16 and FP8 weights on an RTX 4090 and an L40S. FP8 cut the median time per token by 35–42% on both cards, but it raised capacity within our latency target only on the 4090, where BF16 left too little memory for the KV cache. On the L40S, capacity didn't move.
FP8 weights take half the memory of BF16 and half the bandwidth to read, so the usual expectation is a big speedup for LLM serving. We measured what it actually changes, on the two Ada-generation GPUs in our cheaper-GPU comparison, which support FP8 natively.
Results for Qwen3-8B on vLLM 0.31.0, one GPU each, latency target p99 time to first token ≤ 2 s and p99 time per output token (TPOT) ≤ 50 ms:
| RTX 4090 BF16 | RTX 4090 FP8 | L40S BF16 | L40S FP8 | |
|---|---|---|---|---|
| Median time per output token | 21.1 ms | 13.7 ms (−35%) | 29.1 ms | 16.8 ms (−42%) |
| p99 time per output token | 29.5 ms | 21.6 ms | 37.8 ms | 25.3 ms |
| Capacity within the SLO | 0.62 req/s | 0.74 req/s (+19%) | 0.72 req/s | 0.72 req/s (±0) |
| Output tokens/s at capacity | 223 | 276 | 272 | 268 |
| Cost per million output tokens | $1.11 | $0.90 | $1.11 | $1.13 |
| GSM8K accuracy (50 questions) | 94% | 94% | 92% | 94% |
FP8 made every token faster on both cards, with no measurable accuracy loss on our check. It raised capacity on one of them.
Why the RTX 4090 gained capacity
The 4090 has 24 GB of memory. In BF16, Qwen3-8B's weights take about 16 GB, and vLLM's default memory setting leaves only about 5 GB for the KV cache, the per-request memory that holds each conversation's context. With our chat workload the cache ran full (vLLM reported 97% usage) and new requests had to wait, which pushed up time to first token.
FP8 halves the weights to about 8 GB, so the cache roughly doubles. More requests fit at once, and capacity rose 19%. Cost per token fell from $1.11 to $0.90, the cheapest of every GPU and format we tested.
Why the L40S didn't
The L40S has 48 GB, so memory was never the limit: even in BF16 its KV cache stayed at about 25% full. FP8 made each decoding step faster by halving the bytes read, which is why the median time per token fell 42%.
But capacity is set by the slowest tokens, not the median. Occasionally a long prompt (our workload includes some up to 8,000 tokens) gets processed in the same step as everyone else's next token, and those tokens stall. FP8 speeds up the normal steps, but the stalls remained: at 1.02 requests per second the FP8 run's median time per token was 23 ms while its p99 reached 83 ms. Capacity came out at the same 0.72 requests per second in both formats.
When FP8 is worth it
- Memory-tight cards or models: if the KV cache is full in BF16 (watch for waiting requests and high cache usage in your engine's logs), FP8 weights free memory for more concurrent requests and raise capacity directly.
- Latency-sensitive serving: FP8 cut the median time per token by 35–42%, so users see faster streaming even where capacity doesn't change.
- Check accuracy on your own task. We saw no loss on a 50-question math check. Earlier, an FP8 KV cache (a different setting) failed our accuracy gate on an H100. Test what you ship.
- Not on Ampere. The A100 and A40 have no native FP8 support.
If the stalls from long prompts are what limit you, FP8 alone won't fix them. Spreading prompt processing out (smaller prefill chunks in vLLM) raised capacity by up to 15% on a 70B model in our tests.
Limits
Quick-mode measurements (two 2-minute repeats), one node per configuration, one model, default vLLM settings, 50-question accuracy check. Full per-card data: Qwen3-8B benchmark.
Raw data: RTX 4090 FP8 · RTX 4090 BF16 · L40S FP8 · L40S BF16.