Qwen3-30B-A3B FP8 on 1× H100: BF16 vs FP8 KV cache
Validation run on a Lambda Cloud 1× H100 SXM5 instance with vLLM 0.31.0. Quick mode: 120 s windows, 2 repeats, 50 GSM8K items with one seed. FP8 KV cache raised SLO goodput and tokens per joule but failed the accuracy gate on this small suite.
ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on vllm-default takes: vllm-fp8kv failed the accuracy gate (-8.0 pts) and is excluded from capacity math, despite 5,057 tok/s goodput.
Results
Each value is the median of 2 runs.
vllm-defaultvllm-fp8kv
SLO goodput
output tok/s per node · higher is better
Tokens per joule
tok/J, GPU board power · higher is better
Decode bandwidth utilization
% of rated HBM bandwidth
Axis runs to 100% of rated bandwidth.
GPU board power
kW, steady state
GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.
All numbers
Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.
| Metric | vllm-default | vllm-fp8kv |
|---|---|---|
| Engine build | vLLM 0.31.0 · KV BF16 | vLLM 0.31.0 · KV FP8 |
| Decode bandwidth utilization (MBU) | 57% | 45% |
| GPU board power | 0.64 kW | 0.62 kW |
| SLO goodput | 4,572 tok/s | 5,057 tok/s |
| Tokens per joule | 7.18 | 8.21 |
| TTFT p99 | 340 ms | 352 ms |
| TPOT p99 | 44.6 ms | 43.7 ms |
| SLO met | yes | yes |
| Accuracy (gsm8k-5shot) | 96.0 | 88.0 |
| Accuracy Δ vs. reference | reference | -8.0 pts |
| Cards for equal goodput | 64 | — |
Setup
Any change to these fields makes it a different report.
- Accelerator
- 1× NVIDIA H100 80GB HBM3
- Rated HBM BW
- 3.35 TB/s per GPU
- Driver / runtime
- 580.105.08
- Host
- Lambda Cloud instance
- Engine
- vLLM 0.31.0
- Kernel libraries
- vllm/vllm-openai:latest
- Model
- Qwen3-30B-A3B-FP8 (MoE)
- Quantization
- FP8 weights; KV cache varies by configuration
- Parallelism
- 1 replica × TP1
- Workload trace
- chat-mix-v1
- Input tokens
- median 1,024 · p90 4,096
- Output tokens
- median 256 · p90 1,024
- Arrival
- Poisson, rate swept to SLO boundary
- Power source
- GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
- Runs per config
- 2 (variance band ±1.0%)
- Accuracy gate
- gsm8k-5shot · max drop 1.0 pts
Not covered
Do not extrapolate this report to the following.
- Whole-node power (cloud VM: no BMC access)
- Context lengths beyond the chat-mix-v1 trace
- Quick mode: short windows, few repeats — validation only, not for publishing
- Whole-node power: energy figures are GPU board power only (no CPUs, memory, NICs, fans or PSU losses)
Reproduce
Run on the same hardware and software versions. Results should fall within ±1.0%.
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'vllm-default' --role baseline --out runs/vllm-default.json tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'vllm-fp8kv' --role tuned --out runs/vllm-fp8kv.json tokenwatt report runs/vllm-default.json runs/vllm-fp8kv.json --title 'Qwen3-30B-A3B FP8 on 1× H100: BF16 vs FP8 KV cache'