Lab notes #1: FP8 KV cache on an H100 — 11% more goodput, and it failed our accuracy gate
First real-hardware run of TokenWatt Bench. Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. FP8 KV cache doubled KV capacity and raised SLO goodput 10.6% and tokens per joule 14%, but GSM8K accuracy dropped 8 points with uncalibrated KV scales.
Update (2026-10-07): a full-mode rerun corrected two numbers here. With 10-minute windows, default-configuration goodput is 3,996 tok/s (not 4,572), and the uncalibrated FP8 KV-cache accuracy drop is 4.0 points on 600 samples (not 8 on 50). It still fails the gate.
This is the first run of TokenWatt Bench on real hardware: one NVIDIA H100 on Lambda Cloud, serving a mixture-of-experts model with vLLM. We compared vLLM's default BF16 KV cache against an FP8 KV cache. It's a one-flag change that many operators turn on for capacity.
Short version:
- FP8 KV cache doubled the KV capacity from 445,008 to 890,080 tokens in the same 40.74 GiB.
- It raised SLO goodput by 10.6% (4,572 → 5,057 tok/s) and tokens per joule by 14% (7.18 → 8.21, GPU board power).
- It failed the accuracy gate: 96% → 88% on our GSM8K sample. vLLM warned that it was using an uncalibrated KV scaling factor of 1.0.
- Under our rules, a configuration that fails the gate earns no capacity credit. So the report says no saving, despite the higher throughput.
These are quick-mode numbers: short windows, two repeats, a 50-item accuracy sample. They validate the tool, not a production decision. The draft report has every number and the raw JSON.
Setup
| Instance | Lambda Cloud gpu_1x_h100_sxm5, us-southeast-1 |
| GPU | 1× NVIDIA H100 80GB HBM3 (SXM5), driver 580.105.08, 700 W power limit |
| Engine | vLLM 0.31.0 (vllm/vllm-openai image), TP=1, --max-model-len 12288 |
| Model | Qwen/Qwen3-30B-A3B-FP8 (MoE: 48 layers × 128 experts, top-8; FP8 weights) |
| Configurations | default (BF16 KV cache) vs --kv-cache-dtype fp8 |
| SLO | TTFT p99 ≤ 2 s, TPOT p99 ≤ 50 ms |
| Workload | chat-mix-v1: input median 1,024 / p90 4,096 tokens; output median 256 / p90 1,024; Poisson arrivals |
| Measurement | 45 s warm-up, 60 s search windows, 2 × 120 s repeats at the SLO boundary |
| Power | GPU board power via nvidia-smi at 1 Hz (cloud VM, no BMC access) |
| Accuracy | GSM8K 5-shot, first 50 test items, temperature 0, one seed |
The whole session ran under our launch-run-terminate script: a smoke test plus both configurations. It took 63 minutes and cost about $4.48, and the instance was terminated automatically at the end.
Results
| BF16 KV (default) | FP8 KV | |
|---|---|---|
| SLO boundary | 11.31 req/s | 12.34 req/s |
| SLO goodput | 4,572 tok/s | 5,057 tok/s |
| GPU board power | 0.64 kW | 0.62 kW |
| Tokens per joule | 7.18 | 8.21 |
| TTFT p99 / TPOT p99 | 340 ms / 44.6 ms | 352 ms / 43.7 ms |
| Decode bandwidth utilization | 56.6% | 45.3% |
| GSM8K (50 items) | 96% | 88% — gate failed |
| Failed requests | 0 | 0 |
Both configurations met the SLO in both repeats, with variance bands of ±0.6% and ±1.0%.
Where the extra goodput comes from
The SLO search shows the mechanism. At 16 req/s the two configurations fail in different ways:
| at 16 req/s | BF16 KV | FP8 KV |
|---|---|---|
| TTFT p99 | 31,888 ms | 532 ms |
| TPOT p99 | 48.5 ms | 76.6 ms |
With a BF16 KV cache, the server runs out of KV space first. Requests queue, and time-to-first-token explodes to 32 seconds. With twice the KV capacity, the FP8 configuration admits more concurrent sequences. The limit then shifts to per-token latency instead, as bigger batches make every decode step slower. Between those two failure modes, FP8 KV holds the SLO about one request per second longer. That's our reading of the sweep, not a profiler trace.
Two other things in the table are worth explaining:
- MBU went down while goodput went up. That's expected. MBU counts the bytes the model must read, and an FP8 KV cache halves the KV bytes per token. The derived KV traffic dropped from about 940 GB/s to about 525 GB/s, while weight traffic stayed near 1 TB/s. How MBU is derived.
- Board power fell slightly while throughput rose, which is where the 14% tokens-per-joule gain comes from.
nvidia-smireported the software power-cap throttle reason in 81–85% of 1 Hz samples for the default configuration and 53–58% for FP8 KV, even though mean board power was about 630 W against a 700 W limit. We read that as short power transients above the limit, and the FP8 configuration hitting the cap less often fits its lower power draw. We haven't profiled it further yet.
The accuracy gate
The FP8 KV configuration scored 88% against 96% for the default on the same 50 GSM8K questions. That's 4 more wrong answers, and the drop is 8 points against a gate threshold of 1 point.
The vLLM log points at a likely cause:
Using KV cache scaling factor 1.0 for fp8_e4m3. If this is unintended, verify that
k/v_scale scaling factors are properly set in the checkpoint.
Using uncalibrated q_scale 1.0 and/or prob_scale 1.0 with fp8 attention.
This may cause accuracy issues.
The checkpoint ships FP8 weights but no calibrated KV-cache scales. So the FP8 KV cache ran with a scale of 1.0.
Fifty items and one seed is a small sample. We wouldn't call FP8 KV "8 points worse" from this run. But the gate does exactly what it's for: a configuration that might be trading accuracy for speed doesn't get to claim the speed until it's shown it isn't. The next run uses the full suite (200 items × 3 seeds) and calibrated KV scales.
What the run changed in the tool
The first real-hardware run is also a test of Bench itself. We changed four things:
- KV-cache precision is now a per-configuration setting, not part of the spec fingerprint. Bench originally refused to compare these two runs because their model section differed. Comparing KV precision is a legitimate question, and the accuracy gate already guards it.
- Gate failures are excluded from capacity math everywhere: the terminal summary, report pages and the conclusion sentence.
- Short windows give a misleading MBU. Our 10-minute smoke test reported 73.7% MBU at a lower operating point (batch ≈ 79) versus 56.6% in the quick run (batch ≈ 145). Smoke mode now exists only to check the pipeline; its numbers aren't reported.
- Plumbing: an explicit User-Agent for the Lambda API, which sits behind Cloudflare, and live log streaming from the instance.
Not covered
- Whole-node power. These are GPU board power numbers from a cloud VM.
- Context lengths beyond the
chat-mix-v1trace, other quantizations, and multi-GPU tensor or expert parallelism. - Statistical confidence on accuracy. That needs the full suite.
Reproduce
cd bench
scripts/lambda/cloud.py --api-key-file <key> session --ssh-key <key> --type gpu_1x_h100_sxm5 \
--run "MODE=quick LABEL=vllm-default ROLE=baseline MODEL=Qwen/Qwen3-30B-A3B-FP8 TP=1" \
--run "MODE=quick LABEL=vllm-fp8kv ROLE=tuned MODEL=Qwen/Qwen3-30B-A3B-FP8 TP=1 KV_DTYPE=fp8 VLLM_ARGS='--kv-cache-dtype fp8'" --yes
Next up: the full-mode rerun with calibrated scales, and the same comparison on AMD Instinct.