← Lab notes

Lab notes #2: Short benchmark windows overstated H100 goodput by 14%

Full-mode rerun of Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. With 10-minute windows the SLO boundary fell from 11.3 to 9.5 req/s and goodput from 4,572 to 3,996 tok/s; the FP8 KV-cache accuracy drop shrank from 8 to 4 points on a 600-sample suite.

In lab notes #1 we ran a quick validation: two-minute windows, two repeats, and a 50-question accuracy sample. This time we reran the same model on the same kind of GPU the way a published report is supposed to be run. That meant 10-minute windows, five repeats at the SLO boundary, and 200 GSM8K questions × 3 seeds.

Two of our earlier numbers didn't survive.

  • Goodput fell 13%, from 4,572 to 3,996 tok/s. The sustainable request rate dropped from 11.3 to 9.5 req/s.
  • The FP8 KV-cache accuracy penalty was half what we reported: 4.0 points, not 8. It still fails our 1-point gate.

The result is our first published report: Qwen3-30B-A3B FP8 on 1× H100 SXM with vLLM 0.31.0.

Setup

Instance Lambda Cloud gpu_1x_h100_sxm5, us-southeast-1
GPU 1× NVIDIA H100 80GB HBM3 (SXM5), driver 580.105.08
Engine vLLM 0.31.0, TP=1, --max-model-len 12288, default BF16 KV cache
Model Qwen/Qwen3-30B-A3B-FP8 (MoE, FP8 weights)
SLO TTFT p99 ≤ 2 s, TPOT p99 ≤ 50 ms
Workload chat-mix-v1, Poisson arrivals
Measurement 120 s warm-up; 120 s search windows; 5 × 600 s repeats at the boundary
Accuracy GSM8K 5-shot, 200 items × 3 seeds
Power GPU board power via nvidia-smi, 1 Hz

Results

Quick (lab notes #1) Full (this run)
SLO boundary 11.31 req/s 9.50 req/s
SLO goodput 4,572 tok/s 3,996 tok/s
TTFT p99 / TPOT p99 340 ms / 44.6 ms 327 ms / 37.3 ms
GPU board power 0.64 kW 0.60 kW
Tokens per joule 7.18 6.65
Decode bandwidth utilization 56.6% 58.1%
GSM8K 96% (50 items, 1 seed) 94.17% (200 items × 3 seeds)

All five repeats met the SLO, with goodput between 3,948 and 4,127 tok/s (variance band ±2.2%). Across about 28,400 requests in the measured windows, none failed.

At Lambda's $4.29 per GPU-hour, that's $0.30 per million output tokens at full utilization within the SLO. Our GPU price table now uses this measurement for its cost-per-token column.

Why short windows were optimistic

The search phase uses two-minute windows. In those windows, 11.73 req/s passed with a p99 TTFT of 1.6 s. Then the 10-minute repeats started.

Rate Window TTFT p99 TPOT p99 Result
11.73 req/s 600 s 18,213 ms 48.2 ms fail
10.56 req/s 600 s 2,895 ms 43.9 ms fail
9.50 req/s 600 s × 5 308–361 ms 35.6–39.3 ms pass

Every failure was time-to-first-token; per-token latency stayed inside the SLO. That pattern points to a queue that builds slowly. Near capacity, the KV cache fills and new requests wait for space. Two minutes isn't long enough for the backlog to show up in the p99. Ten minutes is.

Bench handled this the way it's designed to. When a repeat misses the SLO, it backs off the rate by 10% and starts the repeats over. Two back-offs later, all five repeats passed.

The lesson for anyone reading benchmark numbers: a throughput figure is only as good as the window it was measured over. Short runs near saturation measure the queue before it forms.

FP8 KV cache: still fails the gate, by less

KV cache Score Per seed
BF16 (default) 94.17% 94.5 · 94.0 · 94.0
FP8, uncalibrated scale 1.0 90.17% 89.5 · 91.0 · 90.0

The drop is 4.0 points, against a 1-point gate. Every FP8 seed is more than three points below every BF16 seed, so it's not noise. But the 8-point drop we reported from 50 questions was too pessimistic by half. Small accuracy samples swing both ways.

We'd planned to test vLLM's runtime KV-scale calibration next. It's gone:

vllm: error: unrecognized arguments: --calculate-kv-scales

vLLM 0.31.0 no longer accepts the flag, so calibrated FP8 KV on this version means a checkpoint that ships its own k/v scales. That's the next experiment. Bench's gated run skipped the 100-minute load test for the calibrated configuration automatically, because the accuracy check it depended on never ran.

Two more things the cloud taught us

  • Driver versions vary between instances of the same type. Our first launch got a host with NVIDIA driver 570.148.08. vLLM 0.31.0's CUDA-graph capture failed there with CUDA error: operation not permitted. The quick run had used 580.105.08. Bench now checks a minimum driver version at preflight and fails in seconds, and our launcher excludes the region and retries. The failed host cost $0.84.
  • The software power cap is busy. nvidia-smi reported the software power-cap throttle reason in 62% of 1 Hz samples, even though mean board power was about 600 W against a 700 W limit. Short spikes above the limit are the likely cause; we'll look closer when we have whole-node power.

The run took 131 minutes and cost about $9.36 on Lambda. The instance was terminated automatically when the gated run was skipped.

Not covered

  • Whole-node power (cloud VM, no BMC access).
  • Context lengths beyond the chat-mix-v1 trace, other quantizations, multi-GPU parallelism.
  • A calibrated FP8 KV cache. See above.

Reproduce

cd bench
scripts/lambda/cloud.py --api-key-file <key> session --ssh-key <key> --type gpu_1x_h100_sxm5 --yes \
  --run "MODE=full LABEL=vllm-default ROLE=baseline MODEL=Qwen/Qwen3-30B-A3B-FP8 TP=1 MIN_DRIVER=580 SWEEP_START=9 SWEEP_GROWTH=1.25 SEARCH_WINDOW_S=120" \
  --run "MODE=full ACCURACY_ONLY=1 LABEL=acc-fp8kv-uncal MODEL=Qwen/Qwen3-30B-A3B-FP8 TP=1 KV_DTYPE=fp8 VLLM_ARGS='--kv-cache-dtype fp8' MIN_DRIVER=580"

Next: the same measurement on AMD Instinct MI300X, then a calibrated FP8 KV checkpoint.