← Reports

Model benchmark

Qwen3-30B-A3B inference benchmark: H100 vs H200, vLLM vs SGLang, cost per token

Every measurement we've made of Qwen3-30B-A3B (FP8) on one GPU, in one place. Capacity within a latency target, output tokens per second and cost per million tokens on H100 and H200, with vLLM and SGLang.

Qwen3-30B-A3B is a mixture-of-experts model with 30 billion parameters, of which about 3 billion are active for each token. That makes it unusually cheap to serve for its size: each token reads about a tenth of the weights. We use the FP8 version as our standard model on data-center GPUs.

Results

All results on one GPU, chat-like workload (median 1,024 input and 256 output tokens), latency target p99 time to first token ≤ 2 s and p99 time per output token (TPOT) ≤ 50 ms unless noted. Price is the cloud's on-demand price on the test day.

GPU, cloud (price) Engine Mode Capacity Output tokens/s Cost per million output tokens
H100 SXM, RunPod ($3.99/h) vLLM 0.31.0 full 8.92 req/s 3,764 $0.294
H100 SXM, RunPod ($3.99/h) SGLang 0.5.21 full 9.07 req/s 3,814 $0.291
H100 SXM, RunPod ($3.99/h) vLLM 0.31.0 quick 11.18 req/s 4,526 $0.245
H100 SXM, Lambda ($4.29/h) vLLM 0.31.0 quick 11.18 req/s 4,515 $0.264
H200, RunPod ($5.29/h) vLLM 0.31.0 quick 13.22 req/s 5,430 $0.271

With a tighter 30 ms TPOT target:

GPU Engine Mode Capacity Output tokens/s Cost per million output tokens
H100 SXM vLLM 0.31.0 full 6.49 req/s 2,749 $0.403
H100 SXM SGLang 0.5.21 full ~7.9 req/s (provisional) ~3,250 ~$0.34
H100 SXM vLLM 0.31.0 quick 8.03 req/s 3,255 $0.340
H200 vLLM 0.31.0 quick 10.33 req/s 4,172 $0.352

Quick vs full: quick mode uses 2-minute windows, full mode 10-minute windows with five repeats. Full mode reads 15–20% lower capacity because longer windows catch more of the rare slow requests. Plan with full-mode numbers.

Accuracy on our GSM8K check was 91–94% in every configuration.

What we learned

  • H100 vs H200: the H200 served 18% more requests at 50 ms and 29% more at 30 ms, but at RunPod's prices each token cost 11% more at 50 ms. The model is small and leaves most of the H200's extra memory unused. Full comparison.
  • vLLM vs SGLang: the same capacity at 50 ms, but at 9 requests/s SGLang returned the first token in 67 ms against 196 ms and generated tokens about a third faster, so it led by about 21% at 30 ms. Full comparison.
  • Clouds: the same H100 cost $0.245 to $0.264 per million tokens on healthy nodes at RunPod and Lambda. A degraded marketplace node couldn't meet the latency target at all. Full comparison.
  • Kernel tuning: tuning vLLM's MoE kernels for this model on an H100 made no measurable difference. Lab notes #6.
  • Compared with a dense model: this 30B model on an H100 cost about 3× less per token than Qwen3-8B on any cheaper GPU we tested.

How to reproduce

The measurements come from TokenWatt Bench, with the method described in our methodology. Reports with raw data: H100 full mode · vLLM vs SGLang · H200 · MoE kernel tuning. Before benchmarking, check the node with pip install tokenwatt-check && tokenwatt-check.

FAQ

How many requests per second can one H100 serve with Qwen3-30B-A3B?

About 9 requests per second in our full-mode measurement (vLLM or SGLang), with p99 time to first token under 2 seconds and p99 time per output token under 50 ms, for a chat-like workload. That's about 3,800 output tokens per second.

Is an H200 worth it for Qwen3-30B-A3B?

Only at a tight latency target or a low H200 price. At a 50 ms target it served 18% more requests than an H100 but cost 33% more per hour at RunPod. At 30 ms it served 29% more, close to break-even.

What does Qwen3-30B-A3B cost per million tokens on an H100?

About $0.25 to $0.30 per million output tokens at $3.99 to $4.29 an hour, running at capacity. At lower utilization the cost per token rises proportionally.