Qwen3-30B-A3B is a mixture-of-experts model with 30 billion parameters, of which about 3 billion are active for each token. That makes it unusually cheap to serve for its size: each token reads about a tenth of the weights. We use the FP8 version as our standard model on data-center GPUs.
Results
All results on one GPU, chat-like workload (median 1,024 input and 256 output tokens), latency target p99 time to first token ≤ 2 s and p99 time per output token (TPOT) ≤ 50 ms unless noted. Price is the cloud's on-demand price on the test day.
| GPU, cloud (price) | Engine | Mode | Capacity | Output tokens/s | Cost per million output tokens |
|---|---|---|---|---|---|
| H100 SXM, RunPod ($3.99/h) | vLLM 0.31.0 | full | 8.92 req/s | 3,764 | $0.294 |
| H100 SXM, RunPod ($3.99/h) | SGLang 0.5.21 | full | 9.07 req/s | 3,814 | $0.291 |
| H100 SXM, RunPod ($3.99/h) | vLLM 0.31.0 | quick | 11.18 req/s | 4,526 | $0.245 |
| H100 SXM, Lambda ($4.29/h) | vLLM 0.31.0 | quick | 11.18 req/s | 4,515 | $0.264 |
| H200, RunPod ($5.29/h) | vLLM 0.31.0 | quick | 13.22 req/s | 5,430 | $0.271 |
With a tighter 30 ms TPOT target:
| GPU | Engine | Mode | Capacity | Output tokens/s | Cost per million output tokens |
|---|---|---|---|---|---|
| H100 SXM | vLLM 0.31.0 | full | 6.49 req/s | 2,749 | $0.403 |
| H100 SXM | SGLang 0.5.21 | full | ~7.9 req/s (provisional) | ~3,250 | ~$0.34 |
| H100 SXM | vLLM 0.31.0 | quick | 8.03 req/s | 3,255 | $0.340 |
| H200 | vLLM 0.31.0 | quick | 10.33 req/s | 4,172 | $0.352 |
Quick vs full: quick mode uses 2-minute windows, full mode 10-minute windows with five repeats. Full mode reads 15–20% lower capacity because longer windows catch more of the rare slow requests. Plan with full-mode numbers.
Accuracy on our GSM8K check was 91–94% in every configuration.
What we learned
- H100 vs H200: the H200 served 18% more requests at 50 ms and 29% more at 30 ms, but at RunPod's prices each token cost 11% more at 50 ms. The model is small and leaves most of the H200's extra memory unused. Full comparison.
- vLLM vs SGLang: the same capacity at 50 ms, but at 9 requests/s SGLang returned the first token in 67 ms against 196 ms and generated tokens about a third faster, so it led by about 21% at 30 ms. Full comparison.
- Clouds: the same H100 cost $0.245 to $0.264 per million tokens on healthy nodes at RunPod and Lambda. A degraded marketplace node couldn't meet the latency target at all. Full comparison.
- Kernel tuning: tuning vLLM's MoE kernels for this model on an H100 made no measurable difference. Lab notes #6.
- Compared with a dense model: this 30B model on an H100 cost about 3× less per token than Qwen3-8B on any cheaper GPU we tested.
How to reproduce
The measurements come from TokenWatt Bench, with the method described in our methodology. Reports with raw data: H100 full mode · vLLM vs SGLang · H200 · MoE kernel tuning. Before benchmarking, check the node with pip install tokenwatt-check && tokenwatt-check.
FAQ
How many requests per second can one H100 serve with Qwen3-30B-A3B?
About 9 requests per second in our full-mode measurement (vLLM or SGLang), with p99 time to first token under 2 seconds and p99 time per output token under 50 ms, for a chat-like workload. That's about 3,800 output tokens per second.
Is an H200 worth it for Qwen3-30B-A3B?
Only at a tight latency target or a low H200 price. At a 50 ms target it served 18% more requests than an H100 but cost 33% more per hour at RunPod. At 30 ms it served 29% more, close to break-even.
What does Qwen3-30B-A3B cost per million tokens on an H100?
About $0.25 to $0.30 per million output tokens at $3.99 to $4.29 an hour, running at capacity. At lower utilization the cost per token rises proportionally.