Tuned vs default fused-MoE kernels: Qwen3-30B-A3B FP8 on 1× H100 SXM
Lambda Cloud 1× H100 SXM, 2026-10-08: vLLM 0.31.0 with its default fused-MoE Triton configs, then with configs tuned for this GPU and model (2.5 hours of kernel tuning), then default again to bound drift. Quick mode. The tuned kernels made no measurable difference to SLO goodput.
ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on default kernels takes: with tuned kernels, 64 → 64 cards and 3% more GPU board energy per token; with default kernels (repeat), 64 → 64 cards and 1% more GPU board energy per token.
Results
Each value is the median of 2 runs.
default kernelstuned kernelsdefault kernels (repeat)
SLO goodput
output tok/s per node · higher is better
Tokens per joule
tok/J, GPU board power · higher is better
Decode bandwidth utilization
% of rated HBM bandwidth
Axis runs to 100% of rated bandwidth.
GPU board power
kW, steady state
GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.
All numbers
Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.
| Metric | default kernels | tuned kernels | default kernels (repeat) |
|---|---|---|---|
| Engine build | vLLM 0.31.0 · KV BF16 | vLLM 0.31.0 · KV BF16 | vLLM 0.31.0 · KV BF16 |
| Decode bandwidth utilization (MBU) | 56% | 58% | 57% |
| GPU board power | 0.62 kW | 0.64 kW | 0.63 kW |
| SLO goodput | 4,515 tok/s | 4,523 tok/s | 4,516 tok/s |
| Tokens per joule | 7.25 | 7.06 | 7.16 |
| TTFT p99 | 338 ms | 333 ms | 333 ms |
| TPOT p99 | 44.2 ms | 42.3 ms | 43.0 ms |
| SLO met | yes | yes | yes |
| Accuracy (gsm8k-5shot) | 96.0 | 96.0 | 96.0 |
| Accuracy Δ vs. reference | reference | 0.0 pts | 0.0 pts |
| Cards for equal goodput | 64 | 64 | 64 |
Setup
Any change to these fields makes it a different report.
- Accelerator
- 1× NVIDIA H100 80GB HBM3
- Rated HBM BW
- 3.35 TB/s per GPU
- Driver / runtime
- 580.105.08
- Host
- Lambda Cloud instance
- Engine
- vLLM 0.31.0
- Kernel libraries
- vllm/vllm-openai:latest
- Model
- Qwen3-30B-A3B-FP8 (MoE)
- Quantization
- FP8 weights, BF16 KV cache
- Parallelism
- 1 replica × TP1
- Workload trace
- chat-mix-v1
- Input tokens
- median 1,024 · p90 4,096
- Output tokens
- median 256 · p90 1,024
- Arrival
- Poisson, rate swept to SLO boundary
- Power source
- GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
- Runs per config
- 2 (variance band ±0.7%)
- Accuracy gate
- gsm8k-5shot · max drop 1.0 pts
Not covered
Do not extrapolate this report to the following.
- Whole-node power (no BMC access on this host)
- Context lengths beyond the chat-mix-v1 trace
- Quick mode: short windows, few repeats — validation only, not for publishing
Reproduce
Run on the same hardware and software versions. Results should fall within ±0.7%.
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-default' --role baseline --out runs/moe-default.json tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-tuned' --role tuned --out runs/moe-tuned.json tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-default-2' --role other --out runs/moe-default-2.json tokenwatt report runs/moe-default.json runs/moe-tuned.json runs/moe-default-2.json --title 'Tuned vs default fused-MoE kernels: Qwen3-30B-A3B FP8 on 1× H100 SXM'