← All reports

Report · 2026-10-08

Tuned vs default fused-MoE kernels: Qwen3-30B-A3B FP8 on 1× H100 SXM

1× NVIDIA H100 80GB HBM3 Qwen3-30B-A3B-FP8 FP8 weights, BF16 KV cache vLLM 0.31.0 TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Lambda Cloud 1× H100 SXM, 2026-10-08: vLLM 0.31.0 with its default fused-MoE Triton configs, then with configs tuned for this GPU and model (2.5 hours of kernel tuning), then default again to bound drift. Quick mode. The tuned kernels made no measurable difference to SLO goodput.

ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on default kernels takes: with tuned kernels, 64 → 64 cards and 3% more GPU board energy per token; with default kernels (repeat), 64 → 64 cards and 1% more GPU board energy per token.

Results

Each value is the median of 2 runs.

default kernelstuned kernelsdefault kernels (repeat)

SLO goodput

output tok/s per node · higher is better
default kernels4,515default kernels: 4,515 output tok/s per node
tuned kernels4,523tuned kernels: 4,523 output tok/s per node
default kernels (repeat)4,516default kernels (repeat): 4,516 output tok/s per node

Tokens per joule

tok/J, GPU board power · higher is better
default kernels7.25default kernels: 7.25 tok/J, GPU board power
tuned kernels7.06tuned kernels: 7.06 tok/J, GPU board power
default kernels (repeat)7.16default kernels (repeat): 7.16 tok/J, GPU board power

Decode bandwidth utilization

% of rated HBM bandwidth
default kernels56%default kernels: 56 % of rated HBM bandwidth
tuned kernels58%tuned kernels: 58 % of rated HBM bandwidth
default kernels (repeat)57%default kernels (repeat): 57 % of rated HBM bandwidth

Axis runs to 100% of rated bandwidth.

GPU board power

kW, steady state
default kernels0.62default kernels: 0.62 kW, steady state
tuned kernels0.64tuned kernels: 0.64 kW, steady state
default kernels (repeat)0.63default kernels (repeat): 0.63 kW, steady state

GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

Metricdefault kernelstuned kernelsdefault kernels (repeat)
Engine buildvLLM 0.31.0 · KV BF16vLLM 0.31.0 · KV BF16vLLM 0.31.0 · KV BF16
Decode bandwidth utilization (MBU)56%58%57%
GPU board power0.62 kW0.64 kW0.63 kW
SLO goodput4,515 tok/s4,523 tok/s4,516 tok/s
Tokens per joule7.257.067.16
TTFT p99338 ms333 ms333 ms
TPOT p9944.2 ms42.3 ms43.0 ms
SLO metyesyesyes
Accuracy (gsm8k-5shot)96.096.096.0
Accuracy Δ vs. referencereference0.0 pts0.0 pts
Cards for equal goodput646464

Setup

Any change to these fields makes it a different report.

Accelerator
1× NVIDIA H100 80GB HBM3
Rated HBM BW
3.35 TB/s per GPU
Driver / runtime
580.105.08
Host
Lambda Cloud instance
Engine
vLLM 0.31.0
Kernel libraries
vllm/vllm-openai:latest
Model
Qwen3-30B-A3B-FP8 (MoE)
Quantization
FP8 weights, BF16 KV cache
Parallelism
1 replica × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
Runs per config
2 (variance band ±0.7%)
Accuracy gate
gsm8k-5shot · max drop 1.0 pts

Not covered

Do not extrapolate this report to the following.

  • Whole-node power (no BMC access on this host)
  • Context lengths beyond the chat-mix-v1 trace
  • Quick mode: short windows, few repeats — validation only, not for publishing

Reproduce

Run on the same hardware and software versions. Results should fall within ±0.7%.

shell
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-default' --role baseline --out runs/moe-default.json
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-tuned' --role tuned --out runs/moe-tuned.json
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'moe-default-2' --role other --out runs/moe-default-2.json

tokenwatt report runs/moe-default.json runs/moe-tuned.json runs/moe-default-2.json --title 'Tuned vs default fused-MoE kernels: Qwen3-30B-A3B FP8 on 1× H100 SXM'

Download raw JSON