Lab notes #6: We tuned vLLM's MoE kernels on an H100 for 2.5 hours. It made no measurable difference.
vLLM warns "Using default MoE config. Performance might be sub-optimal!" for Qwen3-30B-A3B-FP8 on H100. We ran vLLM's own tuner for 159 minutes, confirmed the tuned config loaded, and measured default, tuned and default again on the same GPU. Kernels got up to 7% faster; end-to-end capacity didn't change, and the latency change was within run-to-run drift.
When vLLM serves a mixture-of-experts model and has no tuned kernel settings for your GPU and model shape, it says so at startup:
WARNING [fused_moe.py] Using default MoE config. Performance might be sub-optimal!
Config file not found at .../fused_moe/configs/E=128,N=768,device_name=NVIDIA_H100_80GB_HBM3,dtype=fp8_w8a8,block_shape=[128,128].json
We saw it while serving Qwen3-30B-A3B-FP8 for six hours on an H100 and called it a free optimization. It sounds like one: vLLM ships a tuner, and the tuned file is a drop-in.
So we tested it.
Result:
- The tuned kernels were up to 7% faster in isolation; at most batch sizes they were 0–3% faster.
- End to end, capacity within the SLO didn't change.
- Time per output token improved by 2–5%, but two runs of the same default config differed by 2.6%.
- Accuracy was identical, and power went up slightly.
For this model on this GPU, the warning can be ignored.
Method
All on one Lambda Cloud H100 SXM (80 GB HBM3), driver 580.105.08, vLLM 0.31.0, in a single session so the hardware didn't change:
- Default: Bench SLO sweep with vLLM's built-in heuristic config.
- Tune: vLLM's own
benchmarks/kernels/benchmark_moe.py --tunefor nine batch sizes (1 to 4,096). For each batch size it searches 1,920 Triton tile configurations. Kernel times were recorded before and after. - Tuned: the same sweep, with the tuned file loaded through
VLLM_TUNED_CONFIG_FOLDER. vLLM loggedUsing configuration from /tuned-moe/E=128,N=768,...json for MoE layer. - Default again: the same sweep without the file, to measure drift.
Each sweep: SLO p99 time to first token ≤ 2 s and p99 time per output token ≤ 50 ms; chat-mix traffic; search from 8 req/s in steps of 1.25×, then bisection; two 120-second repeats at the boundary; a GSM8K gate on 50 items.
Kernel times
Fused-MoE kernel time per call, as measured by vLLM's benchmark:
| Batch size (tokens) | Default | Tuned | Speedup |
|---|---|---|---|
| 1 | 27.9 µs | 27.6 µs | 1.01× |
| 16 | 158.0 µs | 158.3 µs | 1.00× |
| 64 | 230.5 µs | 229.6 µs | 1.00× |
| 128 | 255.8 µs | 238.5 µs | 1.07× |
| 256 | 266.5 µs | 252.4 µs | 1.06× |
| 512 | 284.2 µs | 280.5 µs | 1.01× |
| 1,024 | 344.8 µs | 334.8 µs | 1.03× |
| 2,048 | 506.5 µs | 496.7 µs | 1.02× |
| 4,096 | 864.0 µs | 872.3 µs | 0.99× |
vLLM's default heuristic for FP8 block-quantized weights is already close to the best of the 1,920 configurations the tuner tries. The biggest gain is at batch sizes 128–256, roughly the number of tokens per step during decode under load.
End to end
| Default | Tuned | Default again | |
|---|---|---|---|
| Highest rate within SLO | 11.18 req/s | 11.18 req/s | 11.18 req/s |
| Output tokens/s at that rate | 4,515 | 4,523 | 4,516 |
| Time per output token, p50 | 33.7 ms | 31.9 ms | 32.8 ms |
| Time per output token, p99 | 44.2 ms | 42.3 ms | 43.0 ms |
| Time to first token, p99 | 338 ms | 333 ms | 333 ms |
| GPU board power | 623 W | 641 W | 631 W |
| Output tokens per joule | 7.25 | 7.05 | 7.15 |
| GSM8K (50 items) | 96.0% | 96.0% | 96.0% |
Two things stand out.
Capacity can't move by less than the search resolution. All three sweeps landed on the same step, 11.18 req/s; the next step up, 11.50, failed every time. Under open-loop load, output throughput is set by the arrival rate, so tokens per second is the same whenever the rate is the same. A gain under about 3% can't show up in this number. That's a limit of the method: measuring smaller changes needs a fixed-rate comparison of latency, which is how we'll test the next optimizations.
The latency gain is about the size of the drift. The tuned run was 1.8 ms faster per token than the first default run and 0.9 ms faster than the second. The two default runs differed by 0.9 ms with nothing changed. A real effect of a few percent is possible, but this experiment can't separate it from noise.
Why so little?
A mixture-of-experts kernel is one part of each decode step. In Qwen3-30B-A3B only about 3B of the 30B parameters are active per token, and at these batch sizes much of each step goes to attention over the KV cache, other layers, sampling and scheduling. A 6% faster MoE kernel at the most common batch size might make the whole step 1–2% faster. That's roughly what we saw.
What this means
- For this model on H100, ignore the warning. If you serve another MoE shape or GPU, measure first: the tuner took 159 minutes of GPU time here.
- Tuning has to be measured end to end. A 7% kernel win sounds good, but it disappears in the full step. We'll report end-to-end, accuracy-gated numbers for every optimization we try.
- The bigger levers are elsewhere. In our runs the large effects came from KV-cache precision (Lab notes #1), the admission limit (Lab notes #5), and batching and prefill settings. These come next.
Cost
| Instance time | 234 minutes on gpu_1x_h100_sxm5, $16.74 |
| Failed starts | two instances in regions still on driver 570, caught by preflight in seconds: $1.09 |
Reproduce
From the bench directory:
scripts/lambda/cloud.py session --ssh-key <key> --type gpu_1x_h100_sxm5 --max-minutes 300 --yes \
--run "MODEL=Qwen/Qwen3-30B-A3B-FP8 MIN_DRIVER=580 MODE=quick SWEEP_START=8 SWEEP_GROWTH=1.25 LABEL=moe-default" \
--run "MODEL=Qwen/Qwen3-30B-A3B-FP8 MIN_DRIVER=580 TUNE_MOE=1 TUNE_MINUTES=150 LABEL=moe-tuning" \
--run "MODEL=Qwen/Qwen3-30B-A3B-FP8 MIN_DRIVER=580 MODE=quick SWEEP_START=8 SWEEP_GROWTH=1.25 LABEL=moe-tuned TUNED_MOE_DIR=~/moe-tuned" \
--run "MODEL=Qwen/Qwen3-30B-A3B-FP8 MIN_DRIVER=580 MODE=quick SWEEP_START=8 SWEEP_GROWTH=1.25 LABEL=moe-default-2"