← All reports

Report · 2026-10-10

vLLM 0.31.0 vs SGLang 0.5.21 on 1× H100 SXM, Qwen3-30B-A3B FP8

1× NVIDIA H100 80GB HBM3 Qwen3-30B-A3B-FP8 FP8 weights, BF16 KV cache vLLM varies by configuration TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Full-mode comparison on two RunPod Secure Cloud H100 SXM pods in the same datacenter (US-MO-1), measured at the same time on 2026-10-10: 10-minute windows, 5 repeats at the SLO boundary, GSM8K on 200 items × 3 seeds. Default settings for both engines; vLLM ran with VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0 (required on RunPod). Both nodes passed tokenwatt-check. Energy is GPU board power.

ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on vLLM 0.31.0 takes: with SGLang 0.5.21, 64 → 64 cards and 10% more GPU board energy per token.

Results

Each value is the median of 5 runs.

vLLM 0.31.0SGLang 0.5.21

SLO goodput

output tok/s per node · higher is better
vLLM 0.31.03,764vLLM 0.31.0: 3,764 output tok/s per node
SGLang 0.5.213,814SGLang 0.5.21: 3,814 output tok/s per node

Tokens per joule

tok/J, GPU board power · higher is better
vLLM 0.31.06.69vLLM 0.31.0: 6.69 tok/J, GPU board power
SGLang 0.5.216.06SGLang 0.5.21: 6.06 tok/J, GPU board power

Decode bandwidth utilization

% of rated HBM bandwidth
vLLM 0.31.052%vLLM 0.31.0: 52 % of rated HBM bandwidth
SGLang 0.5.2162%SGLang 0.5.21: 62 % of rated HBM bandwidth

Axis runs to 100% of rated bandwidth.

GPU board power

kW, steady state
vLLM 0.31.00.56vLLM 0.31.0: 0.56 kW, steady state
SGLang 0.5.210.63SGLang 0.5.21: 0.63 kW, steady state

GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

MetricvLLM 0.31.0SGLang 0.5.21
Engine buildvLLM 0.31.0 · KV BF16SGLang 0.5.21 · KV BF16
Decode bandwidth utilization (MBU)52%62%
GPU board power0.56 kW0.63 kW
SLO goodput3,764 tok/s3,814 tok/s
Tokens per joule6.696.06
TTFT p99401 ms249 ms
TPOT p9943.2 ms33.7 ms
SLO metyesyes
Accuracy (gsm8k-5shot)91.891.0
Accuracy Δ vs. referencereference-0.8 pts
Cards for equal goodput6464

Setup

Any change to these fields makes it a different report.

Accelerator
1× NVIDIA H100 80GB HBM3
Rated HBM BW
3.35 TB/s per GPU
Driver / runtime
580.126.09
Host
RunPod secure cloud pod,
Engine
vLLM varies by configuration
Kernel libraries
vllm/vllm-openai:v0.31.0
Model
Qwen3-30B-A3B-FP8 (MoE)
Quantization
FP8 weights, BF16 KV cache
Parallelism
1 replica × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
GPU board power sum (NVIDIA) · 1 Hz · 600 s steady state
Runs per config
5 (variance band ±2.1%)
Accuracy gate
gsm8k-5shot · max drop 1.0 pts

Not covered

Do not extrapolate this report to the following.

  • Whole-node power (no BMC access on this host)
  • Context lengths beyond the chat-mix-v1 trace

Reproduce

Run on the same hardware and software versions. Results should fall within ±2.1%.

shell
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'h100-vllm-full' --role baseline --out runs/h100-vllm-full.json
tokenwatt run --spec lambda-qwen3-30b-a3b-fp8-tp1-dp1.spec.yaml --label 'h100-sglang-full' --role other --out runs/h100-sglang-full.json

tokenwatt report runs/h100-vllm-full.json runs/h100-sglang-full.json --title 'vLLM 0.31.0 vs SGLang 0.5.21 on 1× H100 SXM, Qwen3-30B-A3B FP8'

Download raw JSON