vLLM vs SGLang on an H100: same capacity at 50 ms, but SGLang streams a third faster and leads by ~20% at 30 ms
Two H100s in the same RunPod datacenter, one serving Qwen3-30B-A3B-FP8 with vLLM 0.31.0 and the other with SGLang 0.5.21, measured in full mode (10-minute windows, five repeats, 600 accuracy questions). Under a 50 ms per-token target, both served about 9 requests/s at $0.29 per million tokens. At the same load, SGLang returned the first token 3× sooner and generated tokens about 30% faster, so under a 30 ms target it served about 21% more requests. Accuracy was the same.
After our H100 vs H200 comparison, a reader asked us to add SGLang. vLLM and SGLang are the two most widely used open-source LLM serving engines, and many teams have to choose between them before they deploy. So we ran both on the same kind of GPU, in the same datacenter, at the same time, with the same workload.
Short answer: SGLang is faster per user. Whether it also serves more users depends on your latency target.
- At the same load, SGLang is much faster. At 9 requests/s, it returned the first token in a median of 67 ms against vLLM's 196 ms, and generated each token in 23 ms against 34 ms.
- Under our standard 50 ms per-token target, both served the same number of users: 9.07 requests/s for SGLang, 8.92 for vLLM, about $0.29 per million output tokens either way.
- Under a stricter 30 ms target, SGLang served about 21% more. In vLLM's case the 30 ms limit is what caps capacity, and SGLang's faster token generation leaves more room under it.
- Accuracy was the same: 91–92% on 600 GSM8K questions for both.
Updated 2026-10-10 with full-mode results. Our first, quick-mode run had vLLM 9% ahead at 50 ms. That didn't hold up under longer measurement windows. Details below.
The setup
| vLLM | SGLang | |
|---|---|---|
| Version | 0.31.0 (vllm/vllm-openai:v0.31.0) |
0.5.21 (lmsysorg/sglang:v0.5.21-cu130) |
| GPU | H100 SXM 80 GB, RunPod Secure Cloud | H100 SXM 80 GB, RunPod Secure Cloud |
| Datacenter | US-MO-1 | US-MO-1 |
| Price | $3.99/h | $3.99/h |
| Settings | defaults¹ | defaults |
| Node check (tokenwatt-check) | PASS: 1,367 MHz under load, 681 TFLOPS, 3.05 TB/s | PASS: 1,381 MHz under load, 683 TFLOPS, 3.04 TB/s |
The node check confirms both cards were healthy and equivalent, so the differences below come from the engines.
The model was Qwen3-30B-A3B-FP8, a mixture-of-experts model with about 3 billion active parameters. The workload was a chat-like mix (median ~1,700 input and ~420 output tokens) with requests arriving at random. For each engine and target, TokenWatt Bench searched for the highest request rate where the p99 time to first token (TTFT) stayed under 2 s and the p99 time per output token (TPOT) stayed under the target. It then had to hold that rate for five 10-minute repeats; if any repeat missed, it backed off 10% and started again. Accuracy was checked on 200 GSM8K questions with three seeds. The methodology has the details.
¹ On RunPod, vLLM's default FP8 kernel fails to start, so vLLM ran with VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0 (why). On Lambda, where the default kernel works, vLLM measured almost the same as on RunPod in our earlier tests, so this setting doesn't seem to hold vLLM back.
Capacity and cost per token
| vLLM 0.31.0 | SGLang 0.5.21 | SGLang vs vLLM | |
|---|---|---|---|
| 50 ms target: capacity | 8.92 req/s | 9.07 req/s | +2% (a tie) |
| Output tokens/s | 3,764 | 3,814 | |
| Cost per million output tokens | $0.295 | $0.291 | −1% |
| 30 ms target: capacity | 6.49 req/s | ~7.9 req/s² | ~+21% |
| Output tokens/s | 2,749 | ~3,250² | |
| Cost per million output tokens | $0.403 | ~$0.34² | ~−15% |
| Accuracy (GSM8K, 200 questions × 3 seeds) | 91.8% / 92.2% | 91.0% / 91.5% | same within ±1 point |
² SGLang's 30 ms run hit our 4-hour budget after passing the first of its five confirmation repeats at 7.86 requests/s. In the search phase its limit was 8.74 requests/s, against vLLM's 7.21: the same 21% gap. Treat the SGLang 30 ms row as provisional.
Same load, different experience
Here's what each user experienced at about the same request rate:
| At 9.0 req/s | vLLM | SGLang |
|---|---|---|
| Time to first token, median / p99 | 196 / 403 ms | 67 / 241 ms |
| Time per output token, median / p99 | 33.9 / 48.6 ms | 23.1 / 31.7 ms |
At 9 requests/s, vLLM is already close to the 50 ms per-token limit (p99 48.6 ms), while SGLang has plenty of room left (31.7 ms). So why does SGLang not serve more under the 50 ms target?
Because SGLang hits a different limit. As the rate rises past about 11 requests/s, SGLang still streams tokens within 50 ms, but new requests start to queue: at 11.2 requests/s its p99 time to first token reached 1.96 s, and at 11.9 the median was 3.4 s. The 2-second first-token limit stops it at about the same capacity where vLLM hits the per-token limit. Two engines, two different bottlenecks, the same answer.
Under a 30 ms target, the per-token limit binds first for both engines, and that's where SGLang's faster token generation turns into more capacity.
Which one?
- Interactive products (chat, agents, coding assistants): SGLang. Users see the first token about 3× sooner and text streams about a third faster at the same load. Under a tight per-token target it also serves about 21% more users per GPU.
- Throughput-oriented serving with a relaxed latency target: either. At 50 ms both delivered the same capacity and cost per token on this model.
- Running close to capacity: watch SGLang's queue. Past its limit, its time to first token rose steeply. Leave headroom, or cap concurrency at the gateway.
- Either way, measure your own workload. We ran defaults. Your model, prompt lengths and settings can change the ranking.
What changed from our first run
Our first run used quick mode (1-minute search windows, two 2-minute repeats, 50 accuracy questions). It showed vLLM 9% ahead at 50 ms and SGLang 2–4 points lower in accuracy. Neither held up in full mode:
| Quick mode | Full mode | |
|---|---|---|
| 50 ms capacity, vLLM / SGLang | 11.31 / 10.37 req/s | 8.92 / 9.07 req/s |
| 30 ms capacity, vLLM / SGLang | 6.60 / 8.00 req/s | 6.49 / ~7.9 req/s |
| Accuracy, vLLM / SGLang | 92% / 88–90% (50 questions) | 92% / 91% (600 questions) |
Ten-minute windows catch more of the rare slow requests that push p99 latency over the line, so full-mode capacities run lower, and they're the ones to plan with. The 30 ms result was consistent in both modes. The 50 ms gap and the accuracy gap were noise from short windows and a small question set. This is why we publish quick-mode numbers only with a warning label.
Limits
- One machine per engine, one model (a small mixture-of-experts model), default settings. Dense or larger models may behave differently.
- SGLang's 30 ms capacity is provisional: one of five confirmation repeats passed before our time budget ran out.
- Both engines ran their default scheduling. SGLang's queueing behavior near saturation may be tunable.
Cost of these tests: $39.35 ($15.48 and $15.97 for the two full-mode pods, $7.90 for the earlier quick-mode run, including a failed first SGLang start).