← Lab notes

H100 vs H200 for LLM inference: 18% more capacity, 33% higher price — and where the H200 pulls ahead

We ran the same model, engine and latency target on an H100 SXM and an H200 on RunPod. At a 50 ms-per-token SLO the H200 served 18% more requests for 33% more money, so each token cost 11% more. At a fixed load it generated tokens 36% faster, so under a tighter latency target, or with a bigger model or longer context, the H200 is likely to come out ahead. Measured numbers, cost per million tokens and the break-even price.

The H200 is an H100 with more and faster memory: 141 GB of HBM3e at 4.8 TB/s, against 80 GB at 3.35 TB/s, with the same compute. LLM token generation is mostly limited by memory bandwidth, so the obvious guess is that the H200 is about 40% faster.

We measured it. The answer depends on your model and your latency target.

Results for Qwen3-30B-A3B-FP8 on vLLM 0.31.0, one GPU, both on RunPod Secure Cloud on the same day:

  • At our standard SLO (p99 time to first token ≤ 2 s, p99 time per output token ≤ 50 ms), the H200 served 13.22 requests/s against the H100's 11.18: 18% more.
  • It cost 33% more per hour ($5.29 against $3.99), so each token cost 11% more: $0.271 per million output tokens, against $0.245.
  • At the same load, the H200 generated tokens 36% faster: a median of 16.1 ms per token against 25.1 ms at 8 requests/s. With a tighter latency target, a larger model or long contexts, that advantage turns into capacity, and the H200 is likely to win.

The setup

H100 SXM H200
Where RunPod Secure Cloud, Missouri RunPod Secure Cloud, Colorado
Listed price $3.99/h $5.29/h
GPU memory 80 GB 141 GB
Memory bandwidth (spec) 3.35 TB/s 4.8 TB/s
Power limit 700 W 700 W

Both machines ran the same container (vllm/vllm-openai:v0.31.0) with the same settings and the same workload: a chat-like mix with median ~1,700 input and ~420 output tokens, arriving at random. Bench searched for the highest request rate that still met the SLO and then repeated it twice. The methodology has the details.

On RunPod we had to set VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0 on both GPUs. The default FP8 kernel fails to start there, as we described in our three-platform test.

Before benchmarking, we checked both cards with tokenwatt-check, our open-source node checker. Both passed with full power limits and normal clocks, so the comparison is between healthy cards:

tokenwatt-check H100 SXM H200
BF16 matmul 682 TFLOPS 662 TFLOPS
Clock during matmul 1,359 MHz 1,462 MHz
Memory copy bandwidth 3.05 TB/s 4.3 TB/s (+41%)

The compute is the same, and the measured memory bandwidth is 41% higher, close to the spec ratio.

Capacity and cost per token

H100 SXM H200 H200 vs H100
SLO capacity 11.18 req/s 13.22 req/s +18%
Output tokens/s at capacity 4,526 5,430 +20%
GPU board power 607 W 675 W +11%
Output tokens per joule 7.45 8.05 +8%
Cost per million output tokens $0.245 $0.271 +11%
KV cache available 445,000 tokens 1,054,000 tokens 2.4×
GSM8K (50 items) 92% 92% —

The H200 is faster and more energy-efficient. At these prices, though, it isn't cheaper per token for this model.

Why only 18% more capacity from 41% more bandwidth

The model is small. Qwen3-30B-A3B is a mixture-of-experts model that activates about 3 billion parameters per token. Each decode step reads relatively few bytes, so a good share of each step goes to things that don't scale with memory bandwidth: attention compute, scheduling and sampling. Measured against its own bandwidth, the H200 ran at 45% utilization and the H100 at 56%.

The model doesn't need the memory. Even on the H100 there was room for 445,000 tokens of KV cache, far more than this workload uses at 11 requests per second. The H200's 2.4× larger cache sat mostly empty.

Where the H200 does win: latency

At the same load the H200 is much faster for each user:

At a fixed load H100: time per token (median / p99) H200: time per token (median / p99)
8 requests/s 25.1 / 31.3 ms 16.1 / 21.6 ms
10 requests/s 33.3 / 42.8 ms 23.3 / 31.5 ms

At 8 requests/s the H200 streams answers 36% faster. That gap matters as soon as the latency target is tighter than our 50 ms:

  • With a p99 target of about 30 ms per token (roughly 33 tokens/s per user, typical for agents or chat that should feel instant), the H100 already misses at 8 requests/s (p99 31.3 ms). The H200 meets it at 8 with room to spare (21.6 ms) and only just misses at 10 (31.5 ms). Under that target the H200 carries load the H100 can't. How much more, and whether that makes it cheaper per token, needs a direct measurement with the tighter SLO, which is our next test.
  • With a bigger model, the memory starts to matter. A 70B model in FP8 needs about 70 GB for weights. On an H100 that leaves almost nothing for KV cache. On an H200 it leaves about 60 GB.
  • With long contexts, for example document Q&A or RAG with 32k-token prompts, the KV cache fills up fast. 2.4× more room means more concurrent users before requests start queuing.

The break-even price

For this model at a 50 ms SLO, the H200 delivers 1.20× the tokens of an H100, so it's the better buy when it costs less than 1.2× the H100's price. At RunPod's $3.99 for an H100, that's an H200 at about $4.79 an hour or less.

H200 prices vary widely. On the day of this test the lowest on-demand H200 we tracked was below that threshold, and RunPod's $5.29 was above it. Check current H100 and H200 prices before you choose.

Which one?

  • Small models, or MoE models with few active parameters, under a relaxed latency target: H100. It's cheaper per token unless you find an H200 at less than 1.2× the H100 price.
  • Interactive products that need fast streaming, 70B-class models, or long contexts: H200. It's faster per user and has room for the KV cache; once the latency target or model size bites, it may well be cheaper per token too.
  • Either way: check the card you get. Our three-platform test found one H100 delivering 40% of normal clock speed. tokenwatt-check takes two minutes: pip install tokenwatt-check.

Limits

  • One machine of each, on one platform, on one day, in quick mode (60-second search windows, two 120-second repeats).
  • One model. The tighter-SLO and 70B conclusions follow from the latency and memory numbers above but haven't been measured as SLO capacity yet. That's our next test.
  • Both GPUs ran vLLM's non-FlashInfer FP8 kernel because of the RunPod startup issue. On other platforms vLLM uses the FlashInfer kernel by default.

Cost of this test: $3.45 ($2.64 for the H200 run, $0.81 for the H100 node check). The H100 benchmark comes from our earlier RunPod run with the same settings.