RunPod vs Vast.ai vs Lambda: we benchmarked four rented H100s and found a 2.9× spread
Four H100 SXM GPUs rented by the hour on Lambda, RunPod and Vast.ai (a datacenter host and a community host), each serving Qwen3-30B-A3B-FP8 on vLLM 0.31.0 under the same latency SLO. Cost per million tokens ranged from $0.199 (Vast.ai datacenter) to $0.264 (Lambda); the cheapest host per hour couldn't meet the SLO at all, running at 40% of normal clock speed; both Vast hosts capped GPU power, and the official vLLM image failed to start on RunPod until we changed one kernel.
Disclosure: this post contains referral links to RunPod and Vast.ai. If you sign up through them we may earn a small commission. It doesn't change what we measured, and we paid for every GPU-hour below ourselves.
"An H100 for $2 an hour" and "an H100 for $4 an hour" sound like the same product at different prices. For LLM inference they often aren't. The number that matters is what a million served tokens costs once the GPU has to meet a latency target, and the listing doesn't tell you that.
So we rented four H100 SXM GPUs by the hour and ran the same benchmark on each:
- one from Lambda;
- one from RunPod's Secure Cloud (datacenter);
- two from Vast.ai: a verified datacenter host and a community host.
Results:
- The cheapest per token was the Vast.ai datacenter host, at $0.199 per million output tokens, 25% below Lambda. It got there despite its GPU being power-capped to 525 W.
- RunPod matched Lambda's speed for 7% less. But the official vLLM image would not start on RunPod with default settings, in two regions, until we disabled one FP8 kernel.
- The cheapest per hour, the Vast.ai community host at $2.18, was the worst deal by far. Under load its GPU ran at 787 MHz instead of about 1,900, every token took 2.9× longer to generate, and it couldn't meet our latency target even at our lowest test load.
The test
Every GPU ran the same thing:
- Model: Qwen3-30B-A3B-FP8.
- Engine: vLLM 0.31.0 from the official
vllm/vllm-openaiimage. - Settings: 12,288-token context, one GPU, no tuning.
- Load: Poisson arrivals of a chat-like mix of prompts (median ~1,700 input / ~420 output tokens).
- Measurement: we searched for the highest request rate at which p99 time to first token stays under 2 s and p99 time per output token under 50 ms. That rate, and the tokens per second it delivers, is the GPU's SLO capacity. The method is described here.
- Accuracy check: 50 GSM8K questions on each machine, so a fast but broken setup can't win.
This is a quick-mode run: 60-second search windows and two 120-second repeats at the boundary. Quick mode is good for comparing machines measured the same way. It isn't a substitute for a long run. Each platform is one machine on one day, and other machines on the same platform can behave differently. That variation is part of the point of this post.
Results
| Lambda | RunPod | Vast DC | Vast community | |
|---|---|---|---|---|
| Listed price | $4.29/h | $3.99/h | $2.82/h | $2.18/h |
| Type | cloud | Secure Cloud | verified datacenter | verified community |
| Location | US | US | Czechia | Germany |
| Power limit (max 700 W) | 700 W | 700 W | 525 W | 500 W |
| SLO capacity | 11.18 req/s | 11.18 req/s | 9.72 req/s | below 8 req/s |
| Tokens/s at capacity | 4,515 | 4,526 | 3,945 | — |
| Tokens per joule | 7.25 | 7.45 | 7.67 | — |
| $ per million tokens | $0.264 | $0.245 | $0.199 | — (missed SLO) |
| GSM8K (50 items) | 96% | 92%* | 96% | 96% |
* RunPod ran with one kernel changed; see below. Two questions out of 50 is within the sampling error of a 50-item check.
Cost per million tokens is the hourly price divided by the tokens per hour served within the latency target. That's the price of usable output. The community host never met the target, so it has no SLO capacity. Its true cost per usable token is higher than the hourly price suggests, and for an interactive service it isn't usable at all.
Same load, very different GPUs
At a fixed 8 requests per second, a load every machine received:
| Lambda | RunPod | Vast DC | Vast community | |
|---|---|---|---|---|
| GPU clock under load | 1,961 MHz | 1,966 MHz | 1,812 MHz | 787 MHz |
| GPU board power | 574 W | 544 W | 511 W | 497 W |
| Time per output token, median | 22.5 ms | 25.1 ms | 26.0 ms | 73.0 ms |
| Time to first token, p99 | 0.26 s | 0.32 s | 0.35 s | 12.1 s |
The Vast.ai datacenter host and the community host drew almost the same power, 511 W and 497 W. One ran at 1,812 MHz and the other at 787 MHz. A 500 W cap alone doesn't explain that: the datacenter host, capped at 525 W and drawing 511 W, held 1,812 MHz. That card was getting less than half the work out of each watt, whether from age, power delivery or something else we couldn't see from inside the container. Its listing showed nothing unusual: H100 SXM, verified, 99.5% reliability.
That's the problem with renting from a marketplace. The listing tells you the GPU model, not the GPU.
Things the price page doesn't tell you
Power caps. Both Vast.ai hosts had lowered the GPU's power limit, to 525 W and 500 W out of 700 W. It saves the host electricity and heat. On the datacenter host it cost 13% of capacity and was more than offset by the price. On the community host it came with a much bigger problem. You can see the limit once you're connected:
nvidia-smi --query-gpu=power.limit,power.max_limit,clocks.sm --format=csv
Run it under load, not idle, because clocks only drop when the GPU is busy.
Time to ready. Every platform bills from boot, including the minutes spent downloading the model:
| Lambda | RunPod | Vast DC | Vast community | |
|---|---|---|---|---|
| Order to running | 2.0 min | 0.4 min | 0.9 min | 0.7 min |
| Model ready (download + load) | ~5 min | 2.3 min | 4.9 min | 13.3 min |
| Measured download speed | — | 9.5 Gbit/s | — | 2.1 Gbit/s |
On the community host, loading 29 GB of weights from its slower disk took almost ten minutes. For a one-hour job, that's a sixth of the bill.
What you're actually charged. On Vast.ai our account balance went down by more than the listed hourly price times the minutes the instance ran: $2.29 against $1.09 for the datacenter host, and $0.91 against $0.66 for the community host. The difference is most likely storage and billing granularity. Budget for it on short jobs.
RunPod: the official image didn't start
On RunPod, vllm/vllm-openai:v0.31.0 with default settings crashed while capturing CUDA graphs. We saw it on three pods, including ones in India and Missouri:
[TensorRT-LLM][ERROR] Warning: Failed to copy kernel files to cache: filesystem error: cannot rename: No such file or directory
RuntimeError: Assertion failed: !cubin.empty() || isPathValid(path_)
vLLM picks FlashInfer's DeepGEMM kernel for FP8 block-scaled matrix multiplies, and that kernel compiles at run time. The image has no nvcc, and on RunPod the compiled kernel never arrived in its cache. The same image picked the same kernel on Lambda and on both Vast.ai hosts and worked fine. We couldn't pin down what's different about RunPod's environment.
The workaround is one environment variable, which makes vLLM use a different FP8 kernel:
VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0
With it, RunPod matched Lambda: the same 11.18 req/s and 4,526 tokens per second. If you deploy FP8 models with vLLM on RunPod and the engine dies at startup, try this first.
Each failed attempt cost us about $0.40. A failure like this is cheap when you know what to look for and expensive when you don't.
Which should you rent?
- For a production endpoint where latency matters: RunPod Secure Cloud or Lambda. Both delivered full H100 performance. RunPod was 7% cheaper per token and started fastest, once the kernel issue was out of the way. RunPod
- For the lowest cost per token: a verified Vast.ai datacenter host can be 25% cheaper per token than Lambda even with a power cap, but benchmark the specific host before you commit. Vast.ai
- Community hosts: treat the price as a lottery ticket. Some will be fine. The one we got delivered 40% of normal clock speed and missed every latency target.
Whatever you rent, spend the first ten minutes measuring it. Check the power limit and the clocks under load, and push real traffic at it before you put users on it. That's also why we're working on a way to certify individual nodes rather than GPU models.
Cost of this test
$5.95 in total: $3.20 on Vast.ai and $2.74 on RunPod, including three failed RunPod starts. The Lambda baseline came from our MoE tuning run, with the same settings and method.
Reproduce
From the bench directory (Vast.ai uses its Python SDK; RunPod uses its REST API):
uv run --with vastai python scripts/lambda/vast_session.py --gpu H100_SXM --pick datacenter --ssh-key <key> \
--run "MODE=quick MODEL=Qwen/Qwen3-30B-A3B-FP8 SWEEP_START=8 SWEEP_GROWTH=1.25 LABEL=vast-h100" --yes
python3 scripts/lambda/runpod_session.py --cloud SECURE --country US --ssh-key <key> \
--run "MODE=quick MODEL=Qwen/Qwen3-30B-A3B-FP8 SWEEP_START=8 SWEEP_GROWTH=1.25 LABEL=runpod-h100 VLLM_ENV=VLLM_BLOCKSCALE_FP8_GEMM_FLASHINFER=0" --yes
Both scripts record the node's power limit, clocks, disk and download speed, and time to ready, and they always delete the machine at the end.