← Methodology

Glossary

Cost per token for LLM inference: the formula, and what we measured

What a million tokens costs on your own or rented GPUs depends on the GPU's price and on how many tokens it serves within your latency target. Here's the formula, the trap of using peak throughput, and the costs we measured on real hardware.

Cost per million output tokens is the most useful single number for comparing GPUs, clouds and serving setups, because it combines price and performance.

The formula

Cost per million output tokens = price per GPU-hour ÷ (output tokens per second × 3,600) × 1,000,000

Example: an H100 at $3.99 an hour serving 3,764 output tokens per second within our latency target costs $3.99 ÷ (3,764 × 3,600) × 10⁶ = $0.294 per million tokens.

Two choices in that formula decide whether the number means anything:

  1. Which tokens per second. Use the throughput the GPU sustains while meeting your latency target (SLO goodput), not its peak. At peak, users wait seconds for their first token. We measure capacity at p99 time to first token ≤ 2 s and p99 time per output token ≤ 50 ms.
  2. Utilization. The formula assumes the GPU runs at that capacity every hour you pay for. At 50% average utilization, the real cost per token doubles. Traffic that varies through the day is often the biggest cost factor of all.

Input tokens matter too: long prompts take GPU time to process, and our measured throughput already includes it, since the workload has realistic prompt lengths (median 1,024 input tokens).

What we measured

Cost per million output tokens at each GPU's capacity, at the cloud price on the day of the test:

GPU (price per hour) Model Cost per million output tokens
H100 SXM, RunPod ($3.99) Qwen3-30B-A3B FP8 $0.245 (quick) / $0.294 (full)
H100 SXM, Lambda ($4.29) Qwen3-30B-A3B FP8 $0.264
H200, RunPod ($5.29) Qwen3-30B-A3B FP8 $0.271
RTX 4090, RunPod ($0.89) Qwen3-8B FP8 $0.90
A100 SXM 80 GB, RunPod ($1.79) Qwen3-8B BF16 $1.06
L40S, RunPod ($1.09) Qwen3-8B BF16 $1.11
A40, RunPod ($0.59) Qwen3-8B BF16 $1.61

Sources: RunPod vs Vast.ai vs Lambda (which also covers marketplace hosts), H100 vs H200, SGLang vs vLLM, L40S vs A100 vs RTX 4090 vs A40. "Quick" and "full" are our measurement modes: full mode uses 10-minute windows and reads 15–20% lower capacity.

What moves cost per token most

  • The model. A 30B mixture-of-experts model on an H100 cost about $0.29 per million tokens, about 3× less than an 8B dense model on any cheaper GPU we tested.
  • The node you actually get. The same listing can hide a lowered power limit or a degraded card. One H100 we rented couldn't meet the latency target at any load, so its effective cost per token was infinite. A GPU health check takes two minutes.
  • Cheapest per hour isn't cheapest per token. The A40 was the cheapest GPU per hour and the most expensive per token.
  • Your latency target. A tighter target lowers capacity and raises cost: at 30 ms instead of 50 ms, H100 capacity fell by about 28%.
  • Utilization, as above.

For your own GPU prices, see our daily price index. To measure your nodes, we offer remote evaluations.

FAQ

How do I calculate the cost per million tokens of self-hosting an LLM?

Divide the GPU's price per hour by the output tokens per hour it serves within your latency target, then multiply by one million. Use throughput measured at your latency target, not the peak, and adjust for how much of the time the GPU is actually busy.

Is self-hosting cheaper than an API?

It depends on utilization. A GPU that's busy around the clock can be cheaper per token than many APIs; one that's idle half the time costs twice as much per token. Measure your real capacity and traffic before deciding.

Why do benchmarks report lower cost per token than I see in production?

Many benchmarks use peak throughput, ignore latency targets, and assume 100% utilization. Production traffic varies, and serving at peak throughput makes users wait.