← Lab notes

Does selling tokens pay? Break-even math for an inference provider, from a measured H100

Using measured SLO throughput for Qwen3-30B-A3B on one H100 and live OpenRouter prices, we work out the break-even utilization for rented, owned and idle GPUs. Rented GPUs need 64–77% utilization to match the cheapest price; idle owned GPUs break even at 2%. Includes a calculator.

Companies are buying GPUs as fixed assets. What a GPU earns comes down to three numbers multiplied together:

revenue per GPU = utilization × efficiency (billable tokens per second) × price per token

We can measure efficiency. Prices are public on OpenRouter. So we can work out how much of the time a GPU has to be busy before selling tokens pays for it. The answer depends heavily on whether you rent the GPU, own it, or are filling time it would otherwise sit idle.

Inputs

Input Value Source
Capacity at SLO 9.5 requests/s per H100 SXM (34,207 per hour) measured: Qwen3-30B-A3B FP8, vLLM 0.31.0, p99 TTFT ≤ 2 s, p99 TPOT ≤ 50 ms
Request shape 1,695 input + 421 output tokens on average our chat-mix-v1 trace; output measured at the SLO boundary
Prices (in / out per 1M tokens) StreamLake $0.048 / $0.193 · SiliconFlow $0.09 / $0.30 · DeepInfra $0.12 / $0.50 · Alibaba $0.13 / $0.52 OpenRouter, Qwen3-30B-A3B family, 2026-10-07
Rented H100 SXM $3.56 (market lowest) · $4.19 (median) · $4.29 (Lambda) per GPU-hour our GPU price table, 2026-10-07
Owned, full cost $1.29 per GPU-hour (assumed) $250k 8-GPU server over 4 years ($0.89) + 0.95 kW × PUE 1.3 × $0.08/kWh ($0.10) + $0.30 operations
Owned, idle hours $0.10 per GPU-hour (assumed) energy only: the hardware is paid for whether or not it runs

Revenue per request is input tokens × input price + output tokens × output price. That's $0.00016 at StreamLake's price and $0.00041 at DeepInfra's. Input tokens are a large share of revenue on chat traffic.

Break-even utilization

The share of SLO capacity a GPU has to use, on average, to cover its cost:

Cost scenario $/GPU-hour at StreamLake's price at SiliconFlow's at DeepInfra's
Rent, market lowest $3.56 64% 37% 25%
Rent, Lambda $4.29 77% 45% 30%
Own, full cost $1.29 23% 14% 9%
Own, idle hours $0.10 2% 1% 1%

Profit per GPU-hour at 50% utilization:

Cost scenario at StreamLake's price at DeepInfra's
Rent, Lambda −$1.51 +$2.79
Own, full cost +$1.49 +$5.79
Own, idle hours +$2.68 +$6.98

What it says

  1. Renting GPUs to compete with the cheapest provider doesn't work at this SLO. You'd need 64–77% average utilization for the whole month, which a new provider is unlikely to see. Competing at mid-market prices needs 25–45%, which is reachable but not easy.
  2. Owning the hardware changes the picture. At full cost, 9–23% utilization pays for it.
  3. Idle hours are almost pure margin. An owned GPU sitting idle for its owner costs about the electricity. At 30% utilization of those hours, it breaks even at about 1/17 of the cheapest listed price. A company's private GPUs have roughly 550 idle hours a month (nights and weekends), so selling that time at even a heavily discounted price adds real revenue: about $1,470 per GPU per month at StreamLake's price and 50% utilization of the idle hours.

That last point is why we think idle enterprise GPUs are an underused supply. The hard parts aren't the economics. They're isolation, handing the GPU back instantly when the owner needs it, and certifying hardware that varies from machine to machine. We wrote about what breaks when you assemble these stacks.

Calculator

Revenue per request—
Break-even utilization—
Profit per GPU-hour—
Profit per GPU-month—
Cost per 1M output tokens—

Defaults: rented H100 at Lambda's price, our measured capacity for Qwen3-30B-A3B FP8 at p99 TTFT ≤ 2 s and TPOT ≤ 50 ms, and DeepInfra's OpenRouter price. For idle owned hardware, set the cost to your energy cost per hour and the hours to your idle hours.

What this doesn't cover

  • One model, one GPU, one SLO. Cheaper providers may run looser latency targets, newer GPUs (H200, B200) or more optimized stacks, so they get more throughput per GPU than we measured. "Renting can't match the cheapest price" holds for this configuration, not in general.
  • How much traffic a new provider actually gets. Routing on OpenRouter depends on price, latency and uptime. Utilization is an assumption here, not a measurement.
  • Platform fees, egress, idle capacity kept in reserve, and the cost of handing GPUs back on demand.
  • The owned-hardware costs are assumptions. Change them in the calculator.

The capacity number is the one that matters most, and it's the one that's hardest to get right. Ours comes from 10-minute windows at a strict SLO. A two-minute test overstated it by 14% (lab notes #2).