What we measured, and how.
Runs on real hardware, written up with the setup, the numbers, what surprised us and the commands to reproduce it.
Does selling tokens pay? Break-even math for an inference provider, from a measured H100
Using measured SLO throughput for Qwen3-30B-A3B on one H100 and live OpenRouter prices, we work out the break-even utilization for rented, owned and idle GPUs. Rented GPUs need 64–77% utilization to match the cheapest price; idle owned GPUs break even at 2%. Includes a calculator.
How many GPUs does an internal ChatGPT need? Fewer than you think — and that's the cost problem
Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.
Lab notes #2: Short benchmark windows overstated H100 goodput by 14%
Full-mode rerun of Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. With 10-minute windows the SLO boundary fell from 11.3 to 9.5 req/s and goodput from 4,572 to 3,996 tok/s; the FP8 KV-cache accuracy drop shrank from 8 to 4 points on a 600-sample suite.
Lab notes #3: Nine providers, one Llama 3.3 70B — same accuracy, 18× different speed
We audited every OpenRouter provider of Llama 3.3 70B Instruct with the same GSM8K suite and latency probe. Accuracy was statistically indistinguishable (94.0–98.0%), decode speed ranged from 19 to 347 tokens/s, and output prices varied 7×. A naive first pass wrongly scored two providers at 0% and 44%.
Lab notes #4: Four things that broke a two-node GPU Kubernetes cluster before vLLM would serve
Validating a private-inference stack — k3s, NVIDIA GPU Operator, vLLM, DCGM — on two cloud A10s took five attempts. A CUDA 13 image on a 570 driver, broken -cu129 image variants, a CDI hook missing from the host toolkit, and a Service named vllm. Every failure was a version or naming mismatch between open-source parts, and most left no logs.
Billable tokens per watt: why we're building a vendor-neutral ruler for LLM inference
Power, not GPUs, is now the constraint on AI data centers. We explain the four numbers TokenWatt Bench reports — SLO goodput, tokens per joule, decode bandwidth utilization and an accuracy gate — and why each comparison includes the vendor's own best configuration.
Lab notes #1: FP8 KV cache on an H100 — 11% more goodput, and it failed our accuracy gate
First real-hardware run of TokenWatt Bench. Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. FP8 KV cache doubled KV capacity and raised SLO goodput 10.6% and tokens per joule 14%, but GSM8K accuracy dropped 8 points with uncalibrated KV scales.
How to measure decode bandwidth utilization without hardware counters
A vendor-neutral way to compute memory-bandwidth utilization (MBU) for LLM decode from model geometry and client-side timing — including MoE expert routing and MLA KV caches — and how we validated it against a simulated roofline engine.