Lab notes

What we measured, and how.

Runs on real hardware, written up with the setup, the numbers, what surprised us and the commands to reproduce it.

4 min read

Does selling tokens pay? Break-even math for an inference provider, from a measured H100

Using measured SLO throughput for Qwen3-30B-A3B on one H100 and live OpenRouter prices, we work out the break-even utilization for rented, owned and idle GPUs. Rented GPUs need 64–77% utilization to match the cheapest price; idle owned GPUs break even at 2%. Includes a calculator.

inference economicsOpenRouterH100
4 min read

How many GPUs does an internal ChatGPT need? Fewer than you think — and that's the cost problem

Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.

private inferencecapacity planningH100
4 min read

Lab notes #2: Short benchmark windows overstated H100 goodput by 14%

Full-mode rerun of Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. With 10-minute windows the SLO boundary fell from 11.3 to 9.5 req/s and goodput from 4,572 to 3,996 tok/s; the FP8 KV-cache accuracy drop shrank from 8 to 4 points on a 600-sample suite.

H100vLLMmethodology
4 min read

Lab notes #3: Nine providers, one Llama 3.3 70B — same accuracy, 18× different speed

We audited every OpenRouter provider of Llama 3.3 70B Instruct with the same GSM8K suite and latency probe. Accuracy was statistically indistinguishable (94.0–98.0%), decode speed ranged from 19 to 347 tokens/s, and output prices varied 7×. A naive first pass wrongly scored two providers at 0% and 44%.

OpenRouterLlama 3.3 70Binference providers
4 min read

Lab notes #4: Four things that broke a two-node GPU Kubernetes cluster before vLLM would serve

Validating a private-inference stack — k3s, NVIDIA GPU Operator, vLLM, DCGM — on two cloud A10s took five attempts. A CUDA 13 image on a 570 driver, broken -cu129 image variants, a CDI hook missing from the host toolkit, and a Service named vllm. Every failure was a version or naming mismatch between open-source parts, and most left no logs.

KubernetesGPU OperatorvLLM
3 min read

Billable tokens per watt: why we're building a vendor-neutral ruler for LLM inference

Power, not GPUs, is now the constraint on AI data centers. We explain the four numbers TokenWatt Bench reports — SLO goodput, tokens per joule, decode bandwidth utilization and an accuracy gate — and why each comparison includes the vendor's own best configuration.

methodologyinferenceAMD Instinct
5 min read

Lab notes #1: FP8 KV cache on an H100 — 11% more goodput, and it failed our accuracy gate

First real-hardware run of TokenWatt Bench. Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. FP8 KV cache doubled KV capacity and raised SLO goodput 10.6% and tokens per joule 14%, but GSM8K accuracy dropped 8 points with uncalibrated KV scales.

H100vLLMFP8
4 min read

How to measure decode bandwidth utilization without hardware counters

A vendor-neutral way to compute memory-bandwidth utilization (MBU) for LLM decode from model geometry and client-side timing — including MoE expert routing and MLA KV caches — and how we validated it against a simulated roofline engine.

MBUdecodeMoE