Billable tokens per watt: why we're building a vendor-neutral ruler for LLM inference
Power, not GPUs, is now the constraint on AI data centers. We explain the four numbers TokenWatt Bench reports — SLO goodput, tokens per joule, decode bandwidth utilization and an accuracy gate — and why each comparison includes the vendor's own best configuration.
New AI capacity in the US now waits on power more often than on accelerators. Interconnection queues, transformers and switchgear set the pace of buildout. Once a site is energized, every watt that doesn't turn into a billable token is the most expensive kind of idle.
That changes the question operators should ask about inference. "How fast is this GPU?" matters less than "how many billable tokens does this node produce per joule, at the latency my users will accept?" The answer depends far more on software than spec sheets suggest — and it moves with every engine, driver and kernel-library release.
We're building TokenWatt Bench to answer that question the same way on every vendor's hardware.
Spec sheets aren't SLOs
Peak tokens per second at unbounded latency is easy to publish and hard to use. A serving fleet is sized by how much traffic it can carry while time-to-first-token (TTFT) and time-per-output-token (TPOT) stay inside targets. Two systems with the same peak throughput can differ by a large factor at a 50 ms p99 TPOT.
So Bench measures throughput the way it gets billed:
- SLO goodput — output tokens per second at the highest request rate where p99 TTFT and p99 TPOT both stay inside the SLO. Failed or timed-out requests count against the SLO as infinite latency; they aren't dropped as missing data.
- Tokens per joule — goodput divided by mean power over exactly the same measurement window. Power comes from the BMC (Redfish or IPMI DCMI) or a metered PDU. When only GPU board power is available, as on most cloud VMs, the report says so. Nameplate TDP is never used.
- Decode bandwidth utilization (MBU) — the bytes the model must read from HBM per second, divided by the rated bandwidth. It's the clearest signal of how much headroom the software stack leaves on the table. We derive it from model geometry and measured timing, not hardware counters, so the definition is identical on AMD and NVIDIA. How that works.
- Accuracy gate — a fixed evaluation set, fixed seeds, and a pass threshold written down before the first run. A configuration that is faster but less accurate doesn't count.
Lock everything before the first run
A Bench measurement is defined by a spec file: hardware, engine, model, quantization, workload trace, SLO, accuracy suite and power source. Configurations can only be compared if they share the same spec. The tool refuses to combine runs that don't.
The workload trace matters as much as the SLO. SLO goodput is very sensitive to the request-length distribution, so the trace is part of the spec, and prompts are random token IDs with exact lengths so prefix caching can't flatter a result.
The vendor's best is the bar
Every comparison we publish has at least three columns, measured on the same node in the same session:
| Column | What it is |
|---|---|
| Production | the image and flags running in production today |
| Vendor latest | the hardware vendor's latest recommended image and settings |
| Candidate | the change under evaluation |
The middle column keeps everyone honest, us included. A lot of "speedups" are really just upgrades a customer would get for free. We report those as what they are, and measure any further gain against the vendor's best — never against a configuration picked to lose.
Capacity math, not peak claims
Each report ends with a capacity statement: for a fleet of N cards at reference goodput T₀, a configuration with goodput T₁ needs N′ = ⌈N × T₀ ÷ T₁⌉ cards for the same billable throughput. A 1.4× goodput gain means roughly 29% fewer cards for the same work. We count savings only where they're real: avoided purchases or capacity that can be sold. Idle cards are never counted as savings.
Each report also lists what it does not cover — untested context lengths, quantizations, concurrency ranges and topologies — so nobody extrapolates past the data.
Open method, open data
We also build tuned inference images for AMD Instinct. That is exactly why the method is public, every tuned result sits next to the vendor's latest configuration, and every public report ships with its spec file, raw JSON and the command that reproduces it.
These lab notes are where we write up each run on real hardware: setup, numbers, surprises and what we changed in the tool as a result.
If you operate inference on AMD Instinct or NVIDIA and want the same numbers for your fleet, get in touch.