Principles
- Measure what gets billed. Only tokens served within the latency SLO count. Peak throughput at unbounded latency is not reported.
- Same ruler for every vendor. AMD Instinct and NVIDIA systems use identical definitions, traces and power sources.
- Lock everything up front. SLO, workload, accuracy suite and pass threshold are written into the spec file before the first run.
- Measured watts only. Energy comes from the BMC or a metered PDU, never from nameplate TDP.
- State what isn't covered. Every report lists the context lengths, quantizations and concurrency levels it did not test.
The spec file
A report is defined by its spec file. Changing any field — hardware, driver, engine version, model, quantization, trace, SLO or eval set — produces a different report, not an update to the old one.
hardware: { accelerator: "<model>", count: 8, driver: "<version>" }
engine: { name: vllm, version: "<version>", endpoint: http://localhost:8000/v1 }
model: { name: "<model>", quantization: fp8, kv_cache: fp8 }
slo: { ttft_p99_ms: 2000, tpot_p99_ms: 50 }
workload: { trace: chat-mix-v1, warmup_s: 120, window_s: 600 }
accuracy: { suite: gsm8k-500, max_drop_pts: 0.5, seeds: 3 }
power: { source: redfish, sampling_hz: 1 }
runs: 5SLO goodput
Bench replays the workload trace with Poisson arrivals and sweeps the request rate upward. At each rate it records p99 time-to-first-token (TTFT) and p99 time-per-output-token (TPOT) over the steady-state window.
TTFTp99 ≤ SLOTTFT and TPOTp99 ≤ SLOTPOT
Goodput is reported per node. Requests that fail or time out count against the SLO, not as missing data.
Decode bandwidth utilization (MBU)
During decode, each step reads the active weights and the KV cache from HBM. MBU compares the bytes the model must read with the bandwidth the hardware is rated for. It is derived from model geometry and measured step rate — not from hardware counters — so the same definition applies on every vendor.
- For MoE models, Bweights,active counts only the experts actually routed to in each step, averaged over the window.
- Redundant reads (unfused dequantization, re-reads across kernels) do not raise MBU. Hardware counters are collected as a diagnostic, never as the headline.
- MBU is a diagnostic of software headroom. At large batch sizes decode can become compute- or interconnect-bound, so goodput and tokens per joule remain the primary results.
Power and energy
Node power is read from the BMC (Redfish or IPMI DCMI) or a metered PDU at 1 Hz or faster and averaged over the same steady-state window as goodput. GPU board power from the vendor's management library is reported separately and always labeled as such.
Accuracy gate
Every configuration runs the same evaluation suite with fixed prompts, sampling parameters and seeds. The pass threshold is set in the spec before measurement.
- The reference configuration (usually production) defines the baseline score.
- A configuration fails the gate if its score drops by more than the threshold. Failed configurations are shown, but excluded from capacity math.
- Suites are sized so that seed-to-seed variation is well inside the threshold; the observed spread is published with the report.
Comparisons
A typical report has three columns, measured on the same node in the same session:
| Column | What it is | Why it's there |
|---|---|---|
| Production | The image and flags running in production today | Your real starting point |
| Vendor latest | The hardware vendor's latest recommended image and settings | What you get for free by upgrading |
| Candidate | Any change under evaluation — including TokenWatt Tuned | The increment above the vendor's best |
Cross-vendor reports follow the same rule: each system runs its own vendor's recommended configuration for the same model, quantization class and SLO.
Capacity math
The conclusion of every report is a capacity statement. For a fleet of N cards serving at the reference goodput T0, a configuration with goodput T1 needs:
A 1.4× goodput gain means roughly 29% fewer cards for the same work. Savings only count where they're real — avoided purchases or capacity that can be sold. Idle cards are never counted as savings.
Reproducibility
- Each configuration is measured in at least five runs; the report shows medians and the observed variance band.
- Warm-up is excluded. Clocks, temperatures and throttle events are logged and attached to the raw data.
- Every public report includes the spec file, raw JSON and the exact command used. Another engineer on the same setup should land within the published band.
Scope and limits
A report covers exactly what was measured. In particular, it does not cover untested context lengths, quantization formats, concurrency regions, multi-node topologies or workload traces. These are listed explicitly in every report.
Disclosure
TokenWatt also builds tuned inference images for AMD Instinct. To keep the ruler honest: the measurement method and core are open; every Tuned result is shown next to the vendor's latest recommended configuration; and anyone can rerun a public report.
Questions or corrections: hello@tokenwatt.io.