Decode bandwidth utilization
How much of the rated HBM bandwidth the decode loop actually turns into useful reads. The clearest signal of software headroom.
TokenWatt runs on your nodes, with your model and your latency SLO, and reports what your hardware actually delivers — on AMD Instinct and NVIDIA, with the same ruler.
New capacity waits on interconnection, not on GPUs. Every energized watt that doesn't turn into a billable token is the most expensive kind of idle.
Peak tokens per second at unlimited latency tells you nothing about how many cards you need to serve your users within your TTFT and tail-latency targets.
A new driver, engine version or quantization can swing throughput by double digits. Without a fixed ruler, nobody can tell whether it got better or worse.
Every TokenWatt report pins the hardware, driver, engine, model, quantization, workload trace and SLO — then states exactly what it does and doesn't cover.
How much of the rated HBM bandwidth the decode loop actually turns into useful reads. The clearest signal of software headroom.
From the BMC or a metered PDU, sampled through the steady-state window. GPU board power is reported separately and labeled.
Output tokens per second at the highest load where p99 TTFT and p99 time-per-output-token both stay inside your targets.
A fixed evaluation set with a pass threshold agreed before the run. Speed that costs quality doesn't count.
Bench is a client. It talks to any OpenAI-compatible server — vLLM, SGLang, TensorRT-LLM — and reads telemetry from the node. It never touches your weights.
Declare TTFT and tail-latency targets, the request-length distribution and the eval set up front. The spec file is part of the report.
Production image, vendor's latest recommended image, and any candidate change — measured the same way, on the same node.
Machine-readable JSON plus a one-page summary with the capacity math: how many cards and kilowatts the same billable throughput needs.
# 1. describe what "good" means for you $ cat slo.yaml slo: ttft_p99_ms: 2000 tpot_p99_ms: 50 workload: trace: chat-mix-v1 # input/output length distribution accuracy: suite: gsm8k-500 max_drop_pts: 0.5 power: source: redfish # or ipmi-dcmi, pdu # 2. measure each configuration $ docker run --rm --network host \ -v $PWD:/work tokenwatt/bench run \ --endpoint http://localhost:8000/v1 \ --spec /work/slo.yaml --label production # 3. compare and write the report $ tokenwatt report production vendor-latest tuned ✓ report written: tw-report-2026-10-06.json ✓ summary written: tw-report-2026-10-06.html
We also build tuned inference images for AMD Instinct. That's exactly why the rules below are strict and public.
AMD Instinct and NVIDIA are measured with the same definitions, the same workload traces and the same power sources.
Every comparison includes the vendor's own latest recommended configuration. We never benchmark against a configuration we picked to lose.
Definitions, workload traces and the measurement core are published. Anyone can rerun a report and get the same answer within its stated band.
Bench is read-only. No weights, prompts or logs leave your environment unless you choose to publish a report.
Node power comes from the BMC or a metered PDU. Nameplate TDP is never used to estimate energy.
Every report lists the context lengths, quantizations and concurrency levels it did not test. No extrapolation, no peak-only numbers.
One telemetry core, three products, all priced per GPU and installed by your own engineers.
Vendor-neutral inference measurement for your models, SLOs and nodes.
Open core. Paid tier adds power capture, history and reports.
Find the GPU that isn't broken — just slow.
Read-only agent. Installs in minutes.
Decode-path tuned inference images for AMD Instinct, verified with Bench.
Per-GPU subscription. Your gains, verified by you.
Every public report ships with its spec file, raw JSON and the command that reproduces it.
We're looking for a small number of US inference teams and GPU clouds to work with closely. You get Bench on your fleet and a say in the roadmap; we get honest feedback.