Open measurement for LLM inference

Billable tokens per watt. Measured, not claimed.

TokenWatt runs on your nodes, with your model and your latency SLO, and reports what your hardware actually delivers — on AMD Instinct and NVIDIA, with the same ruler.

Runs as one container Read-only, no data leaves your cluster Measured power, never TDP math
tokenwatt bench · moe-decode · 8 GPUs Illustrative
Decode bandwidth utilizationMBU vs. rated HBM bandwidth
31%
Node powerBMC · 1 Hz · steady state
6.1kW
SLO goodputoutput tok/s · TTFT ≤ 2 s · p99 TPOT ≤ 50 ms
2,400tok/s
TTFT within SLO p99 TPOT within SLO
Tokens per joulegoodput ÷ node power
0.39tok/J
Accuracy gatefixed eval set · Δ vs. production
reference
Production image as deployed today.
The constraint

Power is the bottleneck now.

New capacity waits on interconnection, not on GPUs. Every energized watt that doesn't turn into a billable token is the most expensive kind of idle.

The gap

Spec sheets aren't SLOs.

Peak tokens per second at unlimited latency tells you nothing about how many cards you need to serve your users within your TTFT and tail-latency targets.

The drift

Every release moves the number.

A new driver, engine version or quantization can swing throughput by double digits. Without a fixed ruler, nobody can tell whether it got better or worse.

What we measure

Four numbers. One page. Reproducible by your own engineers.

Every TokenWatt report pins the hardware, driver, engine, model, quantization, workload trace and SLO — then states exactly what it does and doesn't cover.

01 · Bandwidth

Decode bandwidth utilization

How much of the rated HBM bandwidth the decode loop actually turns into useful reads. The clearest signal of software headroom.

MBU = bytes per step × steps/s ÷ peak BW
02 · Power

Measured node power

From the BMC or a metered PDU, sampled through the steady-state window. GPU board power is reported separately and labeled.

source ∈ {Redfish, IPMI DCMI, PDU}
03 · Goodput

Throughput within SLO

Output tokens per second at the highest load where p99 TTFT and p99 time-per-output-token both stay inside your targets.

tok/J = goodput ÷ node power
04 · Accuracy

Accuracy gate

A fixed evaluation set with a pass threshold agreed before the run. Speed that costs quality doesn't count.

pass ⇔ Δscore ≥ −threshold
How it works

From docker pull to a signed-off report in an afternoon.

  1. Point it at your endpoint

    Bench is a client. It talks to any OpenAI-compatible server — vLLM, SGLang, TensorRT-LLM — and reads telemetry from the node. It never touches your weights.

  2. Lock the SLO and the workload

    Declare TTFT and tail-latency targets, the request-length distribution and the eval set up front. The spec file is part of the report.

  3. Compare configurations side by side

    Production image, vendor's latest recommended image, and any candidate change — measured the same way, on the same node.

  4. Get the report

    Machine-readable JSON plus a one-page summary with the capacity math: how many cards and kilowatts the same billable throughput needs.

bench · preview
# 1. describe what "good" means for you
$ cat slo.yaml
slo:
  ttft_p99_ms: 2000
  tpot_p99_ms: 50
workload:
  trace: chat-mix-v1        # input/output length distribution
accuracy:
  suite: gsm8k-500
  max_drop_pts: 0.5
power:
  source: redfish           # or ipmi-dcmi, pdu

# 2. measure each configuration
$ docker run --rm --network host \
    -v $PWD:/work tokenwatt/bench run \
    --endpoint http://localhost:8000/v1 \
    --spec /work/slo.yaml --label production

# 3. compare and write the report
$ tokenwatt report production vendor-latest tuned
✓ report written: tw-report-2026-10-06.json
✓ summary written: tw-report-2026-10-06.html
Neutral by design

A ruler is only worth something if everyone trusts it.

We also build tuned inference images for AMD Instinct. That's exactly why the rules below are strict and public.

One method for every vendor

AMD Instinct and NVIDIA are measured with the same definitions, the same workload traces and the same power sources.

The vendor's best is the bar

Every comparison includes the vendor's own latest recommended configuration. We never benchmark against a configuration we picked to lose.

Open method, open core

Definitions, workload traces and the measurement core are published. Anyone can rerun a report and get the same answer within its stated band.

Your data stays put

Bench is read-only. No weights, prompts or logs leave your environment unless you choose to publish a report.

Measured watts only

Node power comes from the BMC or a metered PDU. Nameplate TDP is never used to estimate energy.

Scope is written down

Every report lists the context lengths, quantizations and concurrency levels it did not test. No extrapolation, no peak-only numbers.

Products

Start with the ruler. Grow into the whole fleet.

One telemetry core, three products, all priced per GPU and installed by your own engineers.

Early access AMD Instinct · NVIDIA

Bench

Vendor-neutral inference measurement for your models, SLOs and nodes.

  • SLO goodput, tokens per joule, MBU, accuracy gate
  • Any OpenAI-compatible engine
  • Regression checks in CI before every upgrade
  • Capacity math: cards and kW per billable token

Open core. Paid tier adds power capture, history and reports.

In development AMD Instinct · NVIDIA

Health

Find the GPU that isn't broken — just slow.

  • Per-GPU performance baselines, fleet-wide
  • Throttling, memory errors, degraded links
  • Pre-flight checks before jobs land
  • One view across mixed-vendor fleets

Read-only agent. Installs in minutes.

Design partners AMD Instinct

Tuned

Decode-path tuned inference images for AMD Instinct, verified with Bench.

  • Dequant, GEMV, attention, MoE dispatch
  • Version-locked images, one-step rollback
  • Measured against the vendor's latest, not a strawman
  • Tracks new engine and driver releases

Per-GPU subscription. Your gains, verified by you.

Public reports

Results, with the receipts attached.

Every public report ships with its spec file, raw JSON and the command that reproduces it.

All reports

Running inference on AMD Instinct or NVIDIA? Become a design partner.

We're looking for a small number of US inference teams and GPU clouds to work with closely. You get Bench on your fleet and a say in the roadmap; we get honest feedback.