← All reports

Report · 2026-10-06

Mid-size MoE decode on an 8-GPU node

8× Accelerator A (example) Example-MoE-30B-A3B FP8 weights, FP8 KV cache vLLM 0.x (example) TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Format example. Three configurations of the same node, model and SLO: the production image as deployed, the vendor's latest recommended image, and a tuned decode path.

ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on Production takes: with Vendor latest, 64 → 47 cards and 21% less node energy per token; with Tuned, 64 → 35 cards and 39% less node energy per token.

Results

Each value is the median of 5 runs.

ProductionVendor latestTuned

SLO goodput

output tok/s per node · higher is better
Production2,400Production: 2,400 output tok/s per node
Vendor latest3,300Vendor latest: 3,300 output tok/s per node
Tuned4,500Tuned: 4,500 output tok/s per node

Tokens per joule

tok/J at the node · higher is better
Production0.39Production: 0.39 tok/J at the node
Vendor latest0.50Vendor latest: 0.50 tok/J at the node
Tuned0.64Tuned: 0.64 tok/J at the node

Decode bandwidth utilization

% of rated HBM bandwidth
Production31%Production: 31 % of rated HBM bandwidth
Vendor latest44%Vendor latest: 44 % of rated HBM bandwidth
Tuned63%Tuned: 63 % of rated HBM bandwidth

Axis runs to 100% of rated bandwidth.

Node power

kW, steady state
Production6.10Production: 6.10 kW, steady state
Vendor latest6.60Vendor latest: 6.60 kW, steady state
Tuned7.00Tuned: 7.00 kW, steady state

Higher power is expected when the node does more work; tokens per joule is the efficiency measure.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

MetricProductionVendor latestTuned
Engine build———
Decode bandwidth utilization (MBU)31%44%63%
Node power6.10 kW6.60 kW7.00 kW
SLO goodput2,400 tok/s3,300 tok/s4,500 tok/s
Tokens per joule0.390.500.64
TTFT p991,710 ms1,650 ms1,590 ms
TPOT p9947.0 ms46.0 ms48.0 ms
SLO metyesyesyes
Accuracy (gsm8k-500)88.488.388.2
Accuracy Δ vs. referencereference-0.1 pts-0.2 pts
Cards for equal goodput644735

Setup

Any change to these fields makes it a different report.

Accelerator
8× Accelerator A (example)
Rated HBM BW
5.0 TB/s per GPU
Driver / runtime
example-driver 1.2.3
Host
2-socket server, 8 accelerators
Engine
vLLM 0.x (example)
Kernel libraries
vendor kernel library (example)
Model
Example-MoE-30B-A3B (MoE)
Quantization
FP8 weights, FP8 KV cache
Parallelism
8 replicas × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
BMC via Redfish · 1 Hz · 600 s steady state
Runs per config
5 (variance band ±3%)
Accuracy gate
gsm8k-500 · max drop 0.5 pts

Not covered

Do not extrapolate this report to the following.

  • Context lengths above 8,192 input tokens
  • Quantization formats other than FP8 weights with FP8 KV cache
  • Concurrency beyond the SLO boundary found in this run
  • Multi-node deployments and prefill/decode disaggregation
  • Workload traces other than chat-mix-v1

Reproduce

Run on the same hardware and software versions. Results should fall within ±3%.

shell
docker run --rm --network host -v $PWD:/work tokenwatt/bench run \
  --endpoint http://localhost:8000/v1 \
  --spec /work/example-moe-8gpu.spec.yaml --runs 5 --label <config>

tokenwatt report production vendor-latest tuned

Download raw JSON