← All reports

Report · 2026-10-11

Qwen3-8B BF16 on 1× A40 with vLLM 0.31.0

1× NVIDIA A40 Qwen3-8B BF16 weights, BF16 KV cache vLLM 0.31.0 TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Quick-mode measurement on a RunPod Secure Cloud A40, 2026-10-11, default vLLM settings. Part of our comparison of cheaper GPUs serving the same 8B model; the node passed tokenwatt-check and was watched during the run (no changes). Energy is GPU board power.

ConclusionSingle configuration measured under SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms).

Results

Each value is the median of 2 runs.

vLLM default

SLO goodput

output tok/s per node · higher is better
vLLM default102vLLM default: 102 output tok/s per node

Tokens per joule

tok/J, GPU board power · higher is better
vLLM default0.36vLLM default: 0.36 tok/J, GPU board power

Decode bandwidth utilization

% of rated HBM bandwidth

Not measured in this report.

GPU board power

kW, steady state
vLLM default0.29vLLM default: 0.29 kW, steady state

GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

MetricvLLM default
Engine buildvLLM 0.31.0 · KV BF16
Decode bandwidth utilization (MBU)—
GPU board power0.29 kW
SLO goodput102 tok/s
Tokens per joule0.36
TTFT p991,501 ms
TPOT p9938.2 ms
SLO metyes
Accuracy (gsm8k-5shot)90.0
Accuracy Δ vs. referencereference
Cards for equal goodput64

Setup

Any change to these fields makes it a different report.

Accelerator
1× NVIDIA A40
Rated HBM BW
—
Driver / runtime
580.159.03
Host
RunPod secure cloud pod,
Engine
vLLM 0.31.0
Kernel libraries
vllm/vllm-openai:v0.31.0
Model
Qwen3-8B (Dense)
Quantization
BF16 weights, BF16 KV cache
Parallelism
1 replica × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
Runs per config
2 (variance band ±1.9%)
Accuracy gate
gsm8k-5shot · max drop 1.0 pts

Not covered

Do not extrapolate this report to the following.

  • Whole-node power (no BMC access on this host)
  • Context lengths beyond the chat-mix-v1 trace
  • Quick mode: short windows, few repeats — validation only, not for publishing

Reproduce

Run on the same hardware and software versions. Results should fall within ±1.9%.

shell
tokenwatt run --spec lambda-qwen3-8b-tp1-dp1.spec.yaml --label 'runpod-a40-qwen8b-r2' --role other --out runs/runpod-a40-qwen8b-r2.json

tokenwatt report runs/runpod-a40-qwen8b-r2.json --title 'Qwen3-8B BF16 on 1× A40 with vLLM 0.31.0'

Download raw JSON