Qwen3-8B FP8 on 1× RTX 4090 with vLLM 0.31.0
Quick-mode measurement on a RunPod Secure Cloud RTX 4090, 2026-10-11, default vLLM settings. Part of our comparison of cheaper GPUs serving the same 8B model; the node passed tokenwatt-check and was watched during the run (no changes). Energy is GPU board power.
ConclusionSingle configuration measured under SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms).
Results
Each value is the median of 2 runs.
vLLM default
SLO goodput
output tok/s per node · higher is better
Tokens per joule
tok/J, GPU board power · higher is better
Decode bandwidth utilization
% of rated HBM bandwidth
Not measured in this report.
GPU board power
kW, steady state
GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.
All numbers
Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.
| Metric | vLLM default |
|---|---|
| Engine build | vLLM 0.31.0 · KV BF16 |
| Decode bandwidth utilization (MBU) | — |
| GPU board power | 0.31 kW |
| SLO goodput | 276 tok/s |
| Tokens per joule | 0.88 |
| TTFT p99 | 1,229 ms |
| TPOT p99 | 21.6 ms |
| SLO met | yes |
| Accuracy (gsm8k-5shot) | 94.0 |
| Accuracy Δ vs. reference | reference |
| Cards for equal goodput | 64 |
Setup
Any change to these fields makes it a different report.
- Accelerator
- 1× NVIDIA GeForce RTX 4090
- Rated HBM BW
- —
- Driver / runtime
- 580.159.04
- Host
- RunPod secure cloud pod,
- Engine
- vLLM 0.31.0
- Kernel libraries
- vllm/vllm-openai:v0.31.0
- Model
- Qwen3-8B-FP8 (Dense)
- Quantization
- FP8 weights, BF16 KV cache
- Parallelism
- 1 replica × TP1
- Workload trace
- chat-mix-v1
- Input tokens
- median 1,024 · p90 4,096
- Output tokens
- median 256 · p90 1,024
- Arrival
- Poisson, rate swept to SLO boundary
- Power source
- GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
- Runs per config
- 2 (variance band ±11.7%)
- Accuracy gate
- gsm8k-5shot · max drop 1.0 pts
Not covered
Do not extrapolate this report to the following.
- Whole-node power (no BMC access on this host)
- Context lengths beyond the chat-mix-v1 trace
- Quick mode: short windows, few repeats — validation only, not for publishing
Reproduce
Run on the same hardware and software versions. Results should fall within ±11.7%.
tokenwatt run --spec lambda-qwen3-8b-fp8-tp1-dp1.spec.yaml --label 'runpod-rtx4090-qwen8b-fp8-r2' --role other --out runs/runpod-rtx4090-qwen8b-fp8-r2.json tokenwatt report runs/runpod-rtx4090-qwen8b-fp8-r2.json --title 'Qwen3-8B FP8 on 1× RTX 4090 with vLLM 0.31.0'