Sample report. A real evaluation of a GPU we rented on RunPod, 2026-10-11. Request an evaluation of your nodes.
TokenWatt remote evaluation · TW-20261011-4090-FP8 · 2026-10-11 08:59 UTC
RunPod Secure · RTX 4090 · Qwen3-8B FP8
1× NVIDIA GeForce RTX 4090 · $0.89 per node-hour
Findings and recommendations
- INFO Capacity was measured in quick mode (2-minute windows). In our full-mode measurements (10-minute windows, 5 repeats) capacity came out 15–20% lower, because longer windows catch more of the rare slow requests. Plan with that headroom, or order a full-mode evaluation.
Node health at start
tokenwatt-check 0.3.1 (re-evaluated; measured with 0.3.0): each GPU is compared with a healthy card of the same model under the same test.
| Result | Check | Finding |
|---|---|---|
| PASS | power limit | Full power limit: 450 W (the board allows up to 500 W). |
| PASS | clocks under load | SM clock during the matmul averaged 2482 MHz — 100% of a healthy RTX 4090 (2480 MHz). |
| INFO | throttling | Running at its power limit for 100% of the load test (normal under full load). |
| PASS | BF16 matmul | 160.3 TFLOPS — 100% of the healthy RTX 4090 reference (160 TFLOPS). |
| PASS | memory copy bandwidth | 0.92 TB/s — 100% of the healthy RTX 4090 reference (0.92 TB/s). |
| PASS | PCIe link | Gen4 x16. |
| PASS | disk | Sequential write 1,145 MB/s. |
| PASS | download | 1,930 Mbit/s from Hugging Face, one stream (~6.9 min per 100 GB of weights). |
power limit pct: 100 · clock under load pct: 100 · tflops: 160.3 · tflops vs reference pct: 100 · tb s: 0.92 · tb s vs reference pct: 100 · disk write mb s: 1145 · download mbit s: 1930
Stability during the test
tokenwatt-check --watch sampled the GPUs every few seconds while the benchmark ran (40 min, 162 samples).
PASS No changes: power limit held at 450 W, no hardware or thermal throttling, no clock collapse, no PCIe or memory errors.
Capacity and latency
| At capacity | |
|---|---|
| Request rate meeting the SLO | 0.74 req/s (2/2 repeats of 2 min passed; spread 11.7%) |
| Output tokens per second | 276 |
| Input tokens per second | 1,100 |
| Time to first token, median / p99 | 156 / 1,229 ms (target p99 ≤ 2,000) |
| Time per output token, median / p99 | 13.1 / 21.6 ms (target p99 ≤ 50) |
| Requests in flight (mean) | 4 |
| Requests served / failed | 158 / 0 |
| Board power · memory-bandwidth utilization | 313 W · — |
| GPU clock · max temperature | 2,758 MHz · 77 °C |
Search for the capacity (each step one window at a fixed request rate):
| req/s | SLO | output tok/s | TTFT p99 ms | TPOT p99 ms |
|---|---|---|---|---|
| 0.25 | PASS | 98 | 686 | 11.6 |
| 0.38 | PASS | 120 | 697 | 13.5 |
| 0.56 | PASS | 193 | 899 | 15.2 |
| 0.84 | PASS | 318 | 860 | 19.7 |
| 1.27 | PASS | 487 | 1,032 | 35.4 |
| 1.90 | FAIL | 759 | 1,531 | 53.9 |
| 1.55 | PASS | 624 | 980 | 37.4 |
| 1.72 | PASS | 697 | 1,141 | 46.9 |
| 1.80 | PASS | 747 | 1,064 | 49.3 |
Cost
| Node price | $0.89 per hour · $641 per 30 days |
| Cost per million output tokens at capacity | $0.897 |
| Output tokens per 30 days at capacity | 0.7 billion |
| At 50% average utilization | $1.794 per million |
Method
- Engine: vLLM 0.31.0 (vllm/vllm-openai:v0.31.0); model: Qwen3-8B-FP8 (FP8 weights, BF16 KV cache).
- Workload: chat-mix-v1, chat-like requests arriving at random (Poisson); mean 1,488 input and 373 output tokens per request at capacity.
- SLO: p99 time to first token ≤ 2,000 ms and p99 time per output token ≤ 50 ms. Capacity is the highest Poisson request rate that met both in every repeat.
- Measurement: 60 s search windows, then 2 repeats of 120 s after 45 s warm-up; a failed repeat backs off 10% and starts over.
- Accuracy gate: gsm8k-5shot, 50 items × 1 seed(s): 94.0%, so the configuration isn't buying speed with quality.
- Power: GPU board power sum (NVIDIA), sampled at 1 Hz.
- Full method: tokenwatt.io/methodology.
Host
| CPU · RAM | AMD EPYC 7282 16-Core Processor · 64 threads · 504 GB |
| Driver | 580.159.04 |
| Disk | 1,145 MB/s sequential write · 129 GB free |
| Download from Hugging Face | 1,930 Mbit/s, one stream (~6.9 min per 100 GB) |