Sample report. A real evaluation of a GPU we rented on RunPod, 2026-10-11. Request an evaluation of your nodes.

TokenWatt remote evaluation · TW-20261011-4090-FP8 · 2026-10-11 08:59 UTC

RunPod Secure · RTX 4090 · Qwen3-8B FP8

1× NVIDIA GeForce RTX 4090 · $0.89 per node-hour

Node health
PASS
check at start + watch during the test
Capacity at SLO
0.74 req/s
276 output tok/s · p99 TTFT ≤ 2 s, TPOT ≤ 50 ms
Cost per 1M output tokens
$0.897
at $0.89/h, running at capacity
Energy
0.88 tok/J
313 W board power at capacity
Accuracy
94.0%
gsm8k-5shot, 50 items × 1

Findings and recommendations

Node health at start

tokenwatt-check 0.3.1 (re-evaluated; measured with 0.3.0): each GPU is compared with a healthy card of the same model under the same test.

ResultCheckFinding
PASSpower limitFull power limit: 450 W (the board allows up to 500 W).
PASSclocks under loadSM clock during the matmul averaged 2482 MHz — 100% of a healthy RTX 4090 (2480 MHz).
INFOthrottlingRunning at its power limit for 100% of the load test (normal under full load).
PASSBF16 matmul160.3 TFLOPS — 100% of the healthy RTX 4090 reference (160 TFLOPS).
PASSmemory copy bandwidth0.92 TB/s — 100% of the healthy RTX 4090 reference (0.92 TB/s).
PASSPCIe linkGen4 x16.
PASSdiskSequential write 1,145 MB/s.
PASSdownload1,930 Mbit/s from Hugging Face, one stream (~6.9 min per 100 GB of weights).

power limit pct: 100 · clock under load pct: 100 · tflops: 160.3 · tflops vs reference pct: 100 · tb s: 0.92 · tb s vs reference pct: 100 · disk write mb s: 1145 · download mbit s: 1930

Stability during the test

tokenwatt-check --watch sampled the GPUs every few seconds while the benchmark ran (40 min, 162 samples).

PASS No changes: power limit held at 450 W, no hardware or thermal throttling, no clock collapse, no PCIe or memory errors.

Capacity and latency

At capacity
Request rate meeting the SLO0.74 req/s (2/2 repeats of 2 min passed; spread 11.7%)
Output tokens per second276
Input tokens per second1,100
Time to first token, median / p99156 / 1,229 ms (target p99 ≤ 2,000)
Time per output token, median / p9913.1 / 21.6 ms (target p99 ≤ 50)
Requests in flight (mean)4
Requests served / failed158 / 0
Board power · memory-bandwidth utilization313 W · —
GPU clock · max temperature2,758 MHz · 77 °C

Search for the capacity (each step one window at a fixed request rate):

req/sSLOoutput tok/sTTFT p99 msTPOT p99 ms
0.25PASS9868611.6
0.38PASS12069713.5
0.56PASS19389915.2
0.84PASS31886019.7
1.27PASS4871,03235.4
1.90FAIL7591,53153.9
1.55PASS62498037.4
1.72PASS6971,14146.9
1.80PASS7471,06449.3

Cost

Node price$0.89 per hour · $641 per 30 days
Cost per million output tokens at capacity$0.897
Output tokens per 30 days at capacity0.7 billion
At 50% average utilization$1.794 per million

Method

Host

CPU · RAMAMD EPYC 7282 16-Core Processor · 64 threads · 504 GB
Driver580.159.04
Disk1,145 MB/s sequential write · 129 GB free
Download from Hugging Face1,930 Mbit/s, one stream (~6.9 min per 100 GB)