← Node evaluation

Guide

GPU health check: how to tell if a rented or new GPU is performing as it should

A GPU can pass every listing check and still deliver half its normal speed. Here is what to check on a rented or newly installed NVIDIA GPU, what healthy values look like, and how to run all of it in about two minutes with an open-source tool.

When we rented four H100s that were all listed as "H100 SXM", one of them ran at 787 MHz under load, about 40% of a healthy card, and two had their power limit lowered by the host. Later, one of four L40S nodes from a "secure" datacenter had its power limit cut too. None of this showed in the listing, and nvidia-smi looked normal at idle.

A GPU health check catches these problems in minutes, before you've paid for hours of slow work. This guide covers what to check, what healthy looks like and how to run the checks.

The short version

pip install tokenwatt-check torch    # torch is usually already installed on GPU images
tokenwatt-check                      # about two minutes; PASS / WARN / FAIL per check
tokenwatt-check --watch              # keep checking while your workload runs

tokenwatt-check is our open-source checker (MIT license). It runs everything below and compares the results with measured healthy cards of the same model. The rest of this page explains each check, so you can run them by hand or understand what the tool reports.

1. Power limit

Hosts can lower a GPU's power limit to save electricity or fit more GPUs on a circuit. The card then runs slower under heavy load, and the listing doesn't change.

nvidia-smi --query-gpu=name,power.limit,power.default_limit,power.max_limit --format=csv

Healthy: power.limit equals power.default_limit (700 W for an H100 SXM, 400 W for an A100 SXM, 350 W for an L40S, 450 W for an RTX 4090). Some boards allow raising the limit above the default; running at the default is still full power.

Warning sign: a limit noticeably below the default, such as 525 W on a 700 W H100. Throughput drops under load, so compare providers on cost per token rather than price per hour.

2. Clocks under load, and why idle clocks tell you nothing

At idle every GPU reports low clocks, and its maximum clock is a spec value. What matters is the clock while the GPU works hard. Run a sustained load (a large matrix multiplication, or your own workload) and sample the clock once a second:

nvidia-smi --query-gpu=clocks.sm,clocks.max.sm,power.draw,temperature.gpu,clocks_event_reasons.active --format=csv -l 1

Healthy: under a dense matrix multiplication, a healthy card sits at its power limit and well below its maximum clock. That's normal. In our measurements a healthy H100 SXM held about 1,360 MHz (of 1,980 max) at 700 W, an A100 SXM about 1,235 MHz, an L40S about 1,215 MHz and an RTX 4090 about 2,480 MHz.

Warning sign: a clock far below what a healthy card of the same model reaches under the same load. The 787 MHz H100 drew almost all of its (already lowered) 500 W limit while serving an LLM, where healthy H100s held 1,800–1,970 MHz: about 40% of normal speed, a sign of a degraded card or poor cooling or power delivery.

3. Throttle reasons

nvidia-smi reports why the clock is being held back (clocks_event_reasons). Decode them, or read them in nvidia-smi -q -d PERFORMANCE:

Reason Meaning Normal under full load?
SW power cap The card is at its power limit Yes, that's how GPUs run flat out
HW slowdown Hardware protection (power or temperature) No
HW thermal / SW thermal slowdown Too hot No
HW power brake The server's power supply asked the GPU to slow down No

A card that spends its time in hardware or thermal slowdown is protecting itself, and your work runs slower.

4. Compute and memory bandwidth against a healthy card

The real test is speed, measured the same way on every card:

  • Compute: sustained BF16 matrix multiplication (8192 × 8192) in TFLOPS. Prompt processing and large batches scale with it.
  • Memory bandwidth: a large device-to-device copy in TB/s. LLM token generation follows it closely.

Healthy values from cards we measured with tokenwatt-check (2–3 nodes each where noted):

GPU BF16 matmul Copy bandwidth Power limit
H100 SXM 680 TFLOPS 3.0 TB/s 700 W
H200 660 TFLOPS 4.3 TB/s 700 W
A100 SXM 80 GB 260 TFLOPS 1.75 TB/s 400 W
L40S 170 TFLOPS 0.65 TB/s 350 W
RTX 4090 160 TFLOPS 0.92 TB/s 450 W
A40 107 TFLOPS 0.56 TB/s 300 W

Copy bandwidth is what this test reaches, not the spec sheet: GDDR6 cards reach roughly 65–75% of their rated bandwidth, HBM cards about 90%. More than 10% below these values deserves a closer look; more than 25% below means something is wrong.

5. Memory errors

nvidia-smi --query-gpu=ecc.mode.current,ecc.errors.uncorrected.volatile.total,retired_pages.pending --format=csv

Healthy: ECC enabled on data-center cards, zero uncorrected errors, no pages pending retirement. Uncorrected errors crash jobs or corrupt results. Ask for another machine.

If NVIDIA DCGM is installed, dcgmi diag -r 1 runs NVIDIA's own quick diagnostics (-r 2 takes a couple of minutes and goes deeper). tokenwatt-check runs it automatically when it's available.

These don't change inference speed once a model is loaded, but they decide how long loading takes:

  • PCIe: nvidia-smi --query-gpu=pcie.link.gen.current,pcie.link.width.current --format=csv under load. Expect the card's full generation and x16. A link that trained down to x8 or Gen3 doubles transfer times.
  • Disk: sequential write well above 1,000 MB/s for comfortable model loading.
  • Download: at 500 Mbit/s, 100 GB of model weights takes about half an hour. Hosts vary from under 1 to over 9 Gbit/s.

7. Keep watching during the rental

A check at the start can't see a power limit lowered halfway through your rental, or a card that starts throttling when the room heats up. Leave a watcher running next to your workload:

tokenwatt-check --watch --interval 30 --hours 24 --log node-a.jsonl

It flags power-limit changes, hardware or thermal throttling, clock collapse under load, temperatures of 87 °C or more, PCIe drops and new ECC errors the moment they happen. Add --on-event to send alerts to a webhook.

What a GPU stress test does and doesn't tell you

Tools like gpu-burn load the GPU for minutes or hours and check for computation errors. They're good at finding unstable cards, but they don't tell you whether the card is as fast as it should be. A capped or slow card passes a burn test. Combine a stress run with the speed checks above, and compare with a healthy card of the same model.

From "healthy" to "how much it serves"

A healthy card isn't the same as a cost-effective one. How many users a GPU serves within a latency target, and what each million tokens costs, depends on the model, the serving engine and the settings. In our tests the same 8B model cost $0.90 to $1.61 per million tokens across healthy cheaper GPUs. If you're about to commit to GPU capacity, we can evaluate your nodes remotely and measure both.

FAQ

How long does a GPU health check take?

About two minutes with tokenwatt-check: a 30-second compute test, a 15-second memory test, and quick disk and download tests. DCGM's level-2 diagnostics add a couple of minutes.

Is it normal for my GPU to run below its maximum clock under load?

Yes. Under heavy load a healthy GPU sits at its power limit, and its clock settles well below the maximum. A healthy H100 SXM holds about 1,360 MHz of 1,980 during a dense matrix multiplication. What isn't normal is a clock far below what a healthy card of the same model reaches in the same test.

How do I know if my GPU is throttling?

Check clocks_event_reasons with nvidia-smi while the GPU is under load. "SW power cap" is normal at full load. Hardware slowdown, thermal slowdown or power brake mean the card is protecting itself and running slower than it should.

Can a cloud provider lower my GPU's power limit?

Yes. The host controls the power limit, and we found lowered limits on rented H100 and L40S nodes, including in a "secure" datacenter. Compare power.limit with power.default_limit in nvidia-smi.

Does a GPU stress test like gpu-burn replace a health check?

No. A burn test finds unstable cards that produce errors, but a capped or slow card passes it. You also need to compare speed against a healthy card of the same model.