Know what your GPU nodes deliver before you commit.
The same H100 listing can hide a capped power limit, a throttling card or a node that runs at 40% of normal speed. We test your nodes remotely and send a one-page report: health against a healthy card of the same model, the load each node sustains within your latency target, and what a million tokens really costs at your price.
Before you sign
Renting for months or reserving a cluster? Check that the nodes perform like healthy cards and compare providers in cost per token, not price per hour.
Acceptance of new hardware
Bought or built new GPU servers? Verify every node against a healthy reference before it goes into production, with a report you can hand to your vendor.
Capacity planning
How many users can one node serve within your latency target, and what does that cost per million tokens? Measured on your hardware, not estimated from a spec sheet.
What the report covers
One page per node, written to be read by whoever makes the decision. Sample: an RTX 4090 on RunPod.
- Node health: power limit, clocks under load, compute, memory bandwidth, PCIe link, memory errors, disk and download speed, each compared with a healthy card of the same model.
- Stability during the test: we keep watching while the node is under load, and flag power-limit cuts, throttling or clock collapse as they happen.
- Capacity at your latency target: the highest request rate that keeps p99 time to first token and time per output token within your SLO, with median and p99 latency.
- Cost per million tokens at the price you pay, plus energy per token and an accuracy check, so speed isn't bought with quality.
- Findings and recommendations: what's wrong, what to ask your provider, and where tuning could add capacity.
Tell us about your nodes
GPU model and count, where they run, and what you want to know. We reply within one business day with scope and timing.
Give us temporary access
A temporary SSH key to a node that isn't serving production traffic. Or run our open tools yourself and send us the results.
We test
About an hour per node for a quick evaluation, three to four hours for a full one. Nothing stays installed afterwards.
You get the report
Usually within one business day. Results are yours: we never publish them without written permission.
What we've found so far
From GPUs we rented ourselves and tested the same way.
40% of normal speed
An H100 on a marketplace ran at 787 MHz under load and couldn't meet a latency target at any rate. Read more
Lowered power limits
Two of four H100 hosts and one of four L40S nodes had their power limit cut below the card's default, including one in a "secure" datacenter. Read more
1.8× in cost per token
Across healthy cheaper GPUs serving the same 8B model, cost per million tokens ranged from $0.90 to $1.61. The cheapest per hour was the most expensive per token. Read more
Pricing
Per node, fixed price, agreed before we start.
Quick evaluation
$300 per node
Health check against a healthy card, stability watch during the test, capacity at one latency target with a standard model, cost per token at your price, one-page report. About an hour of testing.
Full evaluation
$1,000 per node
Everything in Quick, measured with 10-minute windows and five repeats, at two latency targets, with a 600-question accuracy check and the model you serve if we can access it. Three to four hours of testing.
Founding pilots
Free first evaluation
For our first five customers: your first evaluation (Quick or Full, one node) at no cost, in exchange for permission to publish an anonymized case study. Mention it in your request.
Fleets: a health check on every node plus a capacity test on one node per GPU type, quoted per fleet. You run the nodes, so there's no GPU charge from us; if you ask us to rent nodes to compare providers, rental passes through at cost.
Request an evaluation
We're taking a small number of evaluations now, on NVIDIA GPUs (AMD Instinct support is in progress). Tell us about your setup and we'll reply within one business day with scope and timing.
Prefer email? hello@tokenwatt.io
Questions
What access do you need, and is it safe?
A temporary SSH key to one node that isn't serving production traffic, with permission to run containers or Python. We install nothing permanently, don't touch other machines and don't need your data. Revoke the key when we're done. If you'd rather not give access, run our open-source tools yourself and send us the output.
Will it disturb production?
The capacity test drives the GPU to its limit, so use a drained node or one reserved for testing. The health check alone takes about two minutes and can run on any idle node.
Which GPUs and models do you support?
NVIDIA data-center and workstation GPUs today (H100, H200, A100, L40S, RTX 4090 and others). By default we serve a standard open model so results compare across nodes; we can also test the model you serve if its weights are available to us. AMD Instinct support is in progress.
Why should we trust your numbers?
Because you can check them. The method is published, the node checker is open source (tokenwatt-check on GitHub and PyPI), every report states its settings and limits, and we publish our own corrections when a result doesn't hold up.
How much does an evaluation cost?
$300 per node for a Quick evaluation and $1,000 per node for a Full one, fixed and agreed before we start. Our first five customers get their first evaluation free in exchange for an anonymized case study. Fleets are quoted per fleet.
Do you publish our results?
Never without written permission. Reports on this site come from GPUs we rented ourselves.