About

About TokenWatt

GPUs are bought and rented by the hour, but what matters is what they deliver. TokenWatt measures that, independently and in the open, so the people choosing GPUs can decide on evidence instead of listings.

Why TokenWatt exists

The world is building GPU capacity faster than ever, and more of it is being rented, resold and shared through clouds, marketplaces and brokers. But GPUs are still sold on their spec sheets and listings: a model name, a memory size, a price per hour.

What a buyer actually needs to know is different: is this node healthy, how much can it serve within my latency target, and what does each million tokens cost? Two GPUs with the same listing can differ by several times on all three. We've seen an H100 running at 40% of normal speed, power limits quietly lowered by hosts, and the cheapest GPU per hour turn out to be the most expensive per token.

TokenWatt exists to measure that gap and make it visible.

What we do

  • Measure. We rent GPUs on real clouds, at real prices, and measure them the same way every time: node health against a healthy card of the same model, capacity within a latency target, cost per million tokens, energy and accuracy. We publish the results with their settings, raw data and limits, in our lab notes and reports.
  • Track prices. Our GPU price index collects on-demand prices from cloud providers every day, normalized per GPU-hour.
  • Build open tools. tokenwatt-check, our node checker, is open source. Anyone can check a GPU in two minutes, and watch it while it works.
  • Evaluate nodes for teams. Before a team commits to GPU capacity, or when new hardware arrives, we evaluate their nodes remotely and report what they actually deliver.

Our principles

  1. Measured, not claimed. Every number on this site comes from hardware we ran ourselves. When we estimate, we say so.
  2. The same ruler for everyone. One method, applied the same way to every GPU, cloud and engine. We publish the method so you can check it.
  3. Limits are part of the result. Every report says what it didn't cover: one node, one model, short measurement windows.
  4. Corrections in public. When a result doesn't hold up, we correct it on the page and say what changed. Our first SGLang comparison is an example: the full measurement overturned part of it, and the post says so.
  5. Independent. We're not affiliated with any GPU vendor, cloud or engine. Some posts contain referral links to clouds we tested. They're disclosed in the post, we pay for every GPU-hour ourselves, and they never change what we measure.
  6. Read the terms. We only publish data from platforms whose terms allow it.

Where we're going

We think of TokenWatt as Kayak and Consumer Reports for GPUs: a place where anyone choosing GPU capacity can compare options by what they deliver, not what they're called.

The path there:

  1. Evaluation and comparison, today: measured reports, a daily price index, open tools, and node evaluations for teams.
  2. Tuning: getting more out of the GPUs a team already has, through serving engines, settings and configuration, verified with the same measurements.
  3. Managed and shared capacity: helping teams put idle or underused GPUs to work, with the measurements that make shared capacity trustworthy.

Underneath all three is the same idea: GPU capacity should be bought, sold and shared on measured performance.

Who we are

TokenWatt is founded by Bo Shen, who previously founded Codoon, a fitness app with more than 100 million users. It's operated by Teampulse Solutions LLC, a Delaware limited liability company.

Write to us at hello@tokenwatt.io.