<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>TokenWatt lab notes</title>
    <link>https://www.tokenwatt.io/blog/</link>
    <description>LLM inference measurements on real hardware.</description>
    <language>en</language>
    <item>
      <title>Does selling tokens pay? Break-even math for an inference provider, from a measured H100</title>
      <link>https://www.tokenwatt.io/blog/does-selling-tokens-pay.html</link>
      <guid>https://www.tokenwatt.io/blog/does-selling-tokens-pay.html</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <description>Using measured SLO throughput for Qwen3-30B-A3B on one H100 and live OpenRouter prices, we work out the break-even utilization for rented, owned and idle GPUs. Rented GPUs need 64–77% utilization to match the cheapest price; idle owned GPUs break even at 2%. Includes a calculator.</description>
    </item>
    <item>
      <title>How many GPUs does an internal ChatGPT need? Fewer than you think — and that&#x27;s the cost problem</title>
      <link>https://www.tokenwatt.io/blog/how-many-gpus-for-internal-chatgpt.html</link>
      <guid>https://www.tokenwatt.io/blog/how-many-gpus-for-internal-chatgpt.html</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <description>Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.</description>
    </item>
    <item>
      <title>Lab notes #2: Short benchmark windows overstated H100 goodput by 14%</title>
      <link>https://www.tokenwatt.io/blog/lab-notes-2-h100-full-mode.html</link>
      <guid>https://www.tokenwatt.io/blog/lab-notes-2-h100-full-mode.html</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <description>Full-mode rerun of Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. With 10-minute windows the SLO boundary fell from 11.3 to 9.5 req/s and goodput from 4,572 to 3,996 tok/s; the FP8 KV-cache accuracy drop shrank from 8 to 4 points on a 600-sample suite.</description>
    </item>
    <item>
      <title>Lab notes #3: Nine providers, one Llama 3.3 70B — same accuracy, 18× different speed</title>
      <link>https://www.tokenwatt.io/blog/lab-notes-3-openrouter-llama-70b-providers.html</link>
      <guid>https://www.tokenwatt.io/blog/lab-notes-3-openrouter-llama-70b-providers.html</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <description>We audited every OpenRouter provider of Llama 3.3 70B Instruct with the same GSM8K suite and latency probe. Accuracy was statistically indistinguishable (94.0–98.0%), decode speed ranged from 19 to 347 tokens/s, and output prices varied 7×. A naive first pass wrongly scored two providers at 0% and 44%.</description>
    </item>
    <item>
      <title>Lab notes #4: Four things that broke a two-node GPU Kubernetes cluster before vLLM would serve</title>
      <link>https://www.tokenwatt.io/blog/lab-notes-4-gpu-kubernetes-four-failures.html</link>
      <guid>https://www.tokenwatt.io/blog/lab-notes-4-gpu-kubernetes-four-failures.html</guid>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <description>Validating a private-inference stack — k3s, NVIDIA GPU Operator, vLLM, DCGM — on two cloud A10s took five attempts. A CUDA 13 image on a 570 driver, broken -cu129 image variants, a CDI hook missing from the host toolkit, and a Service named vllm. Every failure was a version or naming mismatch between open-source parts, and most left no logs.</description>
    </item>
    <item>
      <title>Billable tokens per watt: why we&#x27;re building a vendor-neutral ruler for LLM inference</title>
      <link>https://www.tokenwatt.io/blog/introducing-tokenwatt-bench.html</link>
      <guid>https://www.tokenwatt.io/blog/introducing-tokenwatt-bench.html</guid>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <description>Power, not GPUs, is now the constraint on AI data centers. We explain the four numbers TokenWatt Bench reports — SLO goodput, tokens per joule, decode bandwidth utilization and an accuracy gate — and why each comparison includes the vendor&#x27;s own best configuration.</description>
    </item>
    <item>
      <title>Lab notes #1: FP8 KV cache on an H100 — 11% more goodput, and it failed our accuracy gate</title>
      <link>https://www.tokenwatt.io/blog/lab-notes-1-h100-fp8-kv-cache.html</link>
      <guid>https://www.tokenwatt.io/blog/lab-notes-1-h100-fp8-kv-cache.html</guid>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <description>First real-hardware run of TokenWatt Bench. Qwen3-30B-A3B-FP8 on one H100 with vLLM 0.31.0. FP8 KV cache doubled KV capacity and raised SLO goodput 10.6% and tokens per joule 14%, but GSM8K accuracy dropped 8 points with uncalibrated KV scales.</description>
    </item>
    <item>
      <title>How to measure decode bandwidth utilization without hardware counters</title>
      <link>https://www.tokenwatt.io/blog/measuring-decode-bandwidth-utilization.html</link>
      <guid>https://www.tokenwatt.io/blog/measuring-decode-bandwidth-utilization.html</guid>
      <pubDate>Tue, 06 Oct 2026 00:00:00 +0000</pubDate>
      <description>A vendor-neutral way to compute memory-bandwidth utilization (MBU) for LLM decode from model geometry and client-side timing — including MoE expert routing and MLA KV caches — and how we validated it against a simulated roofline engine.</description>
    </item>
  </channel>
</rss>
