← Lab notes

Lab notes #5: Serving one H100 as a token API for six hours — 162,040 requests, 9 failures, no request lost when the GPU was handed back

We measured an H100's SLO capacity, set the gateway's admission limit from it, and served Qwen3-30B-A3B-FP8 through it for six hours of day-shaped traffic with hourly overload and GPU-handback drills. Availability 99.994%, 99.97% of requests within the SLO, six drains with no in-flight loss, billing reconciled to the token. Here are the failures as well.

Benchmarks answer "how fast is this GPU for ten minutes?" A token API has to answer a different question: what happens over hours of uneven traffic, when demand goes past capacity, and when the GPU has to be handed back to its owner?

We built the serving path we plan to sell and ran it for six hours on one rented H100. It is the same path a marketplace like OpenRouter or Hugging Face would call: vLLM, then our gateway, then HTTPS.

Result: 162,040 requests.

  • Availability (successes over successes plus failures, not counting deliberate 429s): 99.994%.
  • 99.97% of accepted requests met the SLO.
  • Six capacity handbacks completed with no in-flight request lost.
  • The billing ledger matched the client's token counts exactly once two misclassified requests were accounted for.

There were nine failures. Seven came from our gateway's connection handling; we explain all nine below.

The setup

Machine Lambda Cloud gpu_1x_h100_sxm5 (H100 80 GB HBM3), us-southeast-1, driver 580.105.08
Engine vLLM 0.31.0 (vllm/vllm-openai:latest), Qwen/Qwen3-30B-A3B-FP8, max context 12,288, tool-call parser hermes, reasoning parser qwen3
Gateway TokenWatt provider gateway: OpenRouter-style /v1/models, usage in every response, 429 at the admission limit, drain/resume, per-request billing ledger
TLS Caddy with a Let's Encrypt certificate for an sslip.io name
SLO p99 time to first token ≤ 2 s, p99 time per output token ≤ 50 ms
Traffic Synthetic chat mix (median ~1.7k input / ~420 output tokens), Poisson arrivals
Cost 395 minutes of instance time, $28.24

Step 1: measure, then set the admission limit

Before serving anything we ran a short Bench sweep on the same GPU. The highest rate that met the SLO was 11.3 req/s: 4,572 output tok/s at 0.63 kW, with GSM8K at 96.0% on 50 items. At that rate the server held about 157 requests in flight, so we set the gateway's admission limit to 157.

This is the core of the method. The gateway doesn't queue requests the GPU can't handle within the SLO. It rejects them in a fraction of a second, so a router can send them somewhere else. OpenRouter doesn't count 429s against a provider's uptime, but it does notice slow responses.

That sweep used Bench's quick mode (60-second search windows, two repeats). Our earlier full-mode measurement on another H100 SXM found 9.5 req/s. Quick mode overestimated capacity before, and it probably did here too. The admission limit kept that from hurting the SLO.

Step 2: six hours of traffic

Each hour followed the same pattern:

  • Base load followed a compressed "day": 30% → 90% → 30% of measured capacity over the six hours.
  • Minutes 40–44: overload at 150% of capacity.
  • Minutes 47–49: drain. This simulates the GPU's owner taking it back. New requests get 429; requests already running must finish. Then service resumes.
  • Every minute: a tool-calling request, a structured-output (JSON schema) request, and the same small request sent once straight to vLLM and once through the gateway and HTTPS, to measure their overhead.
  • Every 15 seconds: gateway counters, vLLM queue and KV-cache metrics, GPU power, temperature and clocks.

Results

Requests 162,040
Succeeded 147,434
Rejected with 429 (overload, admission limit, drains) 14,597
Failed 9
Availability (excluding 429s) 99.994%
Accepted requests within SLO 99.97%
Time to first token, p50 / p99 100 ms / 310 ms
Time per output token, p50 / p99 24.6 ms / 38.2 ms
Output tokens served 63.0 million (2,906 tok/s average)

Hour by hour:

Hour Requests 429 Failed In SLO TTFT p99 TPOT p99 Tok/s GPU W °C max
1 17,422 2,164 1 99.99% 271 ms 37.5 ms 1,825 444 58
2 26,748 2,444 3 100.00% 286 ms 37.3 ms 2,861 518 58
3 36,421 3,108 2 99.95% 326 ms 38.6 ms 3,943 592 59
4 36,480 2,834 1 99.95% 338 ms 39.5 ms 3,953 598 60
5 27,347 2,216 2 99.96% 295 ms 36.6 ms 2,994 527 57
6 17,622 1,831 0 99.97% 279 ms 38.0 ms 1,863 447 59

Latency stayed flat over six hours. TTFT p99 tracked load and rose about 25% at the peak, which is expected. It didn't drift upward over time. Temperatures never passed 60 °C. The GPU hit its software power cap in about 44% of samples, as it did in our earlier full run.

Overload

During the 150% windows the gateway turned away 9,032 requests and accepted 15,511. Accepted requests stayed fast: TTFT p99 was 358 ms and TPOT p99 41.4 ms, against 303 ms and 37.1 ms in normal minutes. 99.9% of them met the SLO.

At the peak of the "day" (90% of the quick-mode capacity) the admission limit also kicked in occasionally outside the overload windows, turning away 609 requests. vLLM's KV cache reached 98.7% once, in hour 4, and vLLM preempted 10 requests. That was the only sign of strain all run.

Handing the GPU back

Drain Running at drain Finished Lost New requests rejected GPU empty after
1 26 26 0 561 / 561 12 s
2 60 60 0 962 / 962 14 s
3 150 150 0 1,247 / 1,247 22 s
4 96 96 0 1,058 / 1,058 16 s
5 35 35 0 692 / 692 13 s
6 13 13 0 407 / 407 11 s

Every request that was running when a drain began finished normally, and every new arrival was rejected immediately. The GPU was completely idle 11–22 seconds after the drain began, even with 150 requests in flight. That number determines whether idle enterprise GPUs can be lent out and taken back on short notice. For this model and workload, the answer is about 20 seconds.

Tool calling and structured outputs

Probe Sent Rejected (429, during overload) Accepted and correct
Tool call (get_weather, tool_choice: auto) 342 10 332 / 332
JSON schema (response_format) 342 10 330 / 332

The two JSON failures both hit the 1,024-token limit before producing JSON. Qwen3 "thinks" by default, and on those two requests the reasoning consumed the whole budget. A client that wants JSON should turn thinking off or allow more tokens. Marketplaces test both features, so we now run these probes on every deployment.

Gateway and HTTPS overhead

Across 337 paired requests the gateway added a median of 2 ms to time to first token; at p90 it added 37 ms, mostly when the paired request landed during a busy moment. In the final 2.5 hours we also probed the public HTTPS address from a laptop every 30 seconds, as a marketplace health check would. Results: 296 of 296 model-list calls succeeded, 273 chat calls succeeded, 23 got 429 (all in overload or drain windows), and none failed. Over the public internet, TTFT p50 was 264 ms, and 429s came back in 214 ms median, the same as a round trip.

The nine failures

Seven were HTTP 502 "backend unavailable", about one an hour, at moderate load. The gateway couldn't open or reuse a connection to vLLM, which was running normally the whole time and logged no errors. The most likely cause is a well-known race: the gateway reuses a pooled keep-alive connection just as the server closes it for idling. The gateway didn't log the exception type, so we can't prove it. That logging gap was our mistake.

We changed the gateway in two ways:

  • it now logs the exception;
  • it retries once, immediately, when a connection fails before any response byte arrives. At that point the engine hasn't accepted the request, so retrying is safe.

A test that drops the first connection now passes. Whether the 502s are gone will show in the next run.

Two were "no tokens received" from the client's point of view, but the server completed them, generating 439 and 1,298 tokens. With ignore_eos and random-token prompts, the model sometimes emits only tokens that decode to empty text. This comes from synthetic benchmark traffic and wouldn't happen with real prompts. Our load client now separates these from real failures.

Billing, checked to the token

The gateway records each completed request's token counts and cost in a ledger that holds metadata only. This is what Hugging Face's billing API queries. At the end we compared it with what the client saw. The ledger had 148,100 entries; the client expected 148,098. The extra two were the "no tokens received" requests above: the server finished them and billed them. Their 439 + 1,298 tokens account exactly for the 1,737-token difference. Everything else matched.

Partial streams were never billed. Our paired-latency probe disconnects after the first token on purpose, and none of those 337 requests appear in the ledger.

Energy and economics

Average GPU board power 520 W
Energy over 6 hours (board) 11.3 MJ (3.1 kWh)
Output tokens per joule (whole run) 5.6
Output tokens per joule (at measured capacity) 7.2
GPU rent for the six hours $25.74
Revenue at the prices in our test config ($0.06 / $0.25 per M input / output) $30.71

At an average 62% of measured capacity, this H100 covered its rent at the low prices we used for the test, with about 16% to spare. That's consistent with our break-even model: utilization decides whether selling tokens pays. Board power excludes the host, so energy per token here is a lower bound.

What we changed

  • Gateway: logs backend exceptions; retries once on connection failures that happen before a response.
  • Soak harness: the instance can't reach its own public IP through the cloud's NAT, so it now resolves its HTTPS name to loopback and the main load goes through TLS too. The schema checks now install correctly.
  • Admission limits: we'll set them from full-mode measurements, or about 10% below a quick-mode result.
  • vLLM: it warned Using default MoE config. Performance might be sub-optimal! because no tuned fused-MoE kernel config ships for this shape on H100. That is a free optimization we haven't taken yet.

Limits of this test

  • One GPU, one model, six hours. A 24-hour run and multi-GPU serving are next.
  • Synthetic traffic: random-token prompts with ignore_eos. Real prompts have more prefix-cache hits and natural stopping points.
  • For the first 3.5 hours the main load reached the gateway over local HTTP rather than through TLS (the NAT issue above). The external HTTPS probe covers only the last 2.5 hours.
  • Capacity came from quick mode, which tends to overestimate.
  • Power is GPU board power, not the whole server.

Reproduce

From the bench directory, on Lambda Cloud:

scripts/lambda/cloud.py session --ssh-key <key> --type gpu_1x_h100_sxm5 \
  --max-minutes 450 --open-ports 80,443 --yes \
  --run "MODE=quick MODEL=Qwen/Qwen3-30B-A3B-FP8 MIN_DRIVER=580 \
         VLLM_ARGS='--enable-auto-tool-choice --tool-call-parser hermes --reasoning-parser qwen3' \
         SOAK_HOURS=6 GW_KEY=<random> GW_ADMIN_KEY=<random>"

The session measures capacity, configures and starts the gateway and HTTPS, runs tokenwatt soak, fetches every raw file and terminates the instance. It also restores the cloud firewall it opened.