How many GPUs does an internal ChatGPT need? Fewer than you think — and that's the cost problem
Sizing private LLM inference from a measured H100 benchmark. One H100 can carry the peak chat traffic of roughly 12,000 employees, but at 1,000 employees it sits 99.6% idle and a million tokens costs $67.70 instead of $0.30. Utilization, not GPU count, decides the economics. Includes a calculator.
Companies that can't send their data to an outside API keep asking the same first question: how many GPUs do we need to run our own ChatGPT for our employees?
We can answer it from a measurement rather than a spec sheet. In lab notes #2, one H100 served Qwen3-30B-A3B at 3,996 output tokens per second (9.5 requests per second) while holding p99 time-to-first-token under 2 seconds and p99 time-per-output-token under 50 ms. That's the throughput users actually experience as responsive, not a peak number.
The answer turns out to be "very few GPUs". The harder question is what each token costs once you have them.
The assumptions
| Assumption | Value | Why |
|---|---|---|
| Daily active users | 50% of employees | typical of internal assistant rollouts after the novelty wears off |
| Requests per active user | 20 per workday | chat questions, rewrites, summaries |
| Busiest hour | 20% of the day's requests | traffic bunches up mid-morning |
| Request shape | our chat-mix-v1 trace: input median 1,024 tokens, output median 256 (mean ≈ 421) |
the workload we measured |
| Planning headroom | size GPUs to 70% of the measured SLO capacity, plus one spare | don't run at the edge; survive a failed GPU |
| GPU price | $4.29 per H100-hour, rented around the clock | Lambda on-demand list price on 2026-10-07 |
Change any of these in the calculator below. The only number that comes from our lab is the per-GPU capacity.
The result
| Employees | Requests/day | Peak req/s | H100s (+1 spare) | Average utilization | Cost/month | Cost per 1M output tokens |
|---|---|---|---|---|---|---|
| 1,000 | 10,000 | 0.56 | 1 + 1 | 0.4% | $6,263 | $67.70 |
| 5,000 | 50,000 | 2.78 | 1 + 1 | 2.2% | $6,263 | $13.54 |
| 10,000 | 100,000 | 5.56 | 1 + 1 | 4.4% | $6,263 | $6.77 |
| 25,000 | 250,000 | 13.9 | 3 + 1 | 5.5% | $12,527 | $5.42 |
| 50,000 | 500,000 | 27.8 | 5 + 1 | 7.3% | $18,790 | $4.06 |
| 100,000 | 1,000,000 | 55.6 | 9 + 1 | 8.8% | $31,317 | $3.38 |
At 70% of its SLO capacity, one H100 can carry the peak chat traffic of about 12,000 employees under these assumptions. Most companies need two GPUs, and one of them is the spare.
The same table shows the catch. At full utilization this setup costs $0.30 per million output tokens. At 1,000 employees it costs $67.70, because the GPUs are idle 99.6% of the time. Even at 100,000 employees, average utilization stays under 10%. Chat traffic follows the workday, and the GPUs are paid for around the clock.
What this means
- GPU count is the easy part. For chat on a model of this size, the hardware is small. Don't plan a cluster before you've measured your own traffic.
- Utilization decides the economics. The cost per token falls almost exactly as utilization rises. Ways to raise it:
- Share the GPUs across workloads. Put chat, document summarization, coding assistants and RAG on the same pool instead of buying per project.
- Fill the nights. Batch jobs (classification, extraction, embedding refreshes, evaluation runs) can take the 16 idle hours.
- Rent elastically if you can. On-demand or reserved-plus-burst capacity follows the workday. Owned hardware has to earn its keep 24/7.
- Right-size the model. This is a mixture-of-experts model with about 3 B active parameters. A dense 70 B model needs several times more GPU per request; use it only where the quality difference matters.
- Privacy is the reason to self-host; utilization is how you make it pay. If the data can't leave, compare your effective cost per token, not the full-utilization number, against what the business would accept.
Calculator
Capacity per GPU: 9.5 requests/s and 3,996 output tokens/s at p99 TTFT ≤ 2 s and TPOT ≤ 50 ms — measured on one H100 SXM with Qwen3-30B-A3B-FP8 and vLLM 0.31.0 on our chat-mix trace. Includes one spare GPU. 22 workdays and 730 billed hours per month.
What this doesn't cover
- Other models. The capacity comes from one MoE model with about 3 B active parameters. Dense and larger models serve far fewer requests per GPU, and we haven't measured them yet.
- RAG and agents. Long retrieved contexts and multi-step agent loops change the request shape a lot. Agents can multiply requests per task by ten or more. Measure those traces before sizing.
- Owned hardware. We price rented GPUs. Owned hardware swaps the hourly rate for depreciation, power and operations, but the utilization argument is the same.
- Peaks beyond the busiest hour. Launch days and all-hands demos can spike well above it. The spare GPU and the 70% headroom absorb some of that, not all.
If you're sizing private inference for your own workload, we can measure it on your traffic, your model and your latency targets. Get in touch.