Qwen3-8B is a dense 8-billion-parameter model. In BF16 its weights take about 16 GB, so it fits on every GPU from a 24 GB RTX 4090 up. On Ada-generation cards (RTX 4090, L40S) we also measured the FP8 version, which halves the weights.
Results
One GPU each, vLLM 0.31.0 with default settings, chat-like workload (median 1,024 input and 256 output tokens), latency target p99 time to first token ≤ 2 s and p99 time per output token ≤ 50 ms. Quick mode (2-minute windows); RunPod Secure Cloud prices on 2026-10-11.
| GPU (price) | Weights | Capacity | Output tokens/s | Median / p99 time per token | Cost per million output tokens |
|---|---|---|---|---|---|
| RTX 4090, 24 GB ($0.89/h) | FP8 | 0.74 req/s | 276 | 13.7 / 21.6 ms | $0.90 |
| A100 SXM, 80 GB ($1.79/h) | BF16 | 1.13 req/s | 469 | 16.4 / 33.9 ms | $1.06 |
| RTX 4090, 24 GB ($0.89/h) | BF16 | 0.62 req/s | 223 | 21.1 / 29.5 ms | $1.11 |
| L40S, 48 GB ($1.09/h) | BF16 | 0.72 req/s | 272 | 29.1 / 37.8 ms | $1.11 |
| L40S, 48 GB ($1.09/h) | FP8 | 0.72 req/s | 268 | 16.8 / 25.3 ms | $1.13 |
| A40, 48 GB ($0.59/h) | BF16 | 0.20 req/s | 102 | 34.0 / 38.2 ms | $1.61 |
Accuracy on our GSM8K check was 90–94% for every configuration. Full-mode measurements (10-minute windows) typically read 15–20% lower capacity; the ranking should hold.
What we learned
- Cheapest per token: the RTX 4090 with FP8 weights, 15% below the A100. The A40 was the cheapest per hour and the most expensive per token.
- Capacity is set by the slowest tokens: median time per token was well under the 50 ms target on every card. Occasional stalls, when long prompts are processed in the same step as everyone else's next token, decided capacity.
- FP8 helped the RTX 4090 (+19% capacity) but not the L40S: on the 4090, BF16 weights left only about 5 GB for the KV cache, and FP8 freed it. On the L40S, FP8 made tokens faster on median but the stalls stayed. More on FP8 vs BF16.
- A bigger model can be cheaper: Qwen3-30B-A3B, a mixture-of-experts model with 3 billion active parameters, cost about $0.29 per million tokens on an H100, about 3× less than Qwen3-8B on any of these cards.
Full write-up: L40S vs A100 vs RTX 4090 vs A40. Per-configuration reports with raw data are listed on the reports page.
FAQ
What's the cheapest GPU to run Qwen3-8B?
In our test, an RTX 4090 with FP8 weights at $0.90 per million output tokens at RunPod's price, followed by the A100 at $1.06. The A40 was cheapest per hour but most expensive per token.
How much GPU memory does Qwen3-8B need?
About 16 GB for the weights in BF16 and 8 GB in FP8, plus memory for the KV cache, which grows with the number of concurrent requests and their length. On a 24 GB card, BF16 leaves only about 5 GB for the cache.
How many users can one GPU serve with Qwen3-8B?
Fewer than you might expect within a strict latency target: 0.2 to 1.1 requests per second on these cards for a chat workload. Each answer takes several seconds to stream, so on average 3 (A40) to 9 (A100) requests were being answered at any moment.