← Lab notes

RTX 4090 vs L40S vs A100 vs A40 for LLM inference: what an 8B model really costs per token

We served Qwen3-8B on four cheaper GPUs rented from RunPod Secure Cloud under the same latency target. An RTX 4090 with FP8 weights was cheapest at $0.90 per million output tokens; the A100, the 4090 in BF16 and the L40S landed at $1.06–1.13; the A40 cost $1.61. Every card was limited by occasional slow tokens rather than average speed, and a 30B mixture-of-experts model on an H100 was about 3× cheaper per token than any of them.

Our earlier tests used H100s and H200s. Many teams start on cheaper GPUs, though: an RTX 4090 rents for under a dollar an hour, and an A100 or L40S costs half as much as an H100. So we asked the same question on four of them: what does a million served tokens cost once the GPU has to meet a latency target?

Results for Qwen3-8B on vLLM 0.31.0, one GPU each, RunPod Secure Cloud prices on the day:

GPU Price Weights Capacity at SLO Output tokens/s Cost per million output tokens
RTX 4090 (24 GB) $0.89/h FP8 0.74 req/s 276 $0.90
A100 SXM (80 GB) $1.79/h BF16 1.13 req/s 469 $1.06
RTX 4090 (24 GB) $0.89/h BF16 0.62 req/s 223 $1.11
L40S (48 GB) $1.09/h BF16 0.72 req/s 272 $1.11
L40S (48 GB) $1.09/h FP8 0.72 req/s 268 $1.13
A40 (48 GB) $0.59/h BF16 0.20 req/s 102 $1.61
  • The RTX 4090 with FP8 weights was cheapest, 15% below the A100.
  • Four of the six configurations cost $1.06–1.13 per million tokens. The A100 serves the most users per GPU, but it costs twice as much per hour.
  • The A40 is the one to skip. It's the cheapest per hour and the most expensive per token.
  • For comparison, a 30B mixture-of-experts model on an H100 cost $0.29 per million tokens in our H100 tests. That's about 3× cheaper than any card here, for a bigger model. More on that below.

The setup

Each GPU ran the same container (vllm/vllm-openai:v0.31.0) with default settings, serving Qwen3-8B, the largest common model that fits all four cards in BF16. On the two Ada-generation cards (RTX 4090, L40S), which support FP8 natively, we also ran the FP8 version of the same model. The A100 and A40 don't support FP8.

The workload was a chat-like mix of prompt lengths (median 1,024 input and 256 output tokens; mean about 1,700 and 420) arriving at random. TokenWatt Bench searched for the highest request rate where the p99 time to first token stayed under 2 s and the p99 time per output token under 50 ms, then confirmed it with two 2-minute repeats. The methodology has the details. Accuracy was 90–94% on our 50-question GSM8K check for every configuration, the same within the margin of error.

Before each run, our open-source checker tokenwatt-check tested the card, and its new watch mode kept checking every 15 seconds during the benchmark:

Node check RTX 4090 L40S A100 SXM A40
Power limit 450 W (stock) 350 W 400 W 300 W
BF16 matmul 160 TFLOPS 166–175 TFLOPS 258–261 TFLOPS 107 TFLOPS
Memory copy bandwidth 0.92 TB/s 0.65 TB/s 1.75 TB/s 0.56 TB/s
During the benchmark no changes no changes no changes no changes

Every node we benchmarked was healthy. One of the four L40S pods we rented that day, also on Secure Cloud, had its power limit lowered to 315 W of 350. Its session ended early because of a bug in our own script, so none of the results here come from it. While calibrating, we also found that our checker's provisional L40S reference was about 50% too high and marked healthy L40S cards as FAIL. Version 0.3.1 uses the values measured here.

Why the tail, not the average, sets capacity

None of these cards was close to the 50 ms limit on average. Median time per output token at capacity ranged from 14 ms (RTX 4090, FP8) to 34 ms (A40). What stopped them was the slowest 1% of tokens.

The A100 shows it most clearly. When we first tested it at 1.4 requests/s, its median time per token was 17 ms, but its p99 was 64 ms. Every so often a long prompt (our workload includes some up to 8,000 tokens) gets processed in the same step as everyone else's next token, and those tokens stall. The capacity that holds the p99 under 50 ms is 1.13 requests/s, well below what the average would suggest.

That's also why FP8 made the L40S faster but not more capable: its median time per token dropped from 29 to 17 ms, but capacity stayed at 0.72 requests/s because the stalls from long prompts didn't go away. On the RTX 4090, FP8 did raise capacity by 19%, for a different reason: in BF16 the model's 16 GB of weights leave only about 5 GB of the card's 24 GB for the KV cache, and requests had to queue. FP8 halves the weights and frees memory for more requests at once.

Spreading out long prompts is a serving setting, not a hardware property. On a 70B model, smaller prefill chunks raised capacity by up to 15% in our tests. We ran defaults here, so every card has some headroom left for tuning.

Why a bigger model on an H100 is cheaper per token

An 8B dense model reads all 8 billion parameters for every token. Qwen3-30B-A3B, the mixture-of-experts model we used on H100s, has 30 billion parameters but reads only about 3 billion per token, on a GPU with 1.7–4.8× the memory bandwidth of these cards. Under the same latency target, an H100 served it at about 9 requests/s, for $0.29 per million tokens.

So if the model is your choice to make, a mixture-of-experts model with few active parameters can cut cost per token more than any choice between these GPUs. If you need a specific dense model of this size, the table above is the comparison that matters.

Which one?

  • Cheapest per token for a small model: RTX 4090 with FP8 weights, if your provider offers it in a datacenter with the reliability you need.
  • Most users per GPU, fewer machines to manage: A100 SXM 80 GB. It costs about 15% more per token than the 4090 with FP8, but serves 50% more users per card and has room for bigger models and longer contexts.
  • L40S: about the same cost per token as the A100 here, with 48 GB of memory. Check the power limit before you start: one of the four we rented was capped.
  • A40: skip it for LLM serving. It's the cheapest per hour and the most expensive per token.
  • Whatever you rent: run pip install tokenwatt-check && tokenwatt-check first, and leave tokenwatt-check --watch running alongside your workload.

Current prices: A100, L40S, all GPUs.

Limits

  • Quick mode. Two 2-minute repeats per configuration. On H100s, full mode (10-minute windows, five repeats) found 15–20% lower capacity, so treat these as optimistic by about that much. The ranking should hold, since the same method applied to every card.
  • One node per configuration, one model, default settings, one platform (RunPod Secure Cloud) and its prices on one day.
  • No RTX 5090. None was available on RunPod during our test, on Secure or Community Cloud.

Cost of these tests: $8.07, including two RTX 5090 pods that never got a network address and were deleted after a few minutes, and two earlier attempts whose search started at too high a rate.