← Methodology

Glossary

Time to first token (TTFT): what it measures and what good looks like

Time to first token (TTFT) is the delay between sending a request to an LLM and receiving the first token of the answer. It's what users feel as "thinking time" before text appears. Here's how it's measured, what drives it, and the values we measured on real GPUs.

Time to first token (TTFT) is the time from sending a request to an LLM until the first token of the response arrives. With streaming output, it's the pause before text starts appearing.

What's inside TTFT

TTFT = time waiting in the queue + time to process the prompt (prefill) + network time.

  • Queueing: if the server is busy, new requests wait. Near a GPU's capacity this part dominates and grows fast.
  • Prefill: the model reads the whole prompt before it can produce the first token. Longer prompts take longer; compute speed matters here.
  • Network: usually small compared with the other two.

TTFT is different from time per output token (TPOT), the gap between later tokens, which is mostly limited by memory bandwidth. A server can have fast TPOT and slow TTFT, or the reverse.

Why measure the p99, not the average

Averages hide the bad experiences. We set targets on the 99th percentile: 99 of 100 requests must see their first token within the target. Our default is p99 TTFT ≤ 2 seconds together with p99 time per output token ≤ 50 ms. The highest request rate that keeps both is the GPU's SLO capacity.

Typical values we measured

Setup Load TTFT median / p99
H100, Qwen3-30B-A3B, vLLM 0.31.0 9 req/s 196 / 403 ms
H100, Qwen3-30B-A3B, SGLang 0.5.21 9 req/s 67 / 241 ms
A100 SXM, Qwen3-8B, vLLM 1.13 req/s (capacity) 120 / 987 ms
RTX 4090, Qwen3-8B FP8, vLLM 0.74 req/s (capacity) 178 / 1,229 ms

Sources: SGLang vs vLLM on H100, L40S vs A100 vs RTX 4090 vs A40. Workload: chat-like prompts, median 1,024 input tokens.

Two things stand out. The serving engine matters: SGLang returned the first token about 3× sooner than vLLM on the same GPU at the same load. And near capacity, the p99 climbs far above the median as requests start to queue.

How to improve TTFT

  • Leave headroom: run below capacity, since queueing grows sharply near the limit.
  • Try another engine or scheduler: SGLang's median TTFT was a third of vLLM's in our H100 test.
  • Use prefix caching when many requests share a long system prompt.
  • Faster prefill hardware helps long prompts: compute, not memory bandwidth, limits prefill.

FAQ

What is a good time to first token?

For chat, under about 500 ms at the median feels responsive and under 2 seconds at the 99th percentile is a common target. Agents and coding tools that chain many calls benefit from lower values.

What's the difference between TTFT and TPOT?

TTFT is the wait before the first token. TPOT (time per output token) is the gap between each token after that. TTFT is driven by queueing and prompt processing, TPOT mostly by memory bandwidth and batch size.

Why does TTFT increase under load?

Requests wait in a queue when the GPU is busy. Near the GPU's capacity the queue builds quickly, so the slowest requests (the p99) wait much longer than the median.