← All reports

Report · 2026-10-09

Llama-3.3-70B FP8 on 1× H200: prefill chunk size (max-num-batched-tokens)

1× NVIDIA H200 Llama-3.3-70B-Instruct-FP8-dynamic FP8 weights (compressed-tensors), BF16 KV cache vLLM 0.31.0 TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms

Quick-mode comparison of vLLM's default max_num_batched_tokens (8192) against 2048 and 1024 on one RunPod Secure Cloud H200 (US-CO-1), 2026-10-09. Same node, model and workload for all three. The 15% gain at 1024 is about the size of quick-mode run-to-run variation; treat it as a lead. MBU recomputed with FP8 weight bytes (the original run files mis-detected compressed-tensors FP8).

ConclusionUnder SLO (TTFT p99 ≤ 2.0 s · TPOT p99 ≤ 50 ms), serving the same billable throughput as 64 cards on default (8,192 tokens/step) takes: 2,048 tokens/step failed the accuracy gate (-2.0 pts) and is excluded from capacity math, despite 250 tok/s goodput; 1,024 tokens/step failed the accuracy gate (-2.0 pts) and is excluded from capacity math, despite 269 tok/s goodput.

Results

Each value is the median of 2 runs.

default (8,192 tokens/step)2,048 tokens/step1,024 tokens/step

SLO goodput

output tok/s per node · higher is better
default (8,192 tokens/step)234default (8,192 tokens/step): 234 output tok/s per node
2,048 tokens/step2502,048 tokens/step: 250 output tok/s per node
1,024 tokens/step2691,024 tokens/step: 269 output tok/s per node

Tokens per joule

tok/J, GPU board power · higher is better
default (8,192 tokens/step)0.37default (8,192 tokens/step): 0.37 tok/J, GPU board power
2,048 tokens/step0.392,048 tokens/step: 0.39 tok/J, GPU board power
1,024 tokens/step0.421,024 tokens/step: 0.42 tok/J, GPU board power

Decode bandwidth utilization

% of rated HBM bandwidth
default (8,192 tokens/step)62%default (8,192 tokens/step): 62 % of rated HBM bandwidth
2,048 tokens/step61%2,048 tokens/step: 61 % of rated HBM bandwidth
1,024 tokens/step61%1,024 tokens/step: 61 % of rated HBM bandwidth

Axis runs to 100% of rated bandwidth.

GPU board power

kW, steady state
default (8,192 tokens/step)0.64default (8,192 tokens/step): 0.64 kW, steady state
2,048 tokens/step0.642,048 tokens/step: 0.64 kW, steady state
1,024 tokens/step0.641,024 tokens/step: 0.64 kW, steady state

GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.

All numbers

Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.

Metricdefault (8,192 tokens/step)2,048 tokens/step1,024 tokens/step
Engine buildvLLM 0.31.0 · KV BF16vLLM 0.31.0 · KV BF16vLLM 0.31.0 · KV BF16
Decode bandwidth utilization (MBU)62%61%61%
GPU board power0.64 kW0.64 kW0.64 kW
SLO goodput234 tok/s250 tok/s269 tok/s
Tokens per joule0.370.390.42
TTFT p991,226 ms1,279 ms1,316 ms
TPOT p9937.1 ms38.3 ms36.4 ms
SLO metyesyesyes
Accuracy (gsm8k-5shot)94.092.092.0
Accuracy Δ vs. referencereference-2.0 pts-2.0 pts
Cards for equal goodput64——

Setup

Any change to these fields makes it a different report.

Accelerator
1× NVIDIA H200
Rated HBM BW
4.8 TB/s per GPU
Driver / runtime
580.178.04
Host
RunPod secure cloud pod,
Engine
vLLM 0.31.0
Kernel libraries
vllm/vllm-openai:v0.31.0
Model
Llama-3.3-70B-Instruct-FP8-dynamic (Dense)
Quantization
FP8 weights (compressed-tensors), BF16 KV cache
Parallelism
1 replica × TP1
Workload trace
chat-mix-v1
Input tokens
median 1,024 · p90 4,096
Output tokens
median 256 · p90 1,024
Arrival
Poisson, rate swept to SLO boundary
Power source
GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
Runs per config
2 (variance band ±8.3%)
Accuracy gate
gsm8k-5shot · max drop 1.0 pts

Not covered

Do not extrapolate this report to the following.

  • Whole-node power (no BMC access on this host)
  • Context lengths beyond the chat-mix-v1 trace
  • Quick mode: short windows, few repeats — validation only, not for publishing

Reproduce

Run on the same hardware and software versions. Results should fall within ±8.3%.

shell
tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-default' --role other --out runs/llama70b-default.json
tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-mnbt2048' --role other --out runs/llama70b-mnbt2048.json
tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-mnbt1024' --role other --out runs/llama70b-mnbt1024.json

tokenwatt report runs/llama70b-default.json runs/llama70b-mnbt2048.json runs/llama70b-mnbt1024.json --title 'Llama-3.3-70B FP8 on 1× H200: prefill chunk size (max-num-batched-tokens)'

Download raw JSON