Llama-3.3-70B FP8 on 1× H200: prefill chunk size (max-num-batched-tokens)
Quick-mode comparison of vLLM's default max_num_batched_tokens (8192) against 2048 and 1024 on one RunPod Secure Cloud H200 (US-CO-1), 2026-10-09. Same node, model and workload for all three. The 15% gain at 1024 is about the size of quick-mode run-to-run variation; treat it as a lead. MBU recomputed with FP8 weight bytes (the original run files mis-detected compressed-tensors FP8).
Results
Each value is the median of 2 runs.
SLO goodput
Tokens per joule
Decode bandwidth utilization
Axis runs to 100% of rated bandwidth.
GPU board power
GPU board power only — excludes CPUs, memory, NICs, fans and PSU losses.
All numbers
Cards for equal goodput = 64 × reference goodput ÷ configuration goodput, rounded up.
| Metric | default (8,192 tokens/step) | 2,048 tokens/step | 1,024 tokens/step |
|---|---|---|---|
| Engine build | vLLM 0.31.0 · KV BF16 | vLLM 0.31.0 · KV BF16 | vLLM 0.31.0 · KV BF16 |
| Decode bandwidth utilization (MBU) | 62% | 61% | 61% |
| GPU board power | 0.64 kW | 0.64 kW | 0.64 kW |
| SLO goodput | 234 tok/s | 250 tok/s | 269 tok/s |
| Tokens per joule | 0.37 | 0.39 | 0.42 |
| TTFT p99 | 1,226 ms | 1,279 ms | 1,316 ms |
| TPOT p99 | 37.1 ms | 38.3 ms | 36.4 ms |
| SLO met | yes | yes | yes |
| Accuracy (gsm8k-5shot) | 94.0 | 92.0 | 92.0 |
| Accuracy Δ vs. reference | reference | -2.0 pts | -2.0 pts |
| Cards for equal goodput | 64 | — | — |
Setup
Any change to these fields makes it a different report.
- Accelerator
- 1× NVIDIA H200
- Rated HBM BW
- 4.8 TB/s per GPU
- Driver / runtime
- 580.178.04
- Host
- RunPod secure cloud pod,
- Engine
- vLLM 0.31.0
- Kernel libraries
- vllm/vllm-openai:v0.31.0
- Model
- Llama-3.3-70B-Instruct-FP8-dynamic (Dense)
- Quantization
- FP8 weights (compressed-tensors), BF16 KV cache
- Parallelism
- 1 replica × TP1
- Workload trace
- chat-mix-v1
- Input tokens
- median 1,024 · p90 4,096
- Output tokens
- median 256 · p90 1,024
- Arrival
- Poisson, rate swept to SLO boundary
- Power source
- GPU board power sum (NVIDIA) · 1 Hz · 120 s steady state
- Runs per config
- 2 (variance band ±8.3%)
- Accuracy gate
- gsm8k-5shot · max drop 1.0 pts
Not covered
Do not extrapolate this report to the following.
- Whole-node power (no BMC access on this host)
- Context lengths beyond the chat-mix-v1 trace
- Quick mode: short windows, few repeats — validation only, not for publishing
Reproduce
Run on the same hardware and software versions. Results should fall within ±8.3%.
tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-default' --role other --out runs/llama70b-default.json tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-mnbt2048' --role other --out runs/llama70b-mnbt2048.json tokenwatt run --spec lambda-llama-3.3-70b-instruct-fp8-dynamic-tp1-dp1.spec.yaml --label 'llama70b-mnbt1024' --role other --out runs/llama70b-mnbt1024.json tokenwatt report runs/llama70b-default.json runs/llama70b-mnbt2048.json runs/llama70b-mnbt1024.json --title 'Llama-3.3-70B FP8 on 1× H200: prefill chunk size (max-num-batched-tokens)'