Put a ruler on your fleet.
Bench is opening early access to a small group of US inference teams and GPU clouds. Tell us what you run and we'll get in touch.
Support matrix
What Bench is being built to support in early access. If yours isn't listed, ask.
| Area | Targets | Status |
|---|---|---|
| AMD Instinct | MI300X · MI325X · MI355X | Targeted |
| NVIDIA | H100 · H200 · B200 | Targeted |
| Engines | vLLM · SGLang · TensorRT-LLM · any OpenAI-compatible endpoint | Targeted |
| Node power | Redfish · IPMI DCMI · metered PDU | Targeted |
| Board power & clocks | amd-smi · NVML / DCGM | Targeted |
| Multi-node, disaggregated serving | Prefill/decode split, multi-node expert parallelism | Later |
One container
No agents to install for Bench, no changes to your serving stack. It's a client plus read-only telemetry.
Stays in your cluster
Weights, prompts and logs never leave your environment. Publishing a report is always your choice.
Simple pricing
Open core is free. The paid tier is priced per GPU per month. Design partners get it at no cost during early access.
Questions
Does Bench need access to our model weights?
No. Bench sends requests to your existing inference endpoint and reads node telemetry. For bandwidth utilization it needs the model's architecture (layer shapes, expert count, quantization), which it reads from the public config — never the weights themselves.
What if our nodes don't expose BMC power readings?
Bench will still run and report GPU board power, clearly labeled. Node-level tokens per joule requires Redfish, IPMI DCMI or a metered PDU; the report says which source was used.
Will running Bench disturb production traffic?
Bench drives load to find the SLO boundary, so run it on a node that's drained or reserved for testing. A typical comparison of three configurations takes two to three hours.
You sell tuned images. Why should we trust your numbers?
Because you don't have to. The method and measurement core are open, every Tuned result sits next to the vendor's latest recommended configuration, and you run the measurement yourself on your own hardware.
Do you publish our results?
Never without written permission. Public reports on this site come from our own lab runs or from partners who chose to share.
Tell us what you're running.
Accelerator model and count, inference engine, the models you serve, and what you want to know. We'll reply within two business days.