measured / 2026-08-13
What current hardware serves per node
Published MLPerf Inference v6.0 results for open-weight models on single nodes.
Method
- Hardware
- 8x NVIDIA B200 for Llama 3.1 8B; 4x NVIDIA GB300 for GPT-OSS 120B
- Workload
- MLPerf Offline, Server and Interactive scenarios with fixed accuracy and latency constraints
- Repetitions
- As required by the audited MLPerf submission rules
- Tool
- MLPerf Inference harness with vLLM, submitter tuned
Results
Llama 3.1 8B - offline160,403 tokens/s8x B200, closed division
Llama 3.1 8B - server130,008 tokens/sLatency constrained
Llama 3.1 8B - interactive128,750 tokens/sTighter latency constraint
GPT-OSS 120B - server53,463 tokens/s4x GB300
Takeaway
Capacity is now a concurrency and context-length sizing problem, not a claim that public providers own impossible hardware.