measured / 2026-08-13

What current hardware serves per node

Published MLPerf Inference v6.0 results for open-weight models on single nodes.

Method

Hardware
8x NVIDIA B200 for Llama 3.1 8B; 4x NVIDIA GB300 for GPT-OSS 120B
Workload
MLPerf Offline, Server and Interactive scenarios with fixed accuracy and latency constraints
Repetitions
As required by the audited MLPerf submission rules
Tool
MLPerf Inference harness with vLLM, submitter tuned

Results

Llama 3.1 8B - offline160,403 tokens/s8x B200, closed division
Llama 3.1 8B - server130,008 tokens/sLatency constrained
Llama 3.1 8B - interactive128,750 tokens/sTighter latency constraint
GPT-OSS 120B - server53,463 tokens/s4x GB300

Takeaway

Capacity is now a concurrency and context-length sizing problem, not a claim that public providers own impossible hardware.