modelled / 2026-08-13

Self-hosting usually loses on token cost

A published break-even model for open-weight inference against public API prices.

Method

Hardware
Rented 8x NVIDIA H100 80GB node serving a 671B-class mixture-of-experts model
Workload
About 620 output tokens/s at roughly 100 concurrent requests
Repetitions
None - arithmetic from published prices, not an instrumented run
Tool
Calculation on published hourly rates and vendor price lists

Results

Self-hosted - full utilisation$10 / 1M output tokensOptimistic 24/7 batch case
Self-hosted - 50% utilisation$20 / 1M output tokensSame fixed hardware
Self-hosted - 20% utilisation$50 / 1M output tokensCloser to a bursty internal workload
Budget open-weight API$0.87 / 1M output tokensPublished list price in the source

Takeaway

Sovereignty and evidence are the self-hosting argument. A token-price saving is not.