modelled / 2026-08-13
Self-hosting usually loses on token cost
A published break-even model for open-weight inference against public API prices.
Method
- Hardware
- Rented 8x NVIDIA H100 80GB node serving a 671B-class mixture-of-experts model
- Workload
- About 620 output tokens/s at roughly 100 concurrent requests
- Repetitions
- None - arithmetic from published prices, not an instrumented run
- Tool
- Calculation on published hourly rates and vendor price lists
Results
Self-hosted - full utilisation$10 / 1M output tokensOptimistic 24/7 batch case
Self-hosted - 50% utilisation$20 / 1M output tokensSame fixed hardware
Self-hosted - 20% utilisation$50 / 1M output tokensCloser to a bursty internal workload
Budget open-weight API$0.87 / 1M output tokensPublished list price in the source
Takeaway
Sovereignty and evidence are the self-hosting argument. A token-price saving is not.