Free tool · No signup

Self-Host vs API Break-even

The monthly volume where running your own model starts to beat paying per token, including the utilisation, redundancy and engineering time most comparisons leave out.

Output sets your GPU capacity need.

Affects the API side of the comparison.

GPU and provider

Median rate across dedicated GPU providers.

Open-weight model you would run

A 70B model needs more GPUs each and serves far fewer tokens per second.

A GPU bills the same idle. Peaky traffic rarely clears 40%.

One replica is not a production answer.

Per replica, batched.

Compare like for like: a mid tier, not a frontier model.

At your volume
Keep using the API
by $5,245 a month
Break-even volume
828.5M
output tokens/month · 8.3x your current volume
API$720.00/mo
Self-host$5,965/mo
2 GPUs (2 replicas)$4,365
engineering time$1,600

What is moving this answer

Utilisation
30%

Each replica can serve 2.0B output tokens/month at this rate. The other 70% is paid for and idle.

Engineering
$1,600/mo

27% of the self-host total.

If you stay on API
volume, not util

No utilisation fixes this. You need more volume or cheaper GPUs.

One caveat this model cannot price for you: an open-weight model is not the same model as a frontier API. Compare against the tier you would actually replace, not the most expensive one on the menu, or the break-even will flatter self-hosting.

Why the break-even is higher than it looks

Three things the hourly rate does not tell you, and all three push the same way.

Idle GPUs still bill

Utilisation is the variable that decides most of these cases. At 30% you are paying roughly three times the cost per token the spec sheet implies.

Provider beats silicon

The same H100 is ~$3/hr on a specialist cloud and $11 to $13 on a hyperscaler. That 4x spread moves the break-even more than the GPU model does.

Someone has to run it

Deployment, updates, failover and on-call are real hours. At small scale they routinely cost more than the hardware.

The full write-up is in self-hosting vs API. Before changing where inference runs, it is usually worth checking the four cheaper levers that come first.

Self-hosting FAQ

At what volume does self-hosting an LLM become cheaper than an API?+

On representative mid-2026 numbers, an 8B open-weight model on a specialist-cloud H100 breaks even against a mid-tier API somewhere around 800 million output tokens a month, once you include two replicas for redundancy, 30% average utilisation and twenty hours of engineering time. Below that the API wins comfortably. The exact figure moves a lot with your GPU provider and utilisation, which is what this calculator is for.

Why is my GPU utilisation so important?+

Because a GPU bills the same whether it is at 9% load or 90%. Traffic is peaky, so the capacity you can actually use is far below the capacity you buy, and paying for idle silicon is the single most common reason a self-hosting business case fails in practice. At 30% utilisation you are paying roughly three times the notional cost per token you calculated from the spec sheet.

Does it matter which cloud I rent GPUs from?+

Enormously, and usually more than which GPU you pick. The same H100 runs around $3 an hour on specialist GPU clouds and $11 to $13 an hour on the major hyperscalers, roughly a 4x spread for identical silicon. That difference alone can move the break-even volume by a factor of three, so the provider decision often settles the build-versus-buy question before any other variable does.

Should I include engineering time in the comparison?+

Yes, and leaving it out is the most common way these business cases get inflated. Serving infrastructure needs deployment, model updates, capacity planning, failover, monitoring and someone on call. At small scale those hours frequently exceed the GPU bill itself, which is why self-hosting rarely pays for a team serving modest volume no matter how cheap the hardware looks.

Is self-hosting a 70B model worth it?+

Usually not, unless volume is very large or you have a requirement that rules out APIs entirely. A 70B model needs several GPUs per replica and serves far fewer tokens per second than an 8B one, so cost per token rises on both axes at once. Against a cheap mid-tier API it often never breaks even. Against frontier pricing the case is much easier.

What does self-hosting give me besides cost?+

Data residency, predictable spend, no rate limits, no vendor deprecation cycle, and the ability to fine-tune freely on your own data. Those can justify self-hosting well below the cost break-even, and if any of them are hard requirements the arithmetic here is not really the deciding factor. What the calculator prevents is choosing it for cost reasons that do not hold.

Build versus buy, decided on arithmetic

We run the comparison honestly, including the answer where you should stay on the API. Book a free cost review.

Prefer email? [email protected]