What we solve ·Where do I put the budget?
Vanguard · Guide
What does your AI actually cost, depending on where it runs?
Upstairs they want to bring it in-house and I am the one who will have to operate whatever they decide.
Your break-even point calculated with your volumes, and the failure mode you take on if you decide to self host.
Where inference runsSend this page to whoever decides
This sounds like you if
- There is pressure to bring inference in house, for security or for cost.
- There is a hardware purchase proposal whose break-even point has not been calculated.
- The consumption bill grows and the spend is not attributed by use case.
Use this today, without hiring anyone
The napkin calculation, which dismantles or confirms the decision in twenty minutes and you can do today.
- Pull your daily token consumption, the units of text your vendor bills you in. Daily, not the monthly average. From the peak, not the mean: hardware gets sized for the peak, and that is the error in the typical calculation.
- Multiply it by the price per million of the API you use today and annualize. That is your comparison floor.
- To the cost of your own hardware, add the three things that get left out: power and space, three year depreciation and the time of the person who is going to operate it. That third one is what changes the result.
- Compare, and be suspicious if it comes out similar. At moderate volumes, self-hosting comes out more expensive. If it comes out cheap, check whether you counted the operation.
And the part that is not a calculation: if you self-host, you inherit the entire failure mode: alignment, patching, abuse monitoring, incident response and versioning. Write next to each one who would do it in your organization. The ones left without a name are the real cost of the decision.
How we solve it
The method, not the promise.
- Real consumption gets measured, in peaks and not in averages, because hardware gets sized for the peak.
- The three modes get modeled with the operation included: a frontier vendor service, a service over open weight models, and self-hosting.
- Your break-even point gets calculated, with sensitivity over the two or three assumptions that move it.
- The failure mode being transferred gets inventoried and every element gets a named owner.
- The recommendation gets written with the condition that would change it, so it still works a year from now.
What you receive
- The measurement of your real and projected consumption, in peaks.
- The comparative cost model across the three modes, editable and with open assumptions.
- The break-even point with its sensitivity analysis.
- The inventory of the transferred failure mode, with an owner and a staffing cost.
The proof that applies here
- Operations and infrastructure leadership over 60,000 servers, with cost and capacity as a personal responsibility.
- License elimination program worth more than 100 million dollars a year.
- 19 years selling infrastructure from HPE and IBM and buying it from Citi.
Before you hire
Model supply chain security is flagged as a warning in the report and referred to a specialist in that discipline.
Self-hosting may save money or not, and the answer comes out of the arithmetic with your volumes. What matters is that whoever does it has operated their own infrastructure and knows what it really costs to keep it running for a full year.
If you self host, who patches, who monitors abuse and who answers on a Sunday?
If the problem is a different one
