What AI integration really costs in a mid-sized company: the lines nobody quotes
Model calls are the visible part, and usually the smallest. The lines that blow a budget sit elsewhere, and they are predictable, provided you listed them before starting.
Christophe Bellec ·
Short answer
The cost of an AI integration project splits into three blocks: getting the data and the surrounding systems into shape, often the heaviest; building the integration itself; and running it, which includes a per-use variable cost. Model calls are usually a minority share of the total. A proof of concept costing a few days says nothing about the cost of going to production, which is a software project in its own right.
Why estimates slip
A convincing demo can be built in a day. It gives the impression the hard part is done, when in fact it set aside everything expensive: real data, access rights, edge cases, failures, integration with the rest of the system.
The gap between a prototype and an operable service is not specific to AI: it is the same gap as for any software. What is specific to AI is that the gap looks unusually easy to cross, and is unusually expensive to actually cross.
The cost lines, most-forgotten first
1. Getting the data into shape
The most consistently underestimated line. Documents are scattered across a file server, a mailbox and three SaaS tools; reference data contains duplicates; nobody knows which version of a procedure is authoritative. No AI fixes that.
The good news: this work has value regardless of the AI project. The bad news: it has to happen first, and it cannot be fully outsourced, because it requires knowing what is authoritative inside the company.
2. Integration with existing systems
A useful AI feature is not a separate interface: it lives inside the tools teams already use. That means APIs, authentication, error handling, sometimes opening up a system never designed to be called from outside. This is ordinary development work, and usually the largest line of the build phase.
3. Access rights
As soon as the service touches data that is not readable by everyone, you have to decide how permissions apply, implement them, and be able to demonstrate them. Handled up front, it is a design constraint. Handled at the end, it is a rewrite.
4. Evaluation
An AI system has no binary “it works”. You need a reference set of questions with expected answers, otherwise nobody can say whether a change improved or degraded the service. Building that set takes domain time, not engineering time, which is precisely why it gets skipped.
5. Operations
Logging, monitoring, alerting, cost tracking, handling provider outages, model version upgrades. A replaced model can change how the service behaves: without an evaluation set, that upgrade becomes a gamble.
6. The per-use variable cost
The only genuinely new line compared with a classic software project: the service costs money every time it is used. That cost depends on the size of the context sent, the model chosen and the usage volume: three parameters the architecture directly controls.
7. Change management
A tool nobody uses costs one hundred per cent of its budget for zero benefit. People need training, they need the limits explained, and there has to be a plan for when the system gets it wrong, because it will.
How to size it before committing
Without knowing your context, nobody can give an honest figure. These questions, however, place the order of magnitude, and they can be answered in a few days:
- Which data does the use case rest on, and what state is it genuinely in?
- Does an API already exist to reach it, or does it have to be written?
- Who may see what, and is that authorisation model already implemented somewhere?
- How often will the service be used, and by how many people?
- What happens if the answer is wrong? A draft to review and a binding decision do not call for the same level of guarantees.
- Who will operate the service once it ships?
Question 5 moves the budget the most. An assistant whose output a human reviews, and a system that triggers an action without validation, are not the same project or the same order of magnitude.
How to cut the bill without hurting the result
- Start with one measurable use case, with a named user who needs it every week.
- Read data live rather than building an index to maintain, where possible.
- Pick the model after the architecture: a smaller model is enough for most steps of a pipeline.
- Shrink the context you send: retrieval quality drives cost more than model choice does.
- Cache what is stable, and measure cost per use from day one.
- Agree a stopping criterion before starting: if the proof of concept misses it, you stop.
A proof of concept is not a version 1
A proof of concept exists to settle one precise uncertainty, in a few days, on real data: does retrieval find the right documents? Does the model produce something usable? Is the per-use cost acceptable? Success authorises investing in what comes next; it says nothing about what that will cost.
Confusing the two is the most common reason a budget is announced three times too low. Separating the phases explicitly, and budgeting them separately, removes the surprise.
Frequently asked questions
Which cost line is forgotten most often?
Getting the data into shape, followed by evaluation. The first is invisible in a demo; the second interests nobody until the system has produced its first wrong answer in production.
Do model calls make up most of the budget?
Rarely. On an integration project they are usually a minority of the total, far behind development and data preparation. They become significant at high volume, and at that point architecture is what brings them down.
Can we start small?
It is the only sensible approach: one use case, one typical user, one success criterion agreed in advance. The proof of concept runs a few days and settles a precise question, rather than opening a programme with no visible end.
How do we know it is worth it before starting?
By quantifying the time actually spent on the target task today and comparing it with the estimated integration cost. That is what an AI and business audit is for: sorting use cases by impact and complexity, including the ones best left alone.