On-premise, cloud or API: where to run enterprise AI
After the model choice comes the twin question, where to run it, and the answer draws the setup the company will live with for years: the costs that will become structure, the skills to hire, the constraints every new use case will inherit. The options on the table are three, and before the crite...
After the model choice comes the twin question, where to run it, and the answer draws the setup the company will live with for years: the costs that will become structure, the skills to hire, the constraints every new use case will inherit. The options on the table are three, and before the criteria it pays to name them precisely, because committees often blur them. The managed API: the model lives at the provider, the company consumes capacity on demand. The dedicated cloud: the model, open or licensed, runs on infrastructure the company rents from a hyperscaler, inside a perimeter the company controls. On-premise: own hardware, in one's own data centre, under one's own physical responsibility.
The temptation to treat it as a matter of allegiance should be defused immediately, as with models: the right question is per use case, and the target setup of mature organisations is almost always hybrid.
The constraint commands, the rest negotiates
The criterion that decides first, when it exists, is the non-negotiable constraint on data. Some information, by sector regulation, by client clause or by nature (think of certain health, judicial or defence data), cannot cross defined borders, and there the discussion shortens: the choice is between dedicated cloud with contractually solid residency guarantees, sovereign cloud formulas where the sector demands them, or on-premise when even the rented perimeter is not enough. The serious work, in these cases, lies in separating real constraints from presumed ones, because "our data cannot live in the cloud" is a sentence that in assessments turns out, every other time, to be a policy nobody ever wrote that someone remembers badly.
Where the absolute constraint is absent, the negotiable criteria come in, and the first is the load profile. The API shines with variable or unpredictable volumes, because it turns everything into marginal cost; owned or rented infrastructure starts paying for itself with high and constant loads, when GPUs work instead of waiting. Latency adds its weight in use cases inline with operations (a system answering an operator in real time tolerates network round-trips badly), and deep customisation adds its own, because extensive fine-tuning and total version control demand infrastructure you command.
The costs the slides never show
The economic comparison between the three routes suffers from a systematic distortion: visible prices get compared, cost per token against GPU rental fees, and the lines that decide the total get omitted. For dedicated cloud and on-premise, the omitted lines are people (engineers able to serve models with production-grade reliability, the most contested profile on the market), actual hardware utilisation (a GPU paid for and idle is the purest cost in existence), and refresh (models improve at such a pace that infrastructure sized today may be serving an obsolete model in eighteen months). For the API the omitted line is the mirror image, growth: the marginal cost that made starting easy scales linearly with success, and high-volume use cases need their maths redone every few months.
The practical consequence is that the break-even point exists, moves over time, and must be computed on one's own volumes with every line included. In the accounts we review, the most frequent error remains underestimating the people and overestimating the utilisation.
Hybrid as the destination, abstraction as the premise
With the criteria lined up, the architecture mature organisations converge on is layered: managed APIs for the long tail of low-volume use cases and for whatever demands frontier capability, a dedicated perimeter (cloud or, more rarely, on-premise) for constant loads and for data that does not travel, and a written rule assigning every new use case to the right layer based on constraints, volumes and latency, rather than on the preferences of the proposing team.
The technical premise that makes the hybrid governable, and that should be decided at the start when it costs little, is the abstraction layer: applications talk to an internal interface, never directly to the single provider or the single cluster, so that moving a workload from one layer to another is a configuration change and not a rewrite. It is the same insurance against lock-in we discuss for models, applied to infrastructure, and it doubles as a negotiating position: better terms are obtained from anyone when leaving is technically easy.
The setup choice, in short, is an exercise in constraints before preferences, and the fastest way to get it wrong is by imitation, copying the architecture of a company whose constraints differ from yours. If you are drawing yours, or renegotiating an inherited one, the starting point is the inventory of real constraints, volumes and skills: calendly.com/fabiolalli/zerofive, or hello@zerofive.ai. And for the next committee, a control question that requires no slides: of the constraints keeping you out of the cloud today, or inside it, how many exist in writing?