---
title: "Why AI Budgets Blow Up (And It's Not the Tokens)"
url: https://zerofive.ai/en/blog/strategy/ai-cost-nobody-budgeted-for
canonical: https://zerofive.ai/en/blog/strategy/ai-cost-nobody-budgeted-for
language: en
published: 2026-08-27
updated: 2026-09-18
author: "ZeroFive.AI"
tags: AI costs, AI ROI, AI Rating, AI budget, AI governance
abstract: "Usage-based billing, headcount traded for tokens, PoCs scaled without measuring first. Why AI costs escape the budget, and how to see them coming with an AI Rating."
---

# Why AI Budgets Blow Up (And It's Not the Tokens)

In mid-July, a Forrester report picked up by The Register put in writing something several companies already suspected: software bills are set to grow next year, and AI will be the main reason why. Anthropic, OpenAI and GitHub have already moved to usage-based billing, Microsoft has launched a premium license, and the implicit message to enterprise customers is the same everywhere, the infrastructure that runs the models has a cost, and that cost has to land somewhere.

The report adds a line that reads almost like an aside, and it matters more than the rest: success on an AI investment depends on how well a company has built the right foundations, more than on how much it has spent on AI narrowly defined. Those foundations are, in a company that has never measured its own maturity, exactly what tends to be missing.

## Usage-based billing changes who controls the budget

As long as AI was bought on a license, cost was predictable: one number a year, negotiated once, revisited at renewal. Consumption billing works differently. It grows with usage, with context length, with the number of calls a process generates without anyone counting them one by one. An agent running around the clock, calling a model at every step of a workflow, repeating a search because the first attempt missed the answer, produces a spend curve the CFO only sees once the invoice arrives.

This is not a technical problem to leave to IT. It is a governance problem around consumption, one that needs thresholds, alerts and clear ownership of who authorizes what, the same way a company treats any other variable cost line. The optimization techniques exist and are well documented, from prompt caching to right-sizing the model for each task, to targeted retrieval that avoids reloading unnecessary context on every call. They remain optimization techniques, not decisions: they tell you how to spend less on a use case already chosen, not whether that use case made sense in the first place.

## The trade of people for tokens that isn't paying off

A second thread of news, from the same weeks, tells another side of the same problem. Several companies, including well-known names like Klarna, Uber and Walmart, have shifted budget from headcount to tokens, cutting roles to fund AI adoption. The early data coming out of those choices, though, doesn't show the expected return: customer satisfaction down in some cases, processes that still require human oversight, savings on paper that get consumed by rework and correction.

The problem isn't AI itself, it's the sequence. Cutting first and measuring later reverses the order that makes an investment defensible. Our AI Rating measures maturity across four dimensions, Readiness, Delivery, Risk, Confidence, precisely because a company can have a technically sound use case and still lose money, if it lacks the organizational capacity to run it at scale. A model that performs well in the lab and poorly in production is almost always telling you a skipped assessment, more than a technology limit.

## Where the cost actually originates

In our assessments, the question "how much does AI cost" almost always arrives too late, once the project has already been decided and the cost to justify is the one already incurred, not the one still avoidable. The question that actually saves money comes earlier: this use case, with this data, in this organization, where will it generate value and where will it just generate consumption.

A recurring pattern: a company automates a customer service process with a conversational agent, measures a drop in first-response time and presents it as a success. Six months later, operating costs have doubled because of repeated calls on ambiguous cases the agent can't close on its own, and the staff meant to be freed up gets reassigned to checking responses before they reach the customer. The cost didn't appear out of nowhere. It was predictable from month one, if someone had measured the escalation rate before scaling the project across every channel.

## What changes when you measure first

Companies that come to us already at AI Rating class B or above have these conversations differently. They know which processes have data clean enough to support an autonomous agent, which still need oversight, which risk gates block scaling regardless of how promising the use case looks. Cost, at that point, gets discussed by component, build cost, run cost, governance cost, and each one is estimated up front, not reconstructed later from the invoice.

The reverse holds too. A company in class D or C that invests at scale anyway risks paying twice, once for the project that doesn't hold up, and again to rebuild it properly once the failure mode is understood. It's the pattern we see most often in projects that end up with us after starting somewhere else: the budget wasn't the problem, the sequence was.

With 2027 budgets under discussion over the coming months, the useful question isn't how much to allocate to AI. It's whether the organization already knows, with evidence rather than impressions, where that budget will produce measurable value and where it will just produce a bigger bill in December.
