
When AI Spend Outruns Value, Look at the Billing Model First
Most AI budgets are committed before anyone knows which projects will deliver meaningful returns. That makes one early decision particularly important: how a company pays for the models it uses.
The AI Journal recently reported on a GFT survey of roughly 1,000 CTOs and CIOs at large organisations. Nearly nine in ten respondents said they were concerned that AI spending was moving faster than measurable value. More than eight in ten had abandoned at least one AI initiative because it could not be integrated effectively with legacy systems.
The usual response to findings like these is to focus on choosing better use cases. That is certainly part of the problem. But there is another decision that determines how expensive a failed experiment becomes: the billing model.
Too often, companies make that decision before the experiment has produced any evidence at all.
The commitment often comes before the evidence
Enterprise AI projects frequently begin with questions about infrastructure.
How many GPUs will the workload require? Should the company reserve capacity for six or twelve months to secure a better rate? How much headroom will be needed during periods of peak demand?
Those are sensible questions for a mature workload. For an early-stage pilot, however, the answers are largely estimates.
Teams end up buying capacity based on an assumed level of future demand and only afterwards discover how the application actually behaves in production.
If the project succeeds, the estimate may prove reasonable. If it fails, the financial commitment remains.
That creates a second problem: sunk costs can influence product decisions. When expensive reserved capacity is sitting unused, there is an incentive to keep a marginal project running longer than its results justify. Instead of spend following value, the organisation starts looking for ways to justify spending that has already happened.
What changes with per-token billing
Usage-based inference reverses that sequence.
Instead of reserving infrastructure upfront, a company accesses a hosted model through an API and pays for what it consumes. Pricing is generally based on the number of input and output tokens processed, with the two often charged at different rates.
That changes the economics of experimentation in three important ways.
First, failed experiments have a natural financial endpoint. When the application stops making requests, its inference cost stops as well. There is no remaining infrastructure commitment to recover.
Second, spending scales with adoption. If a feature attracts little usage, its inference bill remains small. If usage grows, spending grows with it.
Third, costs become easier to attribute. Token consumption can be tracked by application, department, customer or individual feature. Instead of a broad infrastructure expense, finance and product teams can see which workloads are generating the bill.
Usage-based pricing does not automatically make AI inexpensive. Its advantage during experimentation is that it makes the relationship between usage and cost considerably easier to see.
Unit economics become measurable
Once inference is metered, teams can calculate something much more useful than a monthly infrastructure bill: cost per task.
Consider a hypothetical customer-support assistant. Suppose each request requires around 3,000 input tokens for the customer question, instructions and relevant context, followed by approximately 500 output tokens.
At an illustrative price of €1 per million input tokens and €4 per million output tokens, the inference cost would be €0.003 for the input and €0.002 for the output — roughly half a euro cent per ticket.
Those prices are examples only. Actual rates vary significantly between models and providers, and real applications may consume substantially different numbers of tokens.
What matters is not the specific figure but the fact that the unit cost can be calculated directly.
A support leader can compare the inference cost per resolved ticket with the operational value of automating that ticket. The same exercise with reserved infrastructure requires assumptions about utilisation, peak demand, idle capacity and how hardware costs should be allocated across multiple workloads.
Metered inference also makes optimisation more concrete. Prompt length, context size, caching and model selection become financial variables. Shortening an unnecessarily large system prompt, caching repeated context or routing simpler tasks to a smaller model can produce a measurable change in the bill.
The infrastructure market has changed
There was previously another argument for committing to dedicated infrastructure early: accessing usage-based inference often meant relying on a small number of major US model providers.
For organisations with strict requirements around data location, infrastructure control or regulatory compliance, that could limit the usefulness of the model.
The market is becoming broader.
Capable open-weight model families, including DeepSeek, GLM and Kimi, can increasingly be deployed and served by infrastructure providers outside the organisations that originally developed them.
Companies can now purchase pay-as-you-go AI inference on modern GPU infrastructure in Europe without necessarily committing to dedicated capacity from the outset. Depending on the provider, this can include usage reporting, low or no minimum consumption commitments and clearer control over where inference requests are processed.
For organisations evaluating both economics and data residency, this creates a middle ground between relying entirely on major model APIs and renting dedicated GPU infrastructure before demand is known.
When usage-based pricing stops making sense
Per-token pricing is not always the cheapest architecture forever. Its strongest advantage is during periods when demand remains uncertain.
Once the workload becomes predictable, the calculation can change.
A high-volume application running at consistently high utilisation may achieve a lower effective cost per token on dedicated infrastructure. Proprietary or heavily customised models may also require dedicated deployments. Applications with strict latency guarantees, security requirements or hardware-isolation policies can create similar constraints.
The important distinction is that these requirements should ideally be demonstrated rather than assumed.
High utilisation, predictable volume and specialised infrastructure needs are things a company can measure after operating a workload. They do not necessarily need to be assumptions embedded in the project’s economics before the first users arrive.
A better sequence for AI budgets
For organisations planning AI investment, the safer approach is to treat infrastructure commitment as something earned by a workload rather than something required to begin one.
A practical sequence looks like this:
- Launch new use cases using metered inference and attribute token consumption to individual applications or features from the beginning.
- Define a measurable value metric before launch, such as tickets resolved, documents processed, transactions completed or employee hours saved.
- Compare cost per task with value per task regularly, and redesign or discontinue applications whose economics do not work.
- Consider dedicated capacity once demand has become sufficiently stable to model utilisation with confidence and the numbers demonstrate a meaningful saving.
This approach does not solve every problem highlighted by the GFT survey. Legacy integration, organisational skills, governance and change management remain significant barriers regardless of how inference is purchased.
What it does change is the cost of being wrong.
If an experiment fails under a usage-based model, the organisation loses what it spent running the experiment. If it fails after substantial infrastructure has already been reserved, the financial consequences can continue long after the project itself has stopped delivering value.
As enterprise AI moves from experimentation into production, infrastructure decisions will increasingly become financial decisions as much as technical ones.
The companies that keep AI spending closest to measurable value may therefore be those that delay their biggest capacity commitments until the workload — rather than the forecast — tells them what they actually need.


