For years Formula One engines were made faster the obvious way: more fuel, more displacement, more of everything. In 2014 the sport changed what it measured. Each car got 100 kilograms of fuel for a whole race, a third less than before, and a second rule capped how fast it was allowed to burn them: 100 kilograms an hour, measured by a sensor fitted to every car [1].
Everyone assumed the racing would get slower. Within a few seasons the cars were quicker than the ones that had been allowed to drink as much as they liked. The rules had simply moved the place where engineers went looking for performance.
Enterprise AI has been through one correction this year and is walking into another. Both rest on the same assumption: that when the bill grows faster than the results, the answer is to spend less. That assumption is wrong, and 2014 is the reason. What moves the needle is deciding in advance where each unit of spend should go.
From somewhere reasonable, at first. Nobody knew what these systems could really do, and the quickest way to find out was to use them heavily. Usage also happened to be the only thing anyone could measure, so it became the stand-in for progress.
Then it stopped being a stand-in and became the target. Teams were ranked on internal leaderboards by tokens consumed. At Amazon, engineers reportedly pointed agents at work nobody had asked for, purely to keep their numbers up. The leaderboards came down eventually[2]. The whole episode picked up a name, tokenmaxxing, and then an ending, for a reason that looks obvious in hindsight: consumption tells you what something cost and never what it was worth[3].
The correction was to measure outcomes instead of activity. Tasks completed, rework avoided, hours handed back to people. That is right, and everything after it depends on it.
Where it stops is at the individual request. Measuring value gives you a portfolio view: it tells you which initiatives deserve next year's budget, and says nothing about how any single job should run. Which model handles it. How many steps it gets. What it is allowed to cost. That gap is where the variability lives. Give the same task to the same agent twice and the second run can cost thirty times what the first one did[4].
Which changes the question you are actually asking. Not how much are we spending on AI, but how is each request being handled.
The default in most companies today is to send everything to the best model available. While you are still learning what the technology can do, that is a sensible way to buy information. Once you know, it is an expensive habit.
Two findings make the case for changing it. Models differ enormously in what they consume on identical work: across the same set of tasks, some burn well over a million more tokens than others. And past a certain point, a job that consumes more stops producing a better answer. Performance climbs with spend and then goes flat[4]. Spend beyond that line and you are buying the same result a second time.
So the next phase is about assignment rather than restraint. Routine classification, extraction and summarising go to a smaller model. Work that genuinely needs reasoning goes to the frontier one. And the choice belongs to the class of task rather than to whoever happens to be at the keyboard. In Europe there is a little more riding on getting it right, since keeping inference inside a defined geography costs roughly 10% more than the option that makes no promise about where the work happens[5].
Three decisions, all taken before the money is spent.
There is a fourth item, and it has to sit alongside the first. Someone needs to be checking that the smaller model still answers as well as the larger one. Assigning models by class of work without measuring quality is a cost cut dressed up as an architecture.
None of this asks Finance to learn anything new. It is what Finance does everywhere else in the business: agree what a unit of work is worth before authorising it. The only unfamiliar part is the size of the unit. Here it is a single request rather than a purchase order.
It is a fair worry. Guardrails placed on an experimental technology can end the experiment.
So look at what happens without them. Nearly half the executives KPMG surveyed this year have already scaled back agent deployments because the running costs outweighed the benefits [6]. That is where an unexplained bill leads: to less AI rather than more, switched off by people who could not defend the invoice. A budget that refuses keeps the work running. An invoice nobody can explain gets the programme cancelled.
It is worth being exact about the friction, too. A policy applied in the path of the request costs milliseconds. What teams actually dread is the approval queue, and an approval queue is what you end up with when the only control anyone has is a human being asked to say yes.
The 2014 season settles that argument. The teams that won under the new rules were not the ones who found a way around them. They were the ones who adapted first.
All three at the same moment. Which model may be used, what the spend is allowed to reach, and whose work this call belongs to share one answer window: when the request is made, before a model has been picked, while the workflow still has a name attached to it. Ask any of them from the finance system and you are already too late, because by the time the invoice lands all three have been settled by default.
The tooling market has a name for that point, the AI gateway, and its useful property is that measurement and enforcement sit in the same place. It can record what every call cost and whose work it was, and it can turn down the one that breaches a limit.
Three questions, then, before the next budget cycle.
In 2014 the fuel limit did not slow the cars down. It changed where the engineering went.
[1] FIA, 2014 Formula One Technical Regulations, Article 5.1.4 (fuel mass flow not to exceed 100 kg/h; 100 kg maximum race fuel allowance; FIA-homologated fuel flow sensor)
[2] Fortune, "Tokenmaxxing is over. It was a flawed way to measure a company's ROI from AI", 28 May 2026
[3] IBM, "Tokenmaxxing is dead, long live valuemaxxing", 25 June 2026
[4] Bai, Huang, Wang, Sun, Mihalcea, Brynjolfsson, Pentland, Pei, "How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks", arXiv:2604.22750, 24 April 2026
[5] Microsoft Learn, Azure AI Foundry deployment types documentation; pricing differential verified against the Azure Retail Prices API, 11 August 2026
[6] KPMG, Global AI Pulse Q2 2026 (2,145 respondents, 20 countries) — 49% of leaders have scaled back AI agent deployments because operating costs outweighed the benefits