The number that matters more than the benchmark
OpenAI's newest model has cut token usage on agentic coding tasks by 54%. It sounds like a technical footnote, but for anyone paying to run AI tools, it may be the most important figure the company has released this year.
Here's the plain-English version. When an AI "agent" tackles a job — say, writing and fixing a piece of software — it doesn't answer once and stop. It works in loops: read the task, make a plan, attempt the work, check the results, spot the errors, try again. Each pass through that loop consumes "tokens", the units of text an AI processes. And every token costs money and adds a delay.
Cutting token use by more than half, then, isn't a minor tweak. It can be the difference between an AI tool that's too expensive to justify and one that's genuinely affordable to keep running day after day.
What efficiency looks like in practice
Consider a business running 1,000 AI-assisted tasks a day, each using roughly 40,000 tokens across its back-and-forth reasoning. Halve that and daily usage drops from 40 million tokens to around 18 million. At scale, that translates into thousands of pounds a month staying in the budget — and every task finishing quicker, because the model achieves more with less effort.
The improvements come from a few sensible changes under the bonnet. The model reasons more directly, reaching the right answer with fewer detours. It makes smarter use of tools, avoiding needless repeat actions. It manages context better, sending only what's changed rather than re-reading everything each time. And it strips out filler, so more of what it produces is actually useful.
This reflects a broader shift in how the AI industry is thinking. For a while, the race was simply about capability — could a model solve the problem at all? Now that most leading models can, the question has moved on: how cheaply and quickly can it solve the problem, especially when it has to repeat the process many times over? Efficiency has become a feature in its own right.
Why this matters for smaller firms
For small business owners, cost has been one of the biggest barriers to using AI beyond simple chatbot tasks. The more ambitious tools — ones that draft code, process documents or handle multi-step admin — have often been priced out of reach once you run them at any real volume.
Falling token costs change that maths. A tool that was uneconomical last quarter may suddenly pay for itself. It also means faster results, which matters if AI is sitting inside a customer-facing process where speed affects the experience.
If you've trialled an AI tool and found it too slow or too costly to roll out properly, efficiency gains like these are worth revisiting. The underlying economics are improving quickly.
What to watch next
Expect rivals to answer with efficiency claims of their own, and expect the tools built on top of these models — the ones most businesses actually buy — to quietly get cheaper or more capable at the same price. Keep an eye on your own usage bills over the coming months: if your provider passes these savings on, the case for adopting more ambitious AI tools gets stronger with each release.
