Stop Using Frontier Models to Wash the Dog!
Arguably the most fun, and the most value, we have had this month came from working out what to do when you cannot use the models you want. Either because you ran out of access or your…
The Token Cost Crisis Is a Maturity Problem
Arguably the most fun, and the most value, we have had this month came from working out what to do when you cannot use the models you want. Either because you ran out of access or your usage limits meant you no longer could do so affordably. Once you dig in think about it, we dont use drinking water to wash the dog/car. Right now in AI we are using our most expensive AI / Fijian Mineral Water to wash the dog.
Anthropic’s Mythos 5 blew us away with its deep thinking while frightening us with the price. A Codex splurge forced a serious conversation about spend caps. Then Gemma 4 showed us another version of the future, running locally on a machine that, by happy coincidence, also plays the latest James Bond game with eye-melting graphics.
Meanwhile, the frontier moved twice in a fortnight. Anthropic’s Fable returned after government intervention had abruptly taken it out of circulation. OpenAI launched Sol, Terra and Luna. By the time firms had finished evaluating one generation, another had arrived.
At the same time, I keep reading that companies are watching their token costs skyrocket and calling it a crisis.
It is all the same story. The token cost crisis is not really a pricing crisis. It is a AI maturity problem.
Start with the people
Every member of staff who can make productive use of frontier models should have access to them.
As a starting budget, circa £100 per knowledge worker per month will cover serious access for most people. Power users should be funded separately and judged on the return they produce.
Some people can spend thousands a month effectively. But they are a special case. Anyone spending that much should be creating enormous value, and if they are, the model bill is probably the cheapest thing on your P&L and the easiest thing to approve.
The budgeting model for people is not especially complicated.Where costs become difficult is automated work. Firms take every task, however routine, and send it through the smartest and most expensive model available. Then they are surprised when the bill arrives.There are four reasons this is a bad operating model.
The first is availability. In June, government action made two leading Anthropic models unavailable almost overnight. Fable returned globally less than three weeks later, while Mythos came back more selectively. If your operation depended entirely on them, you had a very bad few weeks.
The second is price. Frontier pricing changes, and reasoning tokens mean the advertised price tells you less than ever about what a piece of work will actually cost. Carry on using the latest frontier model for all tasks and watch your costs skyrocket.
The third is that leadership changes constantly. Anthropic had an extraordinary year. In my opinion, OpenAI has just taken the crown back. By November, somebody else may hold it. If your tooling is welded to one provider, you are always one release away from being welded to the wrong one.
The fourth is security: where your information goes, what you can interrogate and what you can prove to a client or regulator.
The answer is underneath the frontier
What the token-cost panic misses is that the latest open-weight models are now extremely capable.
They are not always equal to the best frontier models. On the hardest reasoning, the gap still matters. For most high-volume business work, it often does not.
There is now a growing group of very good open-weight models from Google, Meta, Nvidia and Chinese labs. Companies can download them, host them and use them commercially under their respective licences without paying a frontier provider every time the model performs an action.
The one I will name is Google’s Gemma 4, because we use it for a serious amount of our work.It handles documentation, error checking and research. It runs on a single consumer GPU and, outside the hardest reasoning work, it is excellent.At a corporate level, the compute required to run these models is now available to buy or rent at increasingly sensible prices.
Use them for the right workloads and the marginal token bill approaches zero.
That does not mean the work itself is free. You still pay for hardware or hosting, electricity, engineering, monitoring and security. But you stop paying a frontier-model premium on every routine action.
For scale, our entire AI bill for a twenty-person company is less than one junior salary, and we are producing a mountain of code with it.
That is possible because we do not send every task to the most expensive model. If you like Fijian mineral water, drink it. Do not use it to wash the dog.
What mature architecture looks like
Enterprises should lean into Anthropic, OpenAI, Google and Copilot. The frontier is worth paying for.
The mistake is not buying frontier models.The mistake is buying only frontier models, with no ability to route around them.Build your AI tooling to support multiple providers and open-weight models. We run an AI router at TAU for exactly this reason.The difficult reasoning goes to whichever frontier model is best suited to the task. High-volume, repeatable and well-specified work goes to open-weight models running at a fraction of the cost.
When the zeitgeist changes, we move the traffic, not the architecture. My own stack today is a good example.
Fable does much of my planning. ChatGPT is my main harness and advanced coding environment. Gemma 4 handles documentation, error checking and research. Hermes runs my automation.
It is a great stack. But what makes it great is not the specific models in it.
It is that I can replace almost all of it in minutes when the ground moves. And the ground now moves every few months.Staff should also be able to use models in two ways: through company APIs and tools, but also through the providers’ own interfaces.A lot of the value now lives in the harness rather than the raw model. The ChatGPT desktop app is a powerhouse. Many of our team live in Claude, Cowork and Claude Code.Locking people out of the best interfaces to simplify governance is a false economy.
Security and governance matter enormously. But they have to be designed around the way people actually work. Otherwise, staff simply route around the approved system using personal accounts that the security team cannot see.
The architecture is only half the job
Choosing the models is the relatively easy part.The difficult part is getting useful tools adopted by actual people.Too many companies trap their staff in bad tools, no tools or the wrong tools. Copilot is a good example. It is a decent general system and frequently wins the security review.
But for many people working in product or technology development, it is not the tool they need.
A system that wins on security while failing on usefulness has not really succeeded on security. Nobody uses it for the work that matters. They find another route, normally through an unapproved personal account.I see it every day.
The operating model for 2027 is not frontier or open-weight. It is both.Frontier models should handle the work where their reasoning earns its premium. Open-weight models should sit underneath them, eating the volume. Staff should be properly funded to use the best interfaces, while the company maintains control over security, routing and data.And one rule should never bend: do not outsource your entire AI operating layer to a black box you cannot interrogate.
With all this in mind you should be able to walk into your CFO’s office and answer three questions:
Which tasks genuinely needs the frontier costs? Which work tasks can run on open-weight models? How quickly can we switch strategy/model when the market moves?
The firms that win will not be the ones with exclusive access to the smartest AI. They will be the ones that know which work deserves it, and which work does not. Scaled across a wide organisation.