Six weeks ago I published a video arguing that the price you pay for AI is fake. Subsidised. That the companies selling it are losing billions doing it, and that one day the bill could jump ten times.
I still think the prices are fake. But I got the ending wrong.
Watch the video:
Because in February, something happened that I think most business owners missed: Chinese AI models overtook the American ones. Not on a benchmark. In real, paid usage. And last week China launched a model that, on some scoreboards, plays at the same level as the best American ones - and it's open. You can download it.
I'm not watching this from the sidelines. I pay these companies' invoices every month, my dev team builds on their tools every day, and I've put £1.5M of my own money into building a product with AI running through it. So this isn't commentary for me. It's my cost base.
The scoreboard no marketing department controls
OpenRouter is a routing service for AI models. You connect to it once, and it sends your requests to whichever model you choose. Thousands of companies run their AI traffic through it, which makes it a rare thing in this industry: a scoreboard no marketing department controls.
A year ago, American models carried about 70 per cent of that traffic. In February, the Chinese models crossed over. Today the American share is about 30 per cent, and DeepSeek, Kimi and Qwen carry around 60.
And that chart only counts traffic that goes through a marketplace. When a company downloads an open model and runs it on its own servers, nobody counts those tokens at all. No bill, no chart. On All-In they've given that a name: dark tokens. The real shift is bigger than anyone can measure.
The lead has gone the same way. Three years ago the American labs reckoned they were two years ahead of China. After DeepSeek, people said six to nine months. Last week Moonshot launched Kimi K3, China's biggest model yet, and on some scoreboards it's level with the best American models. I went to sign up this week and couldn't: there's a waitlist, because they've sold every scrap of computing power they've got.
Then there's the price. Claude Opus 4.8, the model my dev team rates highest for the hardest work, charges 25 dollars per million tokens of output. xAI's flagship Grok charges two dollars fifty. DeepSeek charges 28 cents - roughly ninety times cheaper than Opus.
Are the cheap ones as good? At the very hardest work, in my experience, no. But most of what a business actually does with AI isn't the hardest work. It's summarising calls, reading documents, drafting, sorting data. Call it 95 per cent of the real tasks, and that 95 per cent can now be done by lots of different models. The premium price isn't defending the whole job any more. It's defending the last five per cent.
One honest caveat, because it's your money: some models burn twice the tokens to finish the same job. The number that counts isn't the price per token, it's the cost per task. But even on cost per task, the gap is enormous.
Free is a strategy
Why would anyone give away a model that cost hundreds of millions to build?
Because free is a strategy - the oldest one in software. Google gave Android away, and now most of the world's phones run it and Google owns the standard. I've played this game from the other side: the platform my own product started on is open-source telecoms software. Free to download, while the company behind it sells hosting and support. When you give the thing away, you're not selling the thing. You're setting the standard.
China wants to set the standard for AI, and the state is backing the labs that give their models away. And because most of these models are open weights, you download the actual model and run it on your own kit. Nobody bills you per token. Which means nobody can ever put the price up on you.
That's the door slamming shut on the old plan. Undercut everyone, outlast everyone, then raise prices - it doesn't work when your competitor is a free download. Not now, not in five years. Never.
And the American industry has been spending like the old ending was guaranteed. OpenAI alone has committed something like 1.4 trillion dollars to computing power, on revenue of about 25 billion a year. America is outspending China on this by seven times or more, for a lead that's now a photo finish.
Watch what they do, not what they say
The people running these companies can do that maths faster than you and I can. So watch what they've started doing.
In June, SpaceX went public in the biggest listing in history. OpenAI and Anthropic both filed to follow in the same window. OpenAI has since pushed theirs back, reportedly to 2027. Anthropic isn't waiting: they're pushing for October, at around a trillion dollars.
Anthropic's position fascinates me, because they mostly rent their computing power - and one of their landlords is xAI, a direct competitor. It came out in the SpaceX listing paperwork: a billion and a quarter dollars a month for a data centre in Memphis, committed through to 2029. That's fifteen billion a year going to a rival, while they build fifty billion dollars of their own data centres on top.
I know what that kind of contract feels like, because we sign them in telecoms. When we order a leased line for a customer, we commit to our provider for the full term. If that customer goes bust - and that's happened to us more than once - we still pay for the line, every month, to the end of the contract. That's what Anthropic has signed, except my lines don't cost a billion and a quarter a month.
Then the second tell: Sam Altman has been in Washington floating a fund where the AI companies hand over equity so the American public can share in the ownership. And a third: some of the American labs are asking Washington to restrict the Chinese models. There are real security questions with open Chinese models, and I'm not dismissing them. But companies that are winning don't usually ask the referee to ban the other team. And you can't un-release an open model anyway.
One more thing about those IPOs, and it's the part I'd want to know if I weren't following any of this. Once these companies are public and big enough, they get added to the index funds. If you hold one, your pension will own a slice of this experiment whether you chose it or not.
Let me be fair, though, because there's a serious argument on the other side. David Sacks, the White House AI adviser, thinks the panic is overdone: one good benchmark isn't the frontier, the labs have unreleased models ahead of anything public, and the revenue backs him up - Anthropic's run rate has reportedly gone from about ten billion dollars in January to around fifty billion by early summer. The technology is real. I've bet my own money on that.
But markets don't price this quarter. They price the next ten years. Can a company priced at a trillion dollars grow into that number when the ceiling on its prices is a free download? Honestly, I can't get to yes. And if the answer's no, somewhere between here and there a lot of air comes out.
The risk has moved
Here's where I've changed my mind since June. Back then, the risk I could see was the bill jumping ten times. Zoom out, and the price of intelligence is only heading one way: down. The cost of running the same level of AI has been falling about ten times a year. Work that cost 20 dollars in 2022 costs about 40 cents today. Chips improve, models get more efficient, and now there's a price war on top.
But the risk didn't disappear. It moved. It's not the bill any more. It's being stuck.
If your product is hard-wired into one AI company, their problems are your problems. If they have to raise prices to survive, that's your margin gone. If they wobble, that's your product down. And there's a third one now that wasn't on my list in June: if their government decides who gets access, that's your business in someone else's hands.
That's not a theory. In June, the US government ordered Anthropic to suspend access to its newest models for every non-US person, over a security worry. Switched off, for everyone outside America, for two and a half weeks. I run a British company building on American AI, so that one's personal.
A closed model isn't yours. It's a licence, and a licence can be switched off by a boardroom or a government. An open model on your own servers can't be.
The internet already ran this experiment. Netscape built the first big web browser and charged for it. Then Microsoft gave its own browser away for free, and Netscape - the company everyone thought would own the internet - was finished as a business within a few years. But look who actually won the internet: not the browser companies. The value went to everything built on top. The shops, the search engines, the millions of ordinary businesses. Same crash, two completely different outcomes. The difference was where you were standing when the price of the layer below you went to zero.
If you build things, a crash in the price of intelligence lands on you as a discount. Not as a bill.
What we're actually doing about it
Three things, inside my companies.
One: we run our own AI where the work is steady. Inside Olatti, our communications product, the transcription and the voice commands run on open models - Meta's Llama and Google's Gemma - on our own servers, in our own UK data centres. It's been built that way from day one. Those jobs don't need a frontier model; they need to be good, fast and cheap. Nobody can reprice that, nobody can switch it off, and our customers' calls never get shipped to a third-party AI company.
I'll be honest about the limit. Our own kit runs the smaller models. The really heavy work can't sensibly be self-hosted by a normal company - we've looked. Renting GPUs only pays if you're hammering them round the clock; if your work comes in bursts, you're paying for silence. So for the heavy stuff we rent by the token like everyone else, which is exactly why this price war pays businesses like ours. And it cuts the other way too: Evalua, our call-scoring product, pays ElevenLabs for premium transcription, because there the quality is the whole product. This isn't ideology. It's value for money, all-in - if switching costs six months of engineering, that's money too.
Two: we never marry one model. For the hardest coding work we use Claude - Opus, and now Fable. I've used OpenAI's 5.6 as well, and honestly it's right up there. The harness we're building, Coet, is designed to route every request to whichever model is best value for that job: frontier model for the hard tasks, a cheap open one for the routine ones. When the maths changes, we change the routing and nobody rebuilds anything. The moment Kimi or Grok will actually take my money, they go into the mix too.
Three: some things, we deliberately don't build yet. I used to think of this as build versus buy. There's a third option now: wait. We've had a voice AI agent, Ringup, designed on paper for a while. I could have built it already, on somebody else's model, at somebody else's prices. I haven't. A thin wrapper around someone else's AI inherits someone else's problems - their price rises, their outages, their rules - and in a shake-out, the wrappers go first. The plan's ready. When the economics settle, we build.
And there's one more reason I want the open models to stay in this race. If two or three companies end up owning all the intelligence, eventually they build every product too. Maybe yours. Real competition at the model layer is what lets the rest of us keep our margins.
The bet underneath the build-out
There's a bigger assumption under all of this spending that's worth naming: that the AI of the future lives in a giant data centre, miles away from you, rented by the token.
I'm not sure it does. Apple is betting the other way - do the everyday work on the device, and only send the hard problems to the cloud. That's the same shape we landed on with Olatti, in miniature: local first, escalate when needed. As the chips in your phone and laptop get stronger and the open models get more efficient, more of the everyday work moves on-device or into a box in your own office. If that's where it goes, the demand for rented, data-centre intelligence grows slower than the trillion-dollar commitments assume - and the overbuild is bigger than it already looks.
We've been here before. The telecoms build-out of the late nineties left the world full of dark fibre, cable laid years ahead of the demand, and the people who made the real money bought it out of the wreckage for pennies and built the internet we use today on top of it. I work in that industry. I run on those leftovers.
When does the air come out this time? I don't know, and nobody honestly does. That's exactly why I'd rather fix my setup than guess the date. Because when the last bubble burst, the winners bought the leftovers for pennies. This time, the leftovers are a free download.
If you run a business on AI and you're working out the same questions - what to own, what to rent, what to wait on - I write a free weekly newsletter about what's actually working inside my companies. You can join it here: axelmolist.com/youtube.
And if your setup is different - if you've gone all-in on one provider, or all-in on local - I'd genuinely like to hear how it's going. Hit reply. I read every email.
Thanks, Axel.
Join 480+ founders and operators getting my learnings from building and running my companies, straight to your inbox.