Open App →
Back to Blog
August 23, 2026 By Steve

Will AI go the way of Mobile Pricing?

...the industry flipped to unlimited for one price a month...

Minutes or Electricity

I used to think I knew how this ended. Most people do. We spent two decades watching one metered technology after another go flat, and mobile minutes were the cleanest example. For years you paid by the minute, rationed your calls, and dreaded the overage line on the bill. Then the industry flipped to unlimited for one price a month, and the meter went away for good. AI pricing, the reasoning goes, is walking the same road. Tokens are today's minutes. Normal people don't know what a token is and never will, and eventually one of the big consumer apps will do the obvious thing and sell unlimited AI for a flat monthly fee.

I assumed that too. Then I put it against the actual cost curves, and I changed my mind about where it lands. What follows is a working thesis, not a verdict. I expect it to move as the technology and the market meet each other. But right now the evidence points somewhere more interesting than a clean flat rate.

Why mobile went unlimited

Start with why the phone analogy worked, because the reasons are the whole argument.

Two things had to be true for unlimited minutes to make sense. First, the cost of one more minute had to fall to almost nothing. It did. The expensive part of a mobile network is building it. Once the towers are up, an extra minute of traffic costs the carrier next to nothing, so giving minutes away is nearly free. Second, and this is the part nobody names, usage was self-capped. A phone call needs a human on the line for its entire length, and there are only so many hours in a day. Even the heaviest talker on the plan was bounded by the clock. The carrier could look at the worst-case user and know the ceiling.

Unlimited works when the marginal unit is cheap and the heaviest user is bounded. Mobile had both. The whole question for AI is whether it has either.

The curve that says yes

Here is the case for the optimists, and it is a strong one.

The cost of intelligence is falling faster than almost anything in the history of computing. Measured at a fixed quality level, the price of running a model has dropped by roughly a factor of ten every year for several years running. A task that cost sixty dollars per million tokens in 2021 costs a few cents today at the same capability. Some analysts put the decline even steeper, closer to a halving every couple of months. That pace beats the cost of computing during the PC era and bandwidth during the dotcom build-out.

If that were the only curve that mattered, the argument would be over. Ride a ten-times-a-year cost decline for a few more years and even a heavy user becomes cheap to serve. Sell them unlimited and let the curve bail you out.

But there is a second curve, and it runs the other way.

The curve that says no

The deflation number carries a quiet condition: it holds at a fixed quality level. Nobody buys a fixed quality level. People buy the best thing available, and the best thing keeps changing.

Two behaviors break the optimistic math. The first is reasoning. Newer models don't just answer, they think on paper first, generating long internal chains before they respond. That hidden thinking can burn many times more tokens than the visible answer, on the order of a hundred times more for a genuinely hard problem. The second is agents. A simple chatbot query is one call to the model. An agent takes a goal and fans it out into a dozen or more calls, reasoning, invoking tools, checking its own work, and trying again. Analysts peg an agentic task at somewhere between five and thirty times the token consumption of a single chat. And it compounds quietly. As an agent's running context fills with history and tool output, the cost per step climbs week over week without anyone touching the code.

So consumption per useful task is rising about as fast as price per token is falling. The two curves are fighting, and at the frontier the consumption curve is winning.

The frontier never got the discount

Here is the sleight of hand that fools the eye. The ten-times-a-year deflation is real, but it describes yesterday's capability getting cheaper. It says nothing about the price of the newest, best model, which is exactly where demand sits.

The clearest tell is at the top of the market. When the first mainstream reasoning model shipped, its price per output token was the same sixty dollars per million that the original GPT-3 charged at its launch years earlier. The frontier price didn't fall. It reset to the old high. Flagship models today still run several dollars to fifteen dollars per million output tokens, and the recent direction is up, not down. At least one major lab raised its standard frontier rate by half in the autumn of 2026, and that increase lands on top of the agent multiplier, not instead of it.

So the headline everyone repeats, that intelligence gets ten times cheaper every year, manages to be true and misleading at once. It is true at a capability tier nobody chooses. At the tier people actually use, the cost of getting real work done is flat or climbing. That is the flattered number of this entire debate. Anchor on it and you conclude unlimited is inevitable. Look at what a frontier task actually costs and you conclude the reverse.

The flat rate already runs underwater

You don't have to model any of this in the abstract, because the providers have already run the experiment on themselves.

The most-quoted admission came from OpenAI, whose chief executive said publicly that the company was losing money on its most expensive subscription, not because the price was too low in theory but because subscribers used it far more than expected. Independent analysts trying to reverse-engineer the economics landed on a startling estimate: on the standard flat plans, the provider's margin goes deeply negative once a single user consumes only a single-digit percentage of the nominal allowance. The plan is profitable only because the overwhelming majority of subscribers barely use it. The light users pay for the heavy ones, and the heavy ones are capable of costing many times their fee.

That estimate comes from one analysis and deserves the caution any single estimate does, but the direction is corroborated by the labs' own admissions and by what shows up in enterprise invoices. Organizations that rolled out agentic coding tools have watched a small group of power users run bills ten times the size of the median, and blow through annual budgets in a single quarter. This is precisely the arithmetic mobile never had to fear, because a phone user could not talk for a thousand hours in a month. An AI user with agents can.

Not a flat line, a barbell

Put the two curves and the economics together and a shape falls out, and it is not the mobile endgame. It is a barbell.

At the floor, providers are going exactly where the phone carriers went. Basic chat is becoming a commodity, handed out unmetered, increasingly carrying ads, used to acquire and lock in the largest possible base. For the median user, the flat-rate future has already arrived.

At the ceiling, they are going the other way. Caps, tiers, and pay-as-you-go overage credits are appearing on the heaviest products, because that is where the unbounded cost lives. One major lab has already realigned a flagship agent product from a flat allowance toward metered usage, and has said plainly that unlimited plans probably will not survive, comparing an unlimited AI plan to an unlimited electricity plan.

That comparison is the punchline, and it is the correction to the phone analogy. AI inference is not minutes. It is electricity. Nobody sells unlimited electricity, because consumption is not self-limiting and every unit drawn has a real cost. In the AI case that is nearly literal, since the dominant cost underneath the token is the power to compute it. The meter, in other words, never disappears. It relocates. It moves off the token bill the consumer was never going to read and reappears at the top of the range, where the agents run overnight.

What would change my mind

I want to name the seam in this argument, because a thesis you can't break isn't worth much.

The deflation curve is ferocious, and so far it has flattened every prediction that the cost of intelligence would plateau. My whole case rests on the consumption curve staying ahead of it at the frontier. If algorithmic efficiency on reasoning and agentic work starts falling as fast as it already fell on plain chat, the ceiling commoditizes too, the economics of unbounded usage stop being frightening, and the flat rate comes back within reach. That is a real possibility. It isn't visible yet, but it is the one development that would flip this, and it is the thing I will be watching.

So take this as a snapshot of a moving target. Today the evidence says the mobile ending only half-arrives: free at the bottom for the many, metered at the top for the few. The useful question isn't who wins the price war. It's what you are actually holding when you look at your plan. Minutes, which run out and then go free, or electricity, which you pay for by the unit no matter what the sticker says.