AI Policy Wiki
Dashboard

Ways to think about token pricing (Benedict Evans, July 2026)

medium confidence · updated 2026-07-25

Essay arguing that the 2026 token supply crunch is transitory and that every currently visible market dynamic points toward frontier models becoming low-margin commodity infrastructure, with value captured further up the stack. Poses four questions that would have to resolve otherwise, and works through fiber, mobile data, and semiconductor-manufacturing comparisons while arguing analogies have no predictive value.

"Ways to think about token pricing" is a July 9, 2026 essay by Benedict Evans published on ben-evans.com. It asks where token prices settle once the current supply crunch eases, and argues that the visible market dynamics point toward frontier models becoming low-margin commodity infrastructure rather than holding sustainable pricing power. It is an opinion piece and is treated here as an argument, not as evidence.

The starting position

Evans opens with the two things he says can be stated with certainty: the market is in a supply crunch, and that state is unstable. He argues the situation is transitory on both sides.

Supply. More than a trillion dollars of data-center capex is in the pipeline, with further semiconductor capex behind it; inference efficiency continues to improve quickly; and new models differ substantially in token efficiency in both directions.

Demand. Although the market has been capacity-constrained since 2022, he attributes the first-half-2026 crunch to sudden product-market fit in a single use case — software development — which he calls "actually a pretty small field." He offers a counterfactual: "imagine if we had product-market fit for a consumer use case with hundreds of millions of DAUs — today's infrastructure couldn't support it at any price." Which use cases scale next, when, and with what token requirements are all unknown.

On unit economics he cites widely reported inference gross margins of 40–50%, noting these include depreciation of associated server costs or rental equivalents but rest on an unknown asset life ("five years? Seven years?") and exclude the cost of training the next model, which he says is currently far larger than revenue. In principle inference is a marginal cost and training a fixed cost, so sufficient revenue reaches profitability — but how training costs will move is unknown, as is how much of the recent usage surge has an ROI that can be quantified to a CFO.

Evans dismisses bottom-up modeling of the equilibrium as achievable but not useful, comparing it to building a five-year broadband forecast in 1998: "the spreadsheet will be very pretty, and you might even get close to the right number for this year," but there are too many unknowns for a longer-term market-structure forecast. His compact statement of the problem: token price is a function of supply and demand at a level between the sellers' marginal cost and the buyers' ROI, "but we don't actually know what supply, demand, marginal cost or ROI will be."

Four questions

In place of a model, Evans poses four questions about the intelligence-versus-cost curve, each a matter of degree rather than a binary and each likely to vary by use case.

  1. How many buyers will pay to sit at the frontier? Some use cases already work with a small, old, possibly open-source model running free on-premises or on a phone; some get better results from the newest and most expensive frontier model; many fall in between. The question is how far up the cost curve results actually improve, for how many use cases that improvement carries an ROI, and how much demand is absorbed by cheaper models that are "good 'enough' and much more commoditised." He notes the optimistic case — that ROI rises with more expensive models because results are better — and asks where that in fact applies.
  2. Does the frontier keep moving significantly? How long capability keeps improving, how long that keeps requiring more compute, and whether it does so fast enough to stay ahead of downward pricing pressure from efficiency and capacity gains — that is, whether "the expensive head of the curve continue[s] to be a thing."
  3. Will competition among frontier models stay fierce? Evans lays out the alternatives: the field shrinks with network effects emerging; models diverge so that different ones lead clearly in different fields; or the current pattern persists, with a mid-single-digit number of companies producing generally equivalent frontier models. He notes that at present "everyone is using mostly the same science and mostly the same training data, and getting mostly the same results," with no known network effect or other winner-takes-all mechanism that would let one company pull ahead sustainably.
  4. How much of the value from high-end use cases does the model capture? How much must be wrapped in tooling, process, proprietary data, go-to-market, networks, and support — the apparatus of a conventional software company — even where an expensive frontier model is required underneath. At the limit, whether a model can invent and build all of that itself and so charge by seat or by outcome, or whether even the most sophisticated use cases must sit inside "hundreds of new companies that can pick and choose which models to use."

He frames the resulting range with two poles: at one extreme "two or three giant minds that run half of everything and have massive pricing power"; at the other, LLMs resembling databases — "there'll be millions of them, some very big and some very small, and the value is in what you build on top - after all, every SaaS company is a 'database wrapper'." Concretely, a future in which Anthropic "(or a company we haven't heard of yet) wins the whole thing and can set its own terms," against one in which "dozens of routers run real-time auctions to allocate your tasks across hundreds of low-margin model-farms and a benchmark company takes a fee on every single one."

Uncertainty and the limits of analogy

Evans states repeatedly that the answer is not yet knowable, and attributes this to the S-curve stage "where it's clear that this is going to be huge but nothing else is clear at all" — the mid-1990s for the internet, 2008 or 2009 for mobile. He notes one position he does hold: that chatbots "are a poor interface that will struggle to capture value up the stack." See How will OpenAI compete? (Benedict Evans, February 2026).

He argues the present uncertainty differs in kind from earlier technology transitions because "we don't have a good theoretical understanding of why these models work so well and so we don't know how much better they can get." In 1995 the evolution of the internet was unknown but the physical limits were not — fewer than 100 million expensive PCs existed, and telcos could not deliver fiber to every home the following year; in 2010 the next iPhone was unknown but it would not have retinal projection. With language models, "next month a new approach could cut inference compute needs by 90%, or double demand, or both."

Three comparisons, and a warning against them. Evans observes that "all conversations about AI end in a hunt for metaphors," and works through the common ones.

  • Fiber. The dot-com fiber overbuild superficially resembles the current infrastructure build-out, but he raises two objections: fiber construction ran far ahead of demand where AI compute runs far behind it; and, more importantly, fiber was mostly fixed cost (digging holes) rather than marginal cost, whereas growth in compute demand requires buying more compute.
  • Mobile data. He treats this as the more fruitful comparison. Mobile networks carry marginal cost for capacity; they saw a usage surge roughly fifteen years ago that overwhelmed capacity and forced carriers to add capacity and rebalance pricing; and selling bits resembles selling tokens as "an opaque measure of marginal cost that doesn't map in any transparent or intuitive way to use cases or value, and needs to be replaced with bundles of some kind." The instructive part is the outcome: cellular data traffic rose by several orders of magnitude over twenty years into an industry with roughly a trillion dollars of annual revenue and $200 billion of capex, "but the stocks have gone nowhere, and all the value was captured by other people further up the stack."
  • Semiconductor manufacturing. This carries the escalating cost and complexity that mobile lacks. Evans cites Rock's Law — the cost of a cutting-edge fab doubling every four years — under which the number of frontier players fell from dozens to a handful and "now really only one, TSMC," raising the question of whether AI becomes so hard and expensive that only a couple of firms can do it even absent network effects. He notes the same price/performance curve applies to semiconductors, many of whose uses sit further back along it. But he adds that even TSMC's de facto frontier monopoly and good margins do not capture a large share of the broader technology economy: net income of $53 billion in the prior year, "less than half of Apple alone."

He notes that Sam Altman has in six months compared OpenAI both to Windows — "a high-margin capital-light monopoly based on network effects" — and to electricity utilities, "natural monopolies but also low-margin regulated utilities selling a pure commodity," and that cloud offers a third comparison with three leading players, good margins, distinguished propositions, and again limited value capture.

Evans then rejects the method he has just demonstrated: "analogies don't have predictive value. You can't prove whether something will have the same outcome as mobile by arguing how much it's like mobile." He cites as the cautionary case the argument fifteen years ago that Android would beat iOS because Android was "open" and "open" Wintel had beaten the "closed" Mac in the 1990s, and adds that it "is also the mistake that doomers make when they claim that AI is 'like' nuclear weapons."

What the comparisons do establish, on his account, is empirical rather than predictive: "something can be very important, very expensive, change the world, and be full of very sophisticated science and engineering, and yet have a wide range of possible outcomes." Price equilibrium is possible at high margins and at low margins, with and without market concentration, and this "can't [be] hand-wave[d] away by talking about AGI and saying 'you don't understand exponentials!'"

The asymmetry

Evans's closing argument is that the uncertainty is not symmetric: "every path to foundation models having market dominance, strategic leverage, value capture, winner-takes-all effects, or anything else other than becoming commodity infrastructure, requires something to change." He lists what would have to change and notes the evidence against each: frontier competition could ease, "yet in the last six months, Mark Zuckerberg and Elon Musk jumped from zero back onto the leaderboards"; network effects could emerge; chatbots could grow into products that need no software wrapper; one lab could pull ahead on execution, as Microsoft, Google, Facebook, and Apple each did before acquiring winner-takes-all effects.

He identifies two possible external interventions, which he calls "not one but two potential dei ex machina — Trump and China": China is reportedly considering regulating open source and some people close to Trump have floated the same, though he notes that "since Meta abandoned Llama the US has no leading open model"; and export controls could expand and become systematic. On the labs' own advocacy he records the two readings without adjudicating: "Many people see the pleas for regulation from Anthropic and (sometimes) OpenAI as a front for regulatory capture, but either way, we can't presume this will remain an entirely free market." See Open-Weight Frontier Models, Chinese AI Policy.

His conclusion returns to the asymmetry: as the supply crunch eases, current dynamics point toward frontier models as commodity infrastructure with the value built on top, "and for a different outcome, something needs to happen that we don't see yet."

Relation to other sources

The essay shares its premise with How will OpenAI compete? (Benedict Evans, February 2026) — no known mechanic gives one frontier lab a durable lead — and extends it from competitive position to price equilibrium. Both invoke Rock's Law and the TSMC case for the same purpose: to argue that frontier-scale capital intensity can consolidate an industry without conferring value capture up the stack.

It reaches the opposite conclusion from Who's Afraid of Chinese Models? (Ben Thompson, Stratechery, July 2026), which accepts that AI is a commodity market but argues the frontier labs likely hold the best cost structure and that "whoever is on the frontier is the best placed to dominate non-frontier markets as well." Evans's first and fourth questions are the ones on which the two disagree: Thompson holds that intelligence rather than tokens is the fungible unit and that token efficiency favors the frontier labs, while Evans treats the share of demand absorbed by cheaper "good enough" models as open. Both accept the framing that AI restores marginal costs to a software industry that had lost them.

Relationships