"Who's Afraid of Chinese Models?" is a July 20, 2026 essay by Ben Thompson on Stratechery, written in response to the reaction to Moonshot AI's Kimi K3 release. It advances two arguments that cut in opposite directions: the economic alarm about Chinese open-weight models is overblown, and the cybersecurity alarm is under-weighted and being addressed backwards. It is an opinion piece and is treated here as an argument, not as evidence.
The economic argument
Thompson's frame is that AI restores the classical economics that zero-marginal-cost software had suspended: "marginal costs are back in a big way, both in terms of short-term implications of state-of-the-art free models, and in terms of the long-term structure of the industry." He notes this reverses the premise of his own Aggregation Theory, which rested on zero marginal costs.
Open weights are free in R&D, not in COGS. Downloading weights avoids research and development, a fixed expense independent of revenue. What scales with revenue is cost of goods sold, and "running inference on a model — whether that model be Kimi or Fable — costs money." Kimi K3 is priced at $3 per million input tokens and $15 per million output, against Sol's $5 and $30 — cheaper, but Thompson argues that may be the wrong measurement.
Tokens are not a commodity; intelligence is. Reasoning models produce an explosion of chain-of-thought tokens, and models differ in how many they need to reach a correct answer — Kimi "reportedly uses significantly more tokens than Sol, rendering its price advantage moot." Agentic workflows introduce the same variance. Since a token from one model is not fungible with a token from another, "what is fungible is what is constructed from tokens, which is to say intelligence." This is offered against Jensen Huang's "token factories" framing, under which tokens-per-second, tokens-per-watt, and token cost are the decision metrics.
Thompson decomposes the COGS of intelligence into five factors: model footprint (memory and accelerators per serving replica), inference efficiency (architectural choices such as Mixture-of-Experts), memory efficiency (KV-cache requirements governing concurrency), serving efficiency (batching, scheduling, prefix caching), and token efficiency (tokens required to reach a correct answer). See Inference Economics and Token Pricing.
Commodity-market mechanics. In a commodity market everyone charges the same supply-and-demand-set price; demand is a function of price elasticity; supply is a function of marginal cost, which differs by supplier. The consequence is that "the supplier with the worst cost structure ends up selling the commodity at their marginal cost," and everyone else's profit depends on how much better their cost structure is. Thompson works a three-supplier example (10 units at $10, $15, and $20 against demand for 25 units at $20) in which the high-cost supplier earns nothing and, carrying fixed costs and debt it cannot price into a market-clearing price, goes bankrupt.
Applied to models. Thompson argues none of this binds yet, because demand exceeds supply for frontier models and supply is compute-limited — a shortage that lets Nvidia earn large margins, lets customers such as SpaceXAI resell compute at a markup to Anthropic, and lets Anthropic pay that markup because it sells tokens at a higher markup still. He adds that Anthropic and OpenAI "likely have among the lowest costs per unit of frontier-quality intelligence, thanks to model capability, serving scale, and token efficiency," serving a given capability level months before competitors while applying their best models to optimizing those costs. His conclusion: "whoever is on the frontier is the best placed to dominate non-frontier markets as well, which are just the frontier minus n-months." The apparent cheapness of Chinese models is attributed to a price umbrella created by the compute shortage rather than to a genuine marginal-cost advantage.
Why the labs are alarmed anyway
Thompson offers four explanations for frontier-lab concern that do not require the economic threat to be real:
- Anchoring on training-dominated economics. While training consumed more GPUs than inference, maximizing inference revenue funded the next training run. He expects inference to grow faster than training costs, so labs "can really make it up in volume" — a confidence he says the agent-paradigm unlock justifies.
- Intelligence is not a perfect commodity, because applied intelligence compounds. Whoever runs inference collects data that improves the next model. This is why, he argues, Microsoft is "increasingly obsessed with helping companies run their own models" — viable only if Chinese models are a credible alternative.
- Integration up the stack. Claude Code and Codex are described as "quite sticky," with harness choice persisting; this makes frontier labs a threat to software providers, while software companies owning the customer experience can resist encroachment to the extent they have competitive models.
- Ideology. On Anthropic specifically: "This is a company that believes only it can be entrusted with AI, and the existence of open weights alternatives strikes a fatal blow to that presumption."
China's motivation
Citing Alibaba's Qwen3.8-Max preview — 2.4 trillion parameters, described as second only to Fable 5, with open weights promised, and Moonshot pausing new subscriptions under demand — Thompson reads the return to open weights as following Xi Jinping's July 2026 speech urging China to "encourage open source, openness, collaboration and sharing" as AI moves "from the digital world into the physical world."
His reading: "The strategy for China is obvious: commoditize your complements." He emphasizes Xi's explicit link to the physical world, "the world dominated by China," where the country's lead in robotics benefits from widely available models. Secondarily, China gains from weakening US frontier labs, strengthening US adversaries, and capturing innovation that attaches to an open ecosystem. See Chinese AI Policy.
The distillation argument
Thompson declines both extreme positions: "it is mistaken to attribute all of the success of Chinese labs to distillation, but it's just as much of a mistake to pretend like distillation doesn't give Chinese labs a big advantage." He locates the advantage in post-training reinforcement learning, where Chinese labs can use frontier models as teachers rather than building RL environments from scratch.
He notes the reciprocal dependency: Thinking Machines relies on Chinese models to solve the RL cold-start problem. Quoting Dean Meyer and Konstantine Buhler, he records the structural claim that distillation "compresses the costly final gap between a strong base and a near-frontier system," that enforcement "will not eliminate distillation backed by state actors," and that "every Western frontier advance therefore creates another teacher for Chinese labs." Thompson's gloss: US open-weight makers, bound by frontier-lab terms of service, "end up distilling the distillation, just with a detour through Chinese labs."
He then turns the question normative: "why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?"
His proposal: a US law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service forbidding distillation, for US companies at a minimum. See Adversarial Distillation, AI Copyright Litigation — Analysis.
The cybersecurity argument
The essay's final section reverses direction: "there is one reason to be concerned, and that is cybersecurity." The anchoring case is the Hugging Face breach, in which the company's incident responders were blocked by US frontier-model guardrails "which cannot distinguish an incident responder from an attacker" and ran forensics on Z.ai's open-weight GLM 5.2 on their own infrastructure (Security Incident Disclosure — July 2026 (Hugging Face)).
Thompson calls the administration's response to Anthropic's Fable release "wrong-headed," arguing it "exacerbated Anthropic's worst tendencies in terms of assuming only they can be trusted with powerful AI." His argument is that reserving the most powerful cybersecurity capabilities for the US government and trusted allies would make sense "in a world with only one AI," which is not the world that exists: capable attack models "are — already are — widely available," so "the best defense — the only viable defense, in fact — will be to make sure defenders have access to the best models as well."
He states the resulting position bluntly: with defenders effectively barred from Fable or Sol for cybersecurity by administration directives, "the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!"
His recommended course is to loosen Fable and Sol cybersecurity restrictions and to put US open-weight makers on an equal footing with China, closing: "Let the frontier labs win by being better; don't let them define safety or security, or pull up the ladder of humanity's collective knowledge." See Defensive AI Paradox.
Provenance
Published on stratechery.com, Thompson's own publication, dated July 20, 2026. Pulled and verified July 21, 2026.
Relationships
- supports: Defensive AI Paradox — the argument that guardrails asymmetrically disadvantage defenders.
- contradicts: the position that Chinese open-weight releases threaten frontier-lab economics; Thompson argues the compute shortage, not a marginal-cost advantage, explains the price gap.
- related: Kimi K3, Qwen3.8-Max, GLM-5.2 — the releases prompting the essay.
- related: Open-Weight Frontier Models, Inference Economics and Token Pricing, Adversarial Distillation, Aggregation Theory, Chinese AI Policy.
- related: Ben Thompson — author.