A July 20, 2026 assessment by Zvi Mowshowitz of Kimi K3, published the same day as Nathan Lambert's essay on the same release.
The assessment
Mowshowitz grants the headline claim and then bounds it. "Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model."
The bound: "Do not get carried away. Do not judge Kimi K3 only on its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now."
That last clause is the load-bearing qualification, and it cuts against the reassurance a shrinking gap would otherwise offer: a narrowing measured in months understates the narrowing measured in capability, because more capability now fits into each month.
Four reasons to discount the benchmarks
- Distillation. "It is somewhat distilled."
- Benchmark-to-practice gap. "It likely outperforms on benchmarks relative to practical performance."
- Effort settings. "All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol" — a like-for-like comparison problem, not a capability claim.
- Jaggedness. "Performance looks jagged. Kimi will be excellent at some things, less so at other things."
He also notes the evidential state at time of writing: "we will know more over the coming weeks. For now access is spotty."
On which benchmarks matter
Mowshowitz argues the open-weight comparison flatters open models where it is most often made. "Narrow and relatively easy coding tasks are where open weights models are at their relative strongest, and the benchmark here is approaching saturation."
He turns instead to UK AISI's longer-horizon cyber-range results, quoting them: "On TLO, GLM-5.2 reaches as far as Opus 4.5, a model released less than 7 months before it, while DeepSeek's V4-Pro falls below Sonnet 4.5 (a sub-cyber-frontier model released 7 months before it). These results are broadly consistent across our other cyber ranges. Notably, GLM-5.2 reached step 7 with marginally fewer tokens than any other model on average, tracking Opus 4.6's trajectory to step 11 before stalling."
His reason for weighting these: "longer tasks are more relevant in terms of both of the most important things to worry about: automation of AI R&D and cyber attacks." He adds a gap in the record — "one might also worry about bio risks, even if that is not as in fashion, and it is a little concerning we don't see standard testing" for them on this release.
Relation to the parallel assessment
Mowshowitz's four-to-six-month estimate is wider than Lambert's 3–5 months for the same model, and the two differ in method rather than only in number: Lambert reads aggregate leaderboard position, while Mowshowitz discounts for effort settings, distillation, and the benchmark-to-practice gap, and weights long-horizon cyber ranges over saturating coding benchmarks. Both agree the gap is narrowing.
Relationships
- related: Kimi K3: The open-weights escalation (Nathan Lambert, July 2026) — same-day assessment of the same release, narrower gap estimate, different weighting of evidence
- supports: Kimi K3, Open-Weight Frontier Models
- related: UK AI Safety Institute (AI Security Institute), GLM-5.2, Adversarial Distillation, Zvi Mowshowitz