AI Policy Wiki
Dashboard

Kimi K3: Open Frontier Intelligence (Moonshot AI, July 2026)

high confidence · updated 2026-07-25

Moonshot AI's launch announcement for Kimi K3, a 2.8-trillion-parameter sparse MoE built on Kimi Delta Attention and Attention Residuals with native vision and a 1M-token context window. Self-described as the world's first open 3T-class model; the primary text for K3's full self-reported benchmark table, architecture, availability, and pricing pending the technical report.

"Kimi K3: Open Frontier Intelligence" is Moonshot AI's launch announcement for Kimi K3, published July 16, 2026. It is the primary text for the model: at launch no technical report had been published, so this post carries the full self-reported benchmark table, the architecture description, and the availability and pricing terms. Moonshot stated a standalone technical report would follow.

Moonshot describes K3 as "the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning," and is explicit about where it does not lead: "While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across our evaluation suite, consistently outperforming other tested models."

Model and scale

K3 is the first open model to reach 2.8 trillion parameters. Moonshot notes that "for nine of the past twelve months, Kimi models have set the upper bound of open-model sizes." Full weights were promised by July 27, 2026. At launch the model used max thinking effort by default, with low- and high-effort modes planned.

Architecture

Two architectural changes are presented as the basis of the model, both concerning how information flows:

  • Kimi Delta Attention (KDA) — described as an efficient foundation for scaling attention.
  • Attention Residuals (AttnRes) — selectively retrieves representations across depth rather than accumulating them uniformly.

MoE sparsity is scaled up to effectively activate 16 of 896 experts under a Stable LatentMoE framework. Moonshot reports these changes, together with refined training and data recipes, yield "an approximate 2.5x improvement in overall scaling efficiency compared to Kimi K2."

Further named components: Quantile Balancing, deriving expert allocation directly from router-score quantiles and eliminating heuristic updates and a sensitive balancing hyperparameter; Per-Head Muon, extending Muon by optimizing attention heads independently; and Sigmoid Tanh Unit (SiTU) and Gated MLA, improving activation control and attention selectivity respectively.

Training and serving infrastructure

K3 applies quantization-aware training from the SFT stage onward, using MXFP4 weights with MXFP8 activations "for broad hardware compatibility." Moonshot describes a fully balanced expert-parallel training method with static shapes and no host synchronization on the critical path, so that expert imbalance does not degrade throughput. It recommends deploying K3 on supernode configurations of 64 or more accelerators.

Because KDA poses new challenges for conventional prefix caching, Moonshot contributed an implementation to the vLLM community for release alongside the model, stating: "KDA with prefill cache allows us to serve Kimi K3 at a highly competitive token price despite its scale and long context."

Reported case studies

The announcement's capability claims rest on four coding demonstrations and several knowledge-work examples, all self-reported:

  • Kernel optimization. Models worked independently in identical sandboxes with up to 24 hours to profile, rewrite, and benchmark four GPU-kernel tasks — spanning AttnRes, KDA, and a 512-head-dimension MLA kernel — across an NVIDIA H200 and a GPGPU from an alternative vendor. Moonshot reports K3 competitive with Fable 5 (with fallback) and substantially ahead of Opus 4.8, GPT 5.6 Sol, and GPT 5.5, and states that in late development "an early version of K3 handled the majority of the team's kernel-optimization work."
  • GPU compiler development. K3 built MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline; on supported roofline benchmarks Moonshot reports it on par with or better than Triton and torch.compile, sustaining end-to-end nanoGPT training with stable convergence.
  • Chip design. In a single 48-hour autonomous run, K3 built, optimized, and verified a chip serving a nano model on its own architecture using open-source EDA tools on the Nangate 45nm library: 4 mm², timing closed at 100 MHz, over 8,700 tokens/s simulated decode throughput, 1.46M standard cells, 0.277 MB SRAM, and an INT4 MAC array with fused dequantization.
  • Coding for research. K3 reproduced the I–Love–Q universal relations in computational astrophysics in about two hours, against an estimated one to two weeks for an experienced researcher: reviewing and cross-validating 20+ papers, implementing the numerical pipeline, evaluating 300+ equations of state, identifying inconsistencies in published formulas, generating 3,000+ lines of Python, and producing an interactive dashboard.

Knowledge-work examples include a 42-year ASIC-industry research report built through 120+ rounds of recursive self-improvement drawing on 2.8k+ web searches and fetches and 1.1k+ terminal data pulls across 11k+ pages, 87 quarterly reports, and 99 original PDFs; a fusion-industry consulting report; and a GWTC-5 gravitational-wave analysis of 391 events using 20+ concurrent subagents. Moonshot also reports K3 edited the model's own teaser video from 56 source clips and produced a motion-graphics explainer of its own architecture.

Benchmarks

All K3 results are reported at reasoning effort "max", temperature 1.0, top-p 1.0, with each model evaluated under one of three agentic harnesses — KimiCode, Claude Code, or Codex. "With fallback" for Claude Fable 5 means requests Fable 5 refuses under its usage policy route automatically to Claude Opus 4.8.

BenchmarkKimi K3Fable 5 (fallback)GPT 5.6 SolOpus 4.8GPT 5.5GLM-5.2
DeepSWE67.570.073.059.067.046.2
Program Bench77.876.877.671.970.863.7
Terminal Bench 2.188.384.688.884.683.482.7
FrontierSWE81.286.671.366.764.967.3
SWE Marathon42.035.039.040.014.013.0
PostTrain Bench36.641.434.634.128.434.3
MLS Bench48.349.946.242.835.540.4
Kimi Code Bench 2.0 (internal)72.976.964.871.769.064.2
GDPval-AA v2 (Elo)166817601748160014941514
BrowseComp91.288.090.484.384.4
DeepSearchQA (F1)95.094.293.1
Toolathlon-Verified73.277.974.976.273.559.9
MCP Atlas84.284.783.683.682.882.6
Automation Bench30.829.129.727.222.712.9
Job Bench52.957.446.548.438.343.4
AA-Briefcase (Elo)154815831495135411581260
APEX-Agents37.643.339.939.438.535.6
Office QA Pro63.369.963.263.960.941.4
SpreadsheetBench 234.834.732.431.629.128.1
DECK-Bench (internal)73.573.074.766.968.268.6
GPQA-Diamond93.592.694.191.093.591.2
HLE-Full43.553.344.549.841.4
HLE-Full with tools56.063.058.057.952.2
MMMU-Pro81.681.283.078.981.2
MMMU-Pro with python83.486.584.682.783.2
CharXiv (RQ)84.888.984.680.584.1

Two K3 benchmarks are Moonshot's own internal suites (Kimi Code Bench 2.0 and DECK-Bench), and several results are marked in the original as evaluated under a differing harness. See AI Benchmarks and Evaluation for the general caution on self-reported leaderboards.

Availability and pricing

K3 launched on the Kimi app (iOS, Android, HarmonyOS) and kimi.com, the Kimi Work desktop app (v3.1.0+, Windows and Apple silicon), the Kimi Code terminal agent, and the Kimi API as kimi-k3. Kimi Enterprise is offered with enterprise data privacy and member management.

TierPrice ($/M tokens)
Cache-hit input$0.30
Cache-miss input$3.00
Output$15.00

Moonshot states the API is powered by Mooncake's disaggregated inference architecture and "achieves a cache hit rate above 90% in coding workloads." See Inference Economics and Token Pricing.

Provenance

Published on kimi.com/blog, Moonshot AI's canonical domain, dated July 16, 2026. Pulled and verified July 16, 2026; core specifications corroborated by Simon Willison, MarkTechPost, and Fortune the same day. The promised technical report had not been published at the time of pull, making this announcement the primary text available at launch.

Relationships