AI Policy Wiki
Dashboard

Frontier Models

medium confidence · updated 2026-06-17

The small set of most-capable, general-purpose AI models at or near the leading edge of capability — typically the largest, most compute-intensive foundation models from a handful of developers. The category that anchors capability-threshold regulation and frontier-specific safety practice.

Frontier models are the most capable general-purpose AI models at or near the leading edge of what the field can produce at a given time — in practice, the largest and most compute-intensive foundation models from a small number of developers. The term names a moving target: as capabilities advance, the frontier moves with them, so membership in the category is defined relative to the current state of the art rather than by a fixed capability level.

Origins of the term

The "frontier AI" framing was popularized by Anderljung et al., "Frontier AI Regulation: Managing Emerging Risks to Public Safety" (2023), which defined frontier AI models as "highly capable foundation models that could possess dangerous capabilities sufficient to pose severe risks to public safety" (Source: arxiv.org). The paper argued that this leading-edge subset raises governance problems distinct from those of AI systems generally: capabilities can be hard to anticipate before deployment, dangerous capabilities can be difficult to remove once a model is broadly available, and a model can proliferate rapidly. The framing was adopted in policy settings over the following year, including the United Kingdom's 2023 AI Safety Summit at Bletchley Park and the institutes and reporting regimes that followed.

Defining the frontier

Because the category is relative, regulators and developers operationalize it through proxies rather than a single capability test. The most common is a compute threshold — a training-compute line (for example, measured in floating-point operations) above which a model is presumed to be at the frontier and attracts additional obligations. Compute thresholds are administratively simple and observable in advance of training, but they are imperfect proxies for capability, and improvements in algorithmic efficiency mean a fixed FLOP line captures different capability levels over time. Other proxies include training cost, benchmark performance, and qualitative assessments of general capability. The instruments that key obligations to these thresholds are treated separately under frontier AI governance.

Frontier models are general-purpose: a single model serves many downstream tasks rather than being trained for one. This generality is the source of both their economic value and the governance concern, since the same system can be adapted to beneficial and harmful uses, a property discussed under dual-use frontier AI. Capability at the frontier is also uneven across tasks — strong in some domains and weak in adjacent ones — a pattern described as the jagged frontier.

Developers and access

As of 2026 the frontier is occupied by a small set of developers, including OpenAI, Anthropic, Google DeepMind, and a number of Chinese laboratories such as DeepSeek, with most leading models released through API access rather than open weights. A separate strand of frontier models is released with downloadable weights; the distinct risk and governance profile of that strand is treated under open-weight frontier models. The concentration of frontier development among a few well-resourced labs — a function of the compute and capital required to train at the leading edge — is itself a recurring subject of competition and governance debate.

Relationships