AI Policy Wiki
Dashboard

Foundation Models

medium confidence · updated 2026-07-28

Models trained on broad data at scale and adaptable to many downstream tasks — the term introduced by Stanford's CRFM in 2021, and the technical category underneath the regulatory terms general-purpose AI model, dual-use foundation model, and frontier model.

Foundation models are models trained on broad data at scale and adaptable to a wide range of downstream tasks. The term names the general technical category — a single model, trained once, then adapted many times — beneath the narrower regulatory categories that have since been layered onto it: the general-purpose AI model of the EU AI Act, the dual-use foundation model of Executive Order 14110, and the frontier model of frontier-safety practice.

Origin of the term

The term was introduced by Stanford's Center for Research on Foundation Models in "On the Opportunities and Risks of Foundation Models," a multi-author report released in August 2021 with Rishi Bommasani as lead author (arXiv:2108.07258) (Source: arxiv.org). The report states: "We call these models foundation models to underscore their critically central yet incomplete character" (Source: arxiv.org). The choice of word was deliberate on both counts — foundation to convey that these models are load-bearing for what is built on them, and the accompanying description of them as "intermediary assets; they are unfinished" to convey that a foundation model is not itself an application (Source: crfm.stanford.edu).

The report frames the category through two properties it calls emergence and homogenization:

"Emergence means that the behavior of a system is implicitly induced rather than explicitly constructed; it is both the source of scientific excitement and anxiety about unanticipated consequences. Homogenization indicates the consolidation of methodologies for building machine learning systems across a wide range of applications; it provides strong leverage towards many tasks but also creates single points of failure." (Source: crfm.stanford.edu)

The report's own summary of the pairing is that foundation models rest on standard deep learning and transfer learning, but that "their scale results in new emergent capabilities, and their effectiveness across so many tasks incentivizes homogenization," with the caution that "the defects of the foundation model are inherited by all the adapted models downstream" (Source: arxiv.org). The report draws the risk conclusion directly from the interaction of the two: "Since emergence generates substantial uncertainty over the capabilities and flaws of foundation models, aggressive homogenization through these models is risky business" (Source: crfm.stanford.edu).

Three further risk characterizations in the report have carried into later policy vocabulary: that foundation models are "a high-leverage single point of failure, making them a prime target for attack," with demonstrated adversarial and training-data-memorization vulnerabilities; that they are "byproducts of computationally expensive training regimes" whose energy demands carry environmental cost; and that their generality implies economic impact "in a wide variety of industries and occupations" (Source: crfm.stanford.edu).

The term was not adopted without dispute. Commentary published by CRFM alongside the report included Jacob Steinhardt's response on emergent behavior, and the coinage drew objections from researchers who considered the existing vocabulary adequate or the new label an implicit endorsement of the paradigm (Source: crfm.stanford.edu).

Relation to the neighbouring terms

The category has four commonly used names, which are not interchangeable. Each was defined for a different purpose, and only one of them — the original — is a purely technical description.

TermDefined byScope
Foundation modelCRFM, 2021Technical and descriptive: trained on broad data at scale, adaptable to many downstream tasks. No capability floor, no compute threshold.
General-purpose AI model[[eu-ai-act\EU AI Act]], Art. 3(63)Regulatory: trained with a large amount of data using self-supervision at scale, displaying significant generality, capable of integration into a variety of downstream systems. A separate systemic-risk tier attaches above 10²⁵ FLOP of cumulative training compute.
Dual-use foundation model[[eo-14110\EO 14110]] (rescinded)Regulatory and risk-scoped: trained on broad data, self-supervised, at least tens of billions of parameters, and capable or easily modifiable to enable CBRN design, offensive cyber operations, or evasion of human control.
Frontier modelAnderljung et al., 2023; frontier-safety practiceRelative: the most capable general-purpose models at or near the leading edge at a given time. A moving target rather than a fixed threshold.

The relationships between them are containment relationships rather than synonymy. Frontier models are the leading-edge subset of foundation models, defined relative to the current state of the art; the category's boundaries move as capability advances, which is treated in detail under frontier models. The EU's general-purpose AI model is a legal category drawn to be operable by a regulator, which is why it substitutes observable training characteristics — broad data, self-supervision at scale, significant generality — for the adaptability criterion that anchors the technical definition; it is treated under general-purpose AI. EO 14110's dual-use foundation model narrowed the category further by risk profile rather than by capability rank, restricting it to models whose capabilities bear on CBRN, offensive cyber, or loss of control, and pairing it with a 10²⁶ FLOP reporting threshold; the order was rescinded in January 2025, but the definition circulated into subsequent US drafting, including California SB 1047's covered-model framing.

Where a source uses one of these terms, the term carries the definition of its own regime. A model can be a foundation model without being a frontier model, and a frontier model without meeting a given compute threshold in a given year, since algorithmic efficiency improvements mean a fixed FLOP line captures different capability levels over time — an issue treated under compute thresholds.

The adaptation layer

The property that distinguishes a foundation model from a task-specific model is that the training and the use are separated: the model is trained once on broad data and then adapted, by fine-tuning, prompting, retrieval augmentation, or integration into a larger system, to tasks that were not specified at training time. This is the structural fact behind most of the governance difficulty. The developer that trains the model does not know the deployment contexts; the deployer that adapts it did not observe the training data; and, on the CRFM report's homogenization argument, a defect present in the base model propagates to every downstream adaptation simultaneously rather than being contained to one application (Source: crfm.stanford.edu).

Regulatory regimes have responded to this by allocating obligations across the layers rather than to a single actor — the EU AI Act's separation of GPAI-model obligations from high-risk-system obligations is the clearest instance — and by keying the heaviest duties to the base-model layer, on the reasoning that upstream intervention is the only point at which a defect can be addressed once rather than many times.

Open questions

  • Whether the technical category retains analytical work now that the regulatory categories are operative, or whether "foundation model" survives mainly as the informal superset of terms that carry legal consequences.
  • Whether adaptability — the criterion in the original definition — can be operationalized well enough to serve as a regulatory test, or whether proxies such as training compute and self-supervision at scale will continue to stand in for it.
  • How the homogenization argument applies where a small number of base models are adapted across critical sectors simultaneously, and whether concentration at the base-model layer is best treated as a safety question, a competition question, or both.

Relationships

Sources