AI Policy Wiki
Dashboard

Classification Institutions

medium confidence · updated 2026-06-06

The Hadfield-Weingast primitive — the institution within a normative social order that produces and maintains the shared classification scheme of punishable and permitted behaviors. The microfoundation beneath 'normative competence' and the 'democratic matrix'.

A classification institution is the structural element, in the Hadfield-Weingast theory of normative social orders, that produces and maintains a society's shared classification scheme — the publicly recognized account of which behaviors are treated as wrongful (and so punishable) and which are treated as permitted. It is the microfoundation the Knight Columbia Building AI for the Democratic Matrix essay (Hadfield, Trivedi & Hadfield-Menell, March 2026) places beneath its two headline concepts, Normative Competence and the The Democratic Matrix.

Origin

The term is drawn from Hadfield & Weingast (2012, 2014), a body of work by Gillian K. Hadfield and the political economist Barry Weingast that asks what law is. Their answer reframes law not as a body of commands backed by a sovereign, nor as a set of internalized moral norms, but as the output of an interactive system whose function is to solve a coordination problem.

In their model, a normative social order is an equilibrium state in which independent actors pursuing ordinary self-interest — not assumed to be innately pro-social — are incentivized and coordinated to take part in costly third-party punishment of wrongful behavior. The set of behaviors classified as wrongful is supplied by a classification institution, producing a reliable, stable pattern of conduct organized around the shared classification scheme.

The analytic move is the focus on punishment rather than compliance. If the punishment problem is solved — people are both motivated to punish and coordinated so that they punish the same things — compliance follows. The classification institution is what makes that coordination possible: it tells dispersed enforcers what counts as a violation, so their punishment efforts converge instead of fragmenting.

A classification institution does more than list rules. It must continuously resolve ambiguity, fill gaps, and adjudicate contested cases, generating a common reference point that decentralized actors can rely on. To perform that role and sustain the equilibrium, the institution's output must satisfy a set of "legal attributes" closely tracking Lon Fuller's rule-of-law criteria: stability, that the scheme does not change unpredictably; generality, that classifications apply to categories of conduct rather than single actors; clarity, that the scheme is legible to those expected to follow and enforce it; neutrality, that classifications are not bent to favor particular parties; and impersonal reasoning, that decisions are justified by reasons that hold regardless of who is involved.

In the Hadfield-Weingast framing these attributes are valuable instrumentally — they are what secure the stability of a normative order coordinated on a shared classification scheme — rather than because they realize the rule of law as a free-standing moral ideal. Courts, regulators, legislatures, and standards bodies are familiar examples of classification institutions, but so are far more informal arrangements: the framework is designed to capture how order is produced even where formal law is weak.

Relation to AI

The classification institution is the microfoundation the Democratic Matrix essay uses to support its larger argument. Normative Competence — the agent-level capacity to detect sanctions, attribute them to behaviors, and adjust conduct — only has content if there is something for a competent agent to read: a moving classification scheme produced by some institution. The The Democratic Matrix is, at the macro level, the co-adapting system of agents and institutions through which classification schemes are continuously revised.

Hadfield, Trivedi & Hadfield-Menell draw a research direction from this: building digital classification institutions, whose outputs are machine-readable, updated in real time, and legible to agentic AI systems operating at the speed and scale of automated decision-making. Their claim is that as AI agents become actors inside the democratic matrix, the classification institutions those agents must read can no longer be assumed to be slow, document-bound, and human-paced; the institutional infrastructure itself has to be re-engineered for an environment where many participants are machines.

The authors present this as a reframing of a familiar alignment question. On their account the problem is not only how to encode the right values into a model, but how to build and maintain the institutions that generate the shared classifications a normatively competent agent would track, and how to make those institutions interpretable to AI.

Relationships

Sources

Page created 2026-05-25 (gap type 4 — a foundational source and two concept pages, The Democratic Matrix and Normative Competence, named Classification Institutions as their microfoundation with no page behind it). Confidence medium — single foundational source, the underlying Hadfield-Weingast theory well established.