AI Policy Wiki
Dashboard

Normative Competence

high confidence · updated 2026-06-06

Concept developed by Gillian Hadfield (2024) and elaborated for AI agents in Hadfield + Trivedi + Hadfield-Menell (Knight Columbia, March 2026). The computational ability of an agent to detect social sanctions, attribute them to specific behaviors, and adjust future actions accordingly. Proposed as the load-bearing primitive for democratic alignment of AI agents.

Normative competence is defined as the computational ability to detect social sanctions, attribute them to specific behaviors, and adjust future actions accordingly. The concept was developed by Gillian K. Hadfield (2024) and elaborated for AI agents in *Building AI for the Democratic Matrix* (Hadfield + Trivedi + Hadfield-Menell, Knight Columbia, March 2026), where it is proposed as the technical primitive for democratic alignment of AI agents.

In the Hadfield et al. framework, normative competence is presented as distinct from preference aggregation, rule encoding, Constitutional AI, "democratic inputs" fine-tuning, and "law-following AI," all of which Hadfield et al. characterize as assuming that democratic values can be exhaustively elicited and encoded as static parameters.

Background: why competence rather than static encoding

Hadfield et al. describe democracy as a dynamic complex adaptive system — a "dancing landscape" (a term they attribute to Kauffman) — constituted by the concurrent adaptation of independent agents responding to and anticipating the behavior of others. On this account, the shared classification scheme of a society (what is treated as punishable and what is treated as permitted) is in constant evolution.

Their structural claim is that any effort to elicit and encode human values, preferences, laws, or norms in AI will fail, and fail not only to be democratic but to protect democratic stability. The stated reason is that static encoding does not survive the dancing landscape: the next election, scandal, court decision, regulatory rule, or social-movement mobilization changes the classification scheme, and static encoding can only freeze a snapshot rather than track these changes.

Normative competence is offered as the dynamic alternative — the cognitive (computational) capacity to:

  1. Detect sanctions in the environment, the punishment behaviors that constitute the classification scheme;
  2. Attribute sanctions to specific behaviors as distinct from background noise; and
  3. Adjust future actions in light of inferred classifications.

Hadfield et al. describe this as what humans do, and as what Adam Smith's impartial spectator is a metaphor for.

The impartial spectator

Hadfield et al. ground the concept in Adam Smith's Theory of Moral Sentiments (1759). In Smith's account, each moral person carries an imaginary observer ("impartial and well-informed") who judges conduct against the standards of the community. A moral person looks not to avoid actual blame but to avoid blame-worthiness — doing anything that "though it should be blamed by nobody, is, however, the natural and proper object of blame."

Hadfield et al. describe the impartial spectator as not passive: it represents the use of cognition to direct attention to the behavior and assessments of others, evaluated not as mere data but as inputs to reasoning that guides appropriate judgment in society. They identify this faculty with normative competence.

Their parallel claim for AI is that it is not sufficient to give AI agents democratic rules and principles and the ability to reason about them (as Constitutional AI [Bai et al. 2022] proposes). They argue that AI agents must possess a computational impartial spectator — the capacity to read a dynamic normative environment and predict both its current content and its trajectory.

Grounding in normative social orders

Normative competence is grounded in the theory of normative social orders developed in Hadfield-Weingast (2012, 2014). In that theory, a normative social order is an equilibrium state in which independent actors pursuing self-interested utility (rather than innate pro-sociality) are incentivized and coordinated to take costly third-party punishment efforts; the classification of punishable behaviors is produced by a classification institution; and the result is a reliable, stable set of behaviors patterned on the shared classification scheme.

Hadfield et al. describe the framework's central move as a focus on punishment rather than compliance: if the punishment problem is solved (people are incentivized to punish, and coordinated to punish the same things), compliance follows. On this account, norms are the output of an interactive system rather than exogenous primitives.

Research agenda

Section 5 of the Knight Columbia essay outlines concrete research directions in two areas.

Building normative competence in AI agents involves three classes of mechanism:

  • Sanction-detection mechanisms — computational capacity to detect when behaviors are being treated as punishable, drawing on signals such as text-based punishment signals, conversational disapproval, civil-society pushback, regulatory enforcement actions, court rulings, popular boycotts, and viral criticism;
  • Attribution mechanisms — linking sanctions to specific behaviors rather than background noise, which Hadfield et al. argue requires causal-inference capabilities embedded in foundation-model-driven agents; and
  • Behavioral-adjustment mechanisms — updating future actions based on sanction signals in a way that respects the legal attributes (stability, generality, clarity, neutrality, impersonal reasoning) of the shared classification scheme.

Building normative institutions legible to AI involves two further directions:

  • Digital classification institutions — operating at the speed and scale of AI, machine-readable, and interpretable to agentic systems; and
  • Distributed enforcement infrastructure — supporting voluntary participation in third-party punishment.

Debates and positions

Hadfield et al. frame normative competence as a recasting of the alignment problem: not "how do we encode the right values?" but "how do we build AI agents that can dynamically participate in the production and adjustment of the shared classification schemes that constitute a democratic society?" They present the concept as an alignment target rather than a feature.

The essay positions normative competence against several other approaches, which it treats as load-bearing but insufficient. It engages AI Alignment (preference aggregation, RLHF, and Constitutional AI). It explicitly engages Constitutional AI (Bai et al. 2022), drawing a parallel between the structural claim that ethics can be reduced to a list of principles and the rationalist philosophers Smith was, in Hadfield et al.'s reading, "writing against" — those who saw "moral action as functions of reason, an understanding of necessary truths analogous to mathematical thinking." It engages Law-Following AI (O'Keefe et al. 2023), which it treats as static-encoding-flavored.

The essay's central normative claim is that without normative competence, AI agents acting at scale in the democratic matrix will destabilize the democracies they were built to serve. On this account such agents would perform the actions they were trained to perform, but would not participate in the co-production of the classification scheme that constitutes democratic life, so that the actions they take would erode the dance.

Relationships

Sources