AI Policy Wiki
Dashboard

Prompt Governance? On Governing Technologies Governed by Natural Language

high confidence · updated 2026-06-06

Neumann, Sargeant & Singh (FAccT'26) — a PRISMA systematic review of 287 papers plus two policy case studies (US EO 14319 + Dec 2025 OMB memo; EU GPAI Code of Practice) arguing that governance regimes treating system prompts as stable behavioural controls rest on assumptions not supported by the technical evidence, producing a 'compliance illusion' / 'false sense of control.'

A peer-reviewed paper by Anna Neumann (Research Centre Trust, UA Ruhr, University of Duisburg-Essen), Holli Sargeant (University of Cambridge), and Jatinder Singh (Research Centre Trust / University of Cambridge), presented at the ACM Conference on Fairness, Accountability, and Transparency (FAccT'26), held June 25–28 2026 in Montréal. The 44-page paper (DOI: 10.1145/3805689.3806763) asks whether system prompts function reliably enough to support the AI-governance frameworks that increasingly target them, and argues that frameworks treating system-prompt disclosure as a stable behavioural control rest on assumptions the authors say are "not consistently supported by technical evidence."

Summary

The paper examines system-level instructions (system prompts), the natural-language text developers and deployers use to constrain model behaviour, and asks whether they function reliably enough to support governance frameworks built around them. It combines a PRISMA systematic literature review (923 records screened down to 287 papers analysed) with a scoping review of policy documents and two case studies. The authors report that the research literature is fragmented and advances divergent and contradictory claims about what system prompts can achieve, while emerging governance instruments treat them as stable, interpretable control mechanisms, an assumption the authors say is not consistently supported by the technical evidence.

Typology of system-instruction goals

Thematic analysis of the extracted claims yields an eight-category typology, split into system goals (what the instruction targets in the model) and prompt goals (the instruction as an artefact). For each category the corpus contains both supportive ("Further") and critical ("Hinder") claims; the table below records the recurring critical finding the authors associate with each.

TypeGoalRecurring critical finding
SystemAlignmentstating values in language does not guarantee behaviour; alignment-faking resists correction
SystemAccessibilityreadable prompts ≠ predictable behaviour; users "operating blind"
SystemAdaptabilitypersona prompts "brittle"; effects "might largely be random"
SystemPerformancefine-tuning can outperform prompting; some studies find no measurable effect
SystemStabilityconstraints violated despite explicit directives; unstable over multi-turn / reordering
SystemSecuritysystem prompts are both a defence and an attack vector; high extraction-success rates
PromptImplementationautomated prompts drift from intended meaning; human authoring lacks systematic improvement paths
PromptAuditabilityeffects fail to transfer across models/domains; frequent iteration defeats fixed audit regimes

Policy case studies

In US federal procurement, EO 14319 "Preventing Woke AI in the Federal Government" (July 2025) lets vendors demonstrate compliance with its "Unbiased AI Principles" "through disclosure of the LLM's system prompt." The December 2025 OMB implementation memo treats system prompts as an optional transparency artefact but excludes them from the "model evaluations" section, which the authors read as embedding the implicit assumption that inspecting prompt language is sufficient.

In the EU, the GPAI Code of Practice folds the system prompt into a required "model specification" disclosed to evaluators, but does not require prompt versioning, change-logs, or re-evaluation triggers.

The authors argue that both regimes treat prompt language as a proxy for model behaviour, and that neither operationalises the conditions under which a disclosed prompt is a reliable lever.

Key arguments

The paper's central thesis is what the authors call a "compliance illusion" or "false sense of control": a prompt may read as aligned with a governance principle while failing to produce stable behaviour across contexts, models, or multi-turn interactions. Conflating linguistic compliance with behavioural compliance, the authors write, "risks regulating what is legible rather than what is operational."

The authors describe a "prompt stack," in which instruction hierarchies span the AI supply chain (System Instructions > System Guidelines > Developer Instructions > User Instructions, comparable to OpenAI's Model Spec "chain of command"). On this view, treating a single system prompt as the governance object "risks misspecifying the system actually being governed."

They also identify what they call function creep and "governance by affordance": because prompts are cheap and easy to edit, prompt-based control tends to expand to objectives it cannot reliably deliver, and the ease of writing a prompt should not determine the scope of governance.

On cross-border spillover, the paper argues that prompt requirements can become de facto global standards through a Brussels Effect dynamic, with multinational vendors applying one jurisdiction's constraints everywhere.

The authors conclude that "Writing rules that govern machines requires different approaches than writing rules that govern humans," and recommend pairing textual requirements with behavioural evaluation, prompt versioning, and specialised intermediary roles.

Relationships