Prompt governance is a regulatory practice that treats system-level instructions (system prompts) — the natural-language text developers and deployers use to constrain model behaviour — as objects of governance: artefacts that can be disclosed, inspected, mandated, or prohibited in order to control how AI systems behave. The term is named and critically examined in Neumann, Sargeant & Singh's FAccT'26 paper.
Background
System prompts have drawn policy attention because, unlike model weights, training data, or inference dynamics, a system prompt is legible to non-technical stakeholders and cheap to edit. Policymakers have begun to treat it as an accessible intervention point. EO 14319 lets vendors evidence compliance "through disclosure of the LLM's system prompt," and the EU GPAI Code of Practice folds the system prompt into a disclosable "model specification."
Underlying assumptions
Neumann, Sargeant & Singh (Prompt Governance? On Governing Technologies Governed by Natural Language) characterize prompt governance as resting on two assumptions: that stakeholders can (i) infer intent from instruction text and (ii) predict behaviour from those instructions. Their systematic review of 287 papers finds that neither assumption reliably holds. The literature advances divergent, contradictory claims across an eight-goal typology — alignment, accessibility, adaptability, performance, stability, security, implementation, and auditability — with findings that prompts are brittle across phrasing, ordering, multi-turn interaction, and model versions.
Governance built on these assumptions, the authors argue, risks a "compliance illusion" or "false sense of control": a prompt that reads as aligned with a principle while failing to yield stable behaviour, letting actors satisfy disclosure requirements with carefully drafted text whose operational effect is unverified. They describe the failure mode as conflating linguistic compliance with behavioural compliance.
Structural issues
The paper identifies several structural complications beyond the prompt-as-proxy problem.
Instruction hierarchies span the AI supply chain — what the authors call the "prompt stack" — including provider System Instructions and System Guidelines, application Developer Instructions, and end-user instructions (compare OpenAI's Model Spec "chain of command"). Treating a single system prompt as the governance object misspecifies the layered system actually being governed.
Because prompts are easy to write, the authors argue, prompt-based control tends to expand to objectives it cannot reliably deliver, a dynamic of governance by affordance or function creep in which the ease of the lever sets the scope of governance.
Prompt mandates can also spill across borders and become de facto global standards (a Brussels Effect dynamic), as multinational vendors may apply one jurisdiction's prompt constraints everywhere.
Relation to policy
Prompt governance sits at the intersection of transparency policy, model-spec-style value specification, and procurement regulation. It is the mechanism behind a class of policy instruments such as "publish your system prompt" and "demonstrate neutrality via prompt disclosure." Neumann et al. reframe those instruments as a starting point for governance and evaluation rather than a substitute for behavioural evidence.
Relationships
- instance-of: AI Governance (umbrella)
- related: AI Transparency, Five Levels of Meaningful Transparency, System Card Due Diligence, Constitutional AI, Prompt Injection, Jailbreaking and Red Teaming, Regulating Under Uncertainty, Executive Order 14319 — Preventing Woke AI in the Federal Government, EU General-Purpose AI Code of Practice (2025)
- supported-by: Prompt Governance? On Governing Technologies Governed by Natural Language