AI Policy Wiki
Dashboard

Prompt Governance

medium confidence · updated 2026-06-06

Governing AI behaviour by regulating the natural-language system prompts that condition it — treating system-level instructions as accessible, inspectable governance levers. Neumann, Sargeant & Singh (FAccT'26) argue this rests on assumptions of stability and control not supported by technical evidence, producing a 'compliance illusion.'

Prompt governance is a regulatory practice that treats system-level instructions (system prompts) — the natural-language text developers and deployers use to constrain model behaviour — as objects of governance: artefacts that can be disclosed, inspected, mandated, or prohibited in order to control how AI systems behave. The term is named and critically examined in Neumann, Sargeant & Singh's FAccT'26 paper.

Background

System prompts have drawn policy attention because, unlike model weights, training data, or inference dynamics, a system prompt is legible to non-technical stakeholders and cheap to edit. Policymakers have begun to treat it as an accessible intervention point. EO 14319 lets vendors evidence compliance "through disclosure of the LLM's system prompt," and the EU GPAI Code of Practice folds the system prompt into a disclosable "model specification."

Underlying assumptions

Neumann, Sargeant & Singh (Prompt Governance? On Governing Technologies Governed by Natural Language) characterize prompt governance as resting on two assumptions: that stakeholders can (i) infer intent from instruction text and (ii) predict behaviour from those instructions. Their systematic review of 287 papers finds that neither assumption reliably holds. The literature advances divergent, contradictory claims across an eight-goal typology — alignment, accessibility, adaptability, performance, stability, security, implementation, and auditability — with findings that prompts are brittle across phrasing, ordering, multi-turn interaction, and model versions.

Governance built on these assumptions, the authors argue, risks a "compliance illusion" or "false sense of control": a prompt that reads as aligned with a principle while failing to yield stable behaviour, letting actors satisfy disclosure requirements with carefully drafted text whose operational effect is unverified. They describe the failure mode as conflating linguistic compliance with behavioural compliance.

Structural issues

The paper identifies several structural complications beyond the prompt-as-proxy problem.

Instruction hierarchies span the AI supply chain — what the authors call the "prompt stack" — including provider System Instructions and System Guidelines, application Developer Instructions, and end-user instructions (compare OpenAI's Model Spec "chain of command"). Treating a single system prompt as the governance object misspecifies the layered system actually being governed.

Because prompts are easy to write, the authors argue, prompt-based control tends to expand to objectives it cannot reliably deliver, a dynamic of governance by affordance or function creep in which the ease of the lever sets the scope of governance.

Prompt mandates can also spill across borders and become de facto global standards (a Brussels Effect dynamic), as multinational vendors may apply one jurisdiction's prompt constraints everywhere.

Relation to policy

Prompt governance sits at the intersection of transparency policy, model-spec-style value specification, and procurement regulation. It is the mechanism behind a class of policy instruments such as "publish your system prompt" and "demonstrate neutrality via prompt disclosure." Neumann et al. reframe those instruments as a starting point for governance and evaluation rather than a substitute for behavioural evidence.

Relationships