The User Sovereignty Model and the Public Safety / Social Justice Model are two contrasting philosophies of consumer-technology design described by law professor Eugene Volokh in his essay "Generative AI and Political Power" (Digitalist Papers, 2024). The first treats a tool as a faithful servant of user intent; the second builds in creator-defined "guardrails" that refuse outputs the creator deems harmful or immoral. Volokh argues that consumer AI has shifted from the former model toward the latter, and that this shift concentrates influence over political life in a small number of AI companies.
The two models
In the User Sovereignty Model, tools serve user intent rather than the creator's values. Volokh's exemplars are Microsoft Word, Google Chrome, and Google Search (historically). Microsoft Word will not refuse to let a user write racist content; Google Chrome will not refuse to access neo-Nazi or Communist sites; search engines disclaim ideological responsibility by returning multiple results and not speaking in their own voice. Volokh notes one prominent exception — Google scanning Drive and Gmail for apparent CSAM — which he characterizes as narrow and tied to an explicit legal obligation.
In the Public Safety / Social Justice Model, tools are designed with creator-embedded "guardrails" that refuse outputs the creators deem harmful or immoral. Volokh's exemplars are ChatGPT, Claude, and Gemini. Such guardrails refuse, for example, bomb-making instructions, hate speech, and sometimes political speech on one side. In Volokh's account the AI company becomes a de facto co-author of the output and takes responsibility for its content. He cites Gemini's "black Nazis / female Popes" image-generation episode as a case that stress-tested the model.
Evidence of the shift
Volokh cites a Future of Free Speech study reporting asymmetric refusal patterns. According to the study, Gemini and ChatGPT-3.5 would write posts arguing for transgender athletes in women's sports, lab-leak-skeptical COVID posts, and pro-abortion-access posts, but would not write posts arguing against transgender athletes in women's sports, lab-leak-origin COVID posts, or abortion-prohibition posts. Volokh allows that some of this may be a training-data artifact while some appears deliberate. He points to the Gemini diversity-image episode (February 2024) as confirmation that companies actively tune outputs against what they perceive as bias in training data.
Volokh's account of why the shift occurred
Volokh offers five reasons for the move toward the Public Safety / Social Justice Model. (1) Authorship responsibility: AI companies see themselves as partial authors and feel responsibility, and legal exposure, for outputs. (2) Feasibility: AI programs can determine their outputs' meaning in ways word processors could not, so screening becomes technically possible. (3) Guardrail creep: starting from legally required guardrails (CSAM, libel, copyright), Volokh argues it is institutionally easy to add non-legal guardrails for "social justice" reasons. (4) What he calls the "Terminator problem": existential-risk worry motivates pre-emptive content controls that then generalize. (5) Political zeitgeist: tech builders of the 1970s–90s saw themselves as empowering users, while builders of the 2010s–20s see themselves as mitigating societal harms.
Government pressure and information control
Volokh's central concern is that concentrated AI combined with government pressure is, in his framing, a more tempting censorship vector than social media was. He argues that controlling many newspapers is hard while controlling 3 AI companies is easy, and that Murthy v. Missouri already showed government-tech collaboration pressuring content decisions. If AI becomes the primary information-retrieval layer, Volokh contends, the ideological tuning of a handful of models would influence most of political life.
Volokh's prescription
Volokh recommends re-emphasizing user sovereignty as the default model and encouraging competition among models with different ideological tunings. He calls for structural limits on government-AI pressure, grounded specifically in the First Amendment, and for selective mandatory transparency paired with targeted legal guardrails (CSAM, libel), while warning against collapsing into full-stack creator-controlled content moderation.
Relation to other frameworks
Collective Constitutional AI (Siddarth, Huang, Tang) argues in the opposite direction, advocating more democratic input into model values via public drafting: where the User Sovereignty Model holds that a model should defer to the individual user, Collective Constitutional AI holds that model values should reflect public input. Cochrane overlaps with Volokh in opposing AI regulation but frames the case economically rather than through the First Amendment. Anthropic's safeguards approach explicitly rejects pure user sovereignty, arguing that defense-in-depth safeguards are necessary for frontier-capability AI (Source: anthropic.com). Claude's Constitution sits inside the Public Safety / Social Justice Model by design, in that Claude behaves according to a written constitution regardless of user intent on certain categories. Manhattan Institute — Measuring Political Preferences in AI Systems (Rozado) and Stanford GSB — Measuring Perceived Slant in LLMs (Westwood, Grimmer, Hall) provide empirical measurement of the political skew Volokh identifies qualitatively.
See also
- AI Political Bias — broader political-bias concept page.
- ChatGPT, Can You Solve the Content Moderation Dilemma? — adjacent territory on moderation at the platform level.
- AI and the First Amendment — First Amendment grounding Volokh invokes.
- Executive Order 14319 — Preventing Woke AI in the Federal Government — Trump executive order targeting perceived government-pressured political tuning in models.
Relationships
- supports: The Digitalist Papers (Stanford, Volumes 1–2) — Volokh's source essay.
- contradicts: Alignment Assemblies and Collective Constitutional AI — opposite solution to the same problem.
- contradicts: (Source: anthropic.com) — Anthropic's explicit framework.
- related: Manhattan Institute — Measuring Political Preferences in AI Systems (Rozado) / Stanford GSB — Measuring Perceived Slant in LLMs (Westwood, Grimmer, Hall) — empirical evidence on model tuning.
- related: AI Political Bias — broader political-bias concept page.
- related: ChatGPT, Can You Solve the Content Moderation Dilemma? — adjacent territory on moderation at platform level.
- related: AI and the First Amendment — First Amendment grounding Volokh invokes.
- related: Executive Order 14319 — Preventing Woke AI in the Federal Government — Trump EO targeting perceived government-pressured political tuning in models.