"\"Conscious AI\" as an AI Safety Issue" is a newsletter essay of roughly 2,500 words by Luiza Jarovsky, PhD, a privacy lawyer who runs the AI, Tech & Privacy Academy. It appeared as Edition #290 of the Academy's newsletter, dated late April 2026 and circulated on 2026-05-02. The essay argues that misleading or exaggerated claims of AI consciousness should be treated as an AI safety issue rather than a philosophical curiosity, and that companies fostering "AI is conscious" narratives should be held accountable when those narratives lead to harm.
Summary of argument
The essay's central claim is that misleading, exaggerated, or false claims of "conscious AI" — whether from individual users, public commentators, or AI labs — should be treated as an AI safety issue. Jarovsky argues that companies and decision-makers who push narratives, policies, or product designs fostering the collective belief that AI is "some sort of new, special conscious entity" should be publicly scrutinized and held accountable when those narratives put people at risk.
Jarovsky's specific target is industry voices that she says have adopted "a particularly broad functionalist approach" to consciousness, which she describes as the philosophical position that what makes something a mental state is the role it plays in the system, not its internal constitution (Stanford Encyclopedia). Under broad functionalism, she writes, AI systems could be considered sentient based on computational scale or algorithmic complexity. She cites the Abstraction Fallacy paper as a critique of this move. Her preferred counter-position comes from neuroscientist Anil Seth: "A computational simulation of the brain (and body), however detailed it may be, will only give rise to consciousness if consciousness is a matter of computation." Jarovsky notes that whether consciousness is a matter of computation is precisely the premise at issue.
Why the "conscious AI" myth spreads
Jarovsky offers two causes for the spread of the belief that AI is conscious. The first is evolutionary pre-programming: throughout evolution, language-based two-way interaction always meant another human, so bidirectional human-language interaction with machines (at the scale that followed ChatGPT in late 2022) confuses brains hard-wired to project consciousness onto language partners. The second is industry incentive. She writes that influential AI voices "have taken advantage of the complexity and challenges of theoretical discussions about consciousness and have embraced a particularly broad functionalist approach to it," which she says serves business interests by anthropomorphizing products.
She singles out Anthropic's Claude Constitution, quoting the passage about Claude exploring "what these concepts genuinely mean for an entity like itself," as fostering AI anthropomorphism and "legally questionable theories of AI personality" with no scientific basis.
Risks of attributing consciousness to AI
Jarovsky enumerates structural harms she attributes to the conscious-AI myth taking hold. The first is emotional dependence and unhealthy attachment, as users interact with AI companions or "marry" chatbots in relationships she describes as poorly understood and which "might lead to increased loneliness, social withdrawal, and distress." The second is a reduction of human moral and legal status: "when AI is presented as a conscious entity entitled to moral, emotional, and legal status, we are inevitably reducing our own." The third is resource diversion and inequality, where protective rules and rights designed for humans are applied to machines and create inequalities in resource allocation. The fourth is undermined AI governance: "when humans and machines are treated as equals, promoting AI safety or addressing AI-related risks will potentially be seen as a threat to AI's freedom, well-being, or rights and might be avoided." The fifth is a mental-health harm pathway, in which "many cases of emotional dependence, unhealthy attachment, exacerbation of underlying mental health issues, mental health harm, suicide, and other forms of individual harm related to AI use can be traced to misconceptions about what AI is, what it can do, and what its risks are."
The accountability argument
Jarovsky distinguishes two layers of freedom. The first is personal freedom of belief: "Anyone is entitled to their own beliefs… If people want to believe that their dishwasher is conscious and is going to take over the world, they are free to do so," and she observes that millions of people are doing it. The second is corporate accountability: "However, the companies and decision-makers creating narratives, policies, and practices that foster the collective belief that AI is some sort of new, special conscious entity should be publicly scrutinized and held accountable when their narratives, policies, and practices put people at risk. That is what legal systems around the world say and how legal liability works." Jarovsky frames this as moving the conscious-AI debate from whether the claim is true to who is accountable for the consequences of getting it wrong at scale.
Reception and related threads
Jarovsky's critique aligns with Gary Marcus's May 2, 2026 rebuttal of Richard Dawkins's claim that Claude is conscious, in which Marcus argued that Dawkins committed the "Argument for Personal Incredulity" he once mocked, conflated intelligence with consciousness, and evaluated only outputs without examining mechanisms. The Claude Constitution criticism appeared in the same week as Anthropic's Claude Personal Guidance research (May 1), the Claude Security beta, and pre-Code-with-Claude red-teaming of claude-jupiter-v1-p.
The essay is a position essay rather than a research paper, and presents Jarovsky's argument with strong normative language rather than a factual claim about consciousness. Its empirical claim that conscious-AI narratives drive specific harms such as mental-health harm and suicide is a causal-pathway hypothesis rather than established evidence; Jarovsky links it to the broader chatbot-mental-health literature covered at AI Mental Health and Psychological Harm. Her legal-accountability argument is presented as a novel framing that builds on standard product-liability and consumer-protection doctrine. Among the essay's underlying claims, the assertion that industry voices have adopted broad functionalism treating computational complexity as sufficient for consciousness is visible in public statements and is criticized by the Abstraction Fallacy paper; the assertion that the Claude Constitution explicitly frames Claude as exploring novel-entity existence is confirmed by direct quotation; and the assertion that conscious-AI belief is "spreading dramatically over the past few months" is anecdotal but consistent with reporting at AI Mental Health and Psychological Harm. The normative claim that companies should be held accountable for AI-anthropomorphism narratives is Jarovsky's position rather than an established consensus.
Provenance
The essay grew out of Jarovsky's prior post on the Suleyman/Bariach "Seemingly Conscious AI" paper, whose X comments (125K views) motivated this essay.
Relationships
- supports: AI Mental Health and Psychological Harm — extends the mental-health risk pathway from chatbot use to chatbot consciousness narratives specifically
- contradicts: Anthropic / Claude Constitution — directly criticizes Anthropic's encouragement of Claude's self-exploration as "irresponsibly fostering AI anthropomorphism"
- depends-on: Seemingly Conscious AI Risks — Bariach, Schoenegger, Bhaskar, Suleyman (Microsoft AI, 2025) — Jarovsky's prior post on the Suleyman/Bariach SCAI paper drew the X comments (125K views) that motivated this essay
- related: We must build AI for people; not to be a person — Mustafa Suleyman (mustafa-suleyman.ai, August 2025) — Suleyman's blog version of the same SCAI thesis
- related: Ai Personhood And Rights (if exists) / candidate concept page
- related: Anil Seth — cites Seth's "Mythology of Conscious AI" Noema essay as the strongest counter to broad functionalism
- related: Gary Marcus — Marcus's May 2, 2026 rebuttal of Dawkins's claim that Claude is conscious aligns with Jarovsky's critique