A 2024 paper by Emmanuel Vargas Penagos (Örebro University), published in the International Journal of Law and Information Technology (doi: 10.1093/ijlit/eaae028). It is a qualitative test of ChatGPT and OpenAI's GPTs service applied to real content moderation decisions, framed within EU law — specifically the EU Digital Services Act (DSA) and the European human rights framework on freedom of expression. The paper argues that while large language models offer some benefits for content moderation, they introduce significant human rights challenges, and that the claim that LLMs can "solve" content moderation is premature.
The paper responds to OpenAI's August 2023 proposal to use GPT-4 for content moderation, evaluating that claim systematically from a legal and human rights perspective.
Summary of argument
The paper defines content moderation as "the systems and rules that determine how [platforms] treat user-generated content on their services." Its central concern is the judgment problem: moderation decisions require contextual interpretation (the same content may be harmful in one context and not another), are culturally contingent (definitions of hate speech vary across jurisdictions), are legally complex (the EU DSA, national laws, the ECHR, and UN human rights norms all apply), and involve value tradeoffs between free expression and harm prevention. The paper cites the UN Special Rapporteur on freedom of expression, who has called platform companies "enigmatic regulators, establishing a kind of 'platform law', in which clarity, consistency, accountability and remedy are elusive."
Against this backdrop, the paper assesses what LLMs can and cannot contribute. Among the benefits it identifies are pattern matching at scale (identifying clear-cut violations faster than human reviewers), consistency within a given prompt framing, and the ability to incorporate legal framework text into reasoning. Against these it sets four human rights challenges:
- Cultural and contextual failure: LLMs trained on predominantly Western and English-language data underperform on non-English content and culturally specific contexts.
- Opacity: explanations for content decisions lack the legal specificity required for appeals under the DSA.
- Bias amplification: existing biases in training data are systematically applied at scale, creating a structural discrimination risk.
- Accountability gap: automated moderation without human review removes the human judgment that rights frameworks require.
The analysis is framed within the EU Digital Services Act, which the paper describes as requiring clear and accessible moderation policies, notice-and-action procedures for illegal content, complaint and redress mechanisms, transparency reporting for large platforms, and human review of automated moderation decisions. Its key finding is that current LLM-based moderation systems do not meet DSA requirements for explainability, appeal rights, or human oversight.
The paper characterizes the result as an instance of AI capability that is technically strong at content detection but weak at legal-quality judgment, and it situates content moderation as a domain where existing laws such as the DSA and the ECHR already govern AI in practice, before any AI-specific regulation.
Relationships
- instance-of: AI as Normal Technology — content moderation as a domain where capability in pattern matching coexists with failure in cultural and legal judgment.
- related: EU AI Act (Regulation 2024/1689) — content moderation systems affecting many users may fall under the EU AI Act's high-risk or transparency requirements; the paper examines the adjacent DSA framework.
- related: AI and the First Amendment — a US/EU comparison point: the EU uses the DSA and the ECHR to govern AI-mediated speech, while the US debate centers on whether AI outputs are themselves speech.
- related: Algorithmic Pricing and Antitrust — like algorithmic pricing, algorithmic content moderation illustrates AI systems making consequential decisions that existing legal frameworks were not designed to govern.
- related: The Emergence of Artificial Intelligence Ethics Auditing — content moderation is the kind of high-stakes deployment that AI auditing frameworks are being built to address.
Provenance
Emmanuel Vargas Penagos (Örebro University), "ChatGPT, Can You Solve the Content Moderation Dilemma?", International Journal of Law and Information Technology, 2024 (doi: 10.1093/ijlit/eaae028). Published 2024-11-26.