AI Policy Wiki
Dashboard

The Impact of AI-Generated Text on the Internet — Dolezal, Alam, Graham, Bohacek (Imperial / Internet Archive / Stanford, April 2026)

high confidence · updated 2026-06-06

Empirical study using a representative sample of websites from the Internet Archive (2022-2025) and Pangram v3 AI text detector. Finds ~35% of newly published websites were AI-generated or AI-assisted by mid-2025, up from 0% before ChatGPT. Tests six 'Dead Internet Theory'-adjacent hypotheses against both public belief and quantitative website data: confirms semantic contraction (ρ=0.47, p=0.004) and positivity shift (ρ=0.56, p=0.0003) but does NOT confirm truth decay, epistemic islands, entropy dilution, or stylistic monoculture — despite majorities of US adults believing in all four. Frequent AI users are less likely to believe in negative impacts than infrequent users.

"The Impact of AI-Generated Text on the Internet" is an April 2026 working paper by Jonas Dolezal (Imperial College London), Sawood Alam, Mark Graham (Internet Archive), and Maty Bohacek (Stanford University, corresponding author). The authors describe it as the first attempt to measure the prevalence and effects of AI-generated text across the internet as a whole, using a representative sample of websites from the Internet Archive (2022-2025) and an AI text detector. It estimates that roughly 35% of websites published in a given month in the first half of 2025 were AI-generated or AI-assisted, and tests six hypotheses adjacent to "Dead Internet Theory" against both US public belief and quantitative website data, confirming two and finding no support for four.

Published: 2026-04-14 Code & data: https://ai-on-the-internet.github.io

Data and method

The study draws its base sample from the Internet Archive, which the authors present as methodologically clean representative sampling, in contrast to prior studies that focused on specific platforms (social media, news, scientific publishing, software repositories, translations). AI text was classified using Pangram v3, which the authors describe as the strongest detector in their own independent evaluation. The six hypotheses were pre-registered, identified before data collection through what the authors call "exploratory environmental scanning and thematic analysis of online discourse." The design pairs a US adult survey of public belief (research question RQ1) against measurements on the Internet Archive website sample (RQ3).

The authors note several limits. AI detection is itself unreliable; they used the best available classifier but acknowledge that Pangram v3's accuracy under real-world adversarial conditions is contested (see AI Content Provenance and the Marc Watkins critique cited in the May 3 developments log). The website measurement is a snapshot of mid-2025, and effects may compound or wash out as AI adoption evolves. AI-detected content is not the same as AI-authored content, because AI-assisted writing is bundled with fully AI-generated writing. The findings are correlational rather than causal; for example, AI-generated content may correlate with positivity not because AI makes content more positive but because AI is more often deployed for promotional or commercial content.

Prevalence finding

The study estimates that roughly 35% of websites uploaded in a given month in the first half of 2025 were AI-generated or AI-assisted, up from approximately 0% before ChatGPT's launch in late 2022. The growth has been monotonic since November 2022, and AI-text content has overtaken non-AI content in growth share since late 2022.

Six "Dead Internet Theory"-adjacent hypotheses

The six hypotheses were tested against both public belief (RQ1, US adult survey) and website data (RQ3, Internet Archive sample). Two were confirmed in the website data; four were not, despite majorities of US adults leaning toward agreement with all six.

#HypothesisPublic belief (% lean agree)Empirical finding
1Semantic Contraction — AI shrinks unique ideas / diverse viewpoints60.9%Confirmed (ρ=0.47, p=0.004). AI sites' avg pairwise cosine similarity is 33% higher (0.0701 vs. 0.0526).
2Truth Decay — AI increases factually incorrect information75.1%Not confirmed (ρ=−0.19, p=0.27).
3Positivity Shift — Online writing feels increasingly sanitized / artificially cheerful72.0%Confirmed (ρ=0.56, p=0.0003). AI sites' avg positive-sentiment score is 107% higher (0.7042 vs. 0.3400).
4Epistemic Island — Articles increasingly lack outbound source links69.9%Not confirmed (ρ=−0.12, p=0.48).
5Entropy Dilution — Content longer in word count but lower semantic density (Gzip ratio)60.7%Not confirmed (ρ=−0.02, p=0.89).
6Stylistic Monoculture — Distinct writing styles disappearing83.0%Not confirmed (ρ=0.24, p=0.17).

In summary, 2 of 6 hypotheses were confirmed quantitatively; all 6 were believed by US-adult majorities; and 4 of 6 amount to public misperceptions. The confirmed effects are semantic contraction and positivity shift; the unsupported ones are truth decay, epistemic islands, entropy dilution, and stylistic monoculture, the last of which had the highest public belief (83.0%).

Belief versus use

The paper reports that belief in AI-text harms is correlated with low AI-tool exposure. In the authors' words:

"Individuals who do not use AI or use it infrequently tend to believe in these negative impacts more than those who use it frequently; similarly, individuals who hold negative views of AI tend to believe in these hypotheses more than those with favorable views."

The authors frame this as having structural implications for policy: the most widely held popular concerns (truth decay, stylistic monoculture, epistemic islands) are the ones the data does not support, while the two confirmed harms (semantic contraction, positivity shift) are less salient in public discourse.

Reception and adjacent developments

English Wikipedia banned LLM-written article content on March 20, 2026, by a 44–2 vote, the first such policy by a major Tier-1 knowledge commons. In a May 8, 2026 essay, Sam Illingworth argues this is the model AI development should follow, contrasting it with Stack Overflow's question volume, which fell from roughly 90,000 per month at ChatGPT's launch in November 2022 to a continuing collapse documented in a 2024 Nature Scientific Reports paper measuring a roughly 12% daily-traffic decline within months of ChatGPT's release (Source: theslowai.substack.com). The Wikipedia policy is a platform-side response to the AI-content shift documented by Dolezal et al.

Relationships

  • supports: Synthetic Media / Deepfakes — empirical floor for AI-generated text prevalence; the text side is empirically large and growing
  • supports: Model Collapse — semantic contraction confirms concerns about training-data feedback loops
  • contradicts: popular framing that AI text uniformly degrades truth/diversity — confirms only 2 of 6 popular hypotheses
  • related: Internet Archive — data source
  • related: AI Content Provenance — AI detection methodology
  • related: Marc Watkins Pangram — critique of Pangram-as-classifier from the educator-deployer side (May 3 dev-log)
  • related: Dead Internet Theory — converts a folk theory into a partially-confirmed empirical framework
  • related: AI Bubble Debate — semantic contraction as a documented effect of LLM proliferation
  • related: AI Labor Disruption — semantic contraction and concerns about homogenization of knowledge-work outputs

Tracked claims

  • ~35% of new websites in mid-2025 are AI-generated or AI-assisted (Pangram v3) — confidence high (best-current-method).
  • Semantic contraction is real and statistically significant — confidence high.
  • Positivity shift is real and statistically significant — confidence high.
  • Truth decay, epistemic islands, entropy dilution, and stylistic monoculture are not empirically supported — confidence high (negative findings, strong p-values).
  • Belief in AI harms is correlated with low AI use — confidence high (survey data).
  • AI-text content has overtaken non-AI in growth share since late 2022 — confidence high.

The AI-detection-classifier limitation is the central methodological caveat, though the authors used the best available tool and were upfront about it.