This episode examines a study analyzing the growth of AI-generated content across the open web, drawing on 33 monthly samples from the Internet Archive's Wayback Machine between August 2022 and May 2025. It highlights the paper's central finding that AI-generated or AI-assisted content on newly published websites rose from zero before ChatGPT's launch to roughly 35 percent by mid-2025, and explores how the authors transform "Dead Internet Theory" from internet folklore into six testable hypotheses — including semantic contraction, truth decay, positivity shift, epistemic islands, entropy dilution, and stylistic monoculture. The discussion covers the methodology behind sampling a representative slice of the internet, including logarithmic downsampling and stratification across time, MIME type, and domain to avoid bias toward heavily-crawled sites. It also connects the findings to the concept of model collapse, framing the 35 percent figure as empirical evidence for a previously theoretical concern about AI models training on their own synthetic output. Listeners interested in web ecosystem health, LLM training data quality, or the intersection of internet culture and rigorous data science will find the episode's blend of meme-to-metric translation particularly compelling. Sources: 1. The Impact of AI-Generated Text on the Internet — Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek, 2026 http://arxiv.org/abs/2604.26965 2. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature — Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, Chelsea Finn, 2023 https://scholar.google.com/scholar?q=DetectGPT%3A+Zero-Shot+Machine-Generated+Text+Detection+using+Probability+Curvature 3. Can AI-Generated Text be Reliably Detected? — Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, Soheil Feizi, 2023 https://scholar.google.com/scholar?q=Can+AI-Generated+Text+be+Reliably+Detected%3F 4. A Watermark for Large Language Models — John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, Tom Goldstein, 2023 https://scholar.google.com/scholar?q=A+Watermark+for+Large+Language+Models 5. GLTR: Statistical Detection and Visualization of Generated Text — Sebastian Gehrmann, Hendrik Strobelt, Alexander M. Rush, 2019 https://scholar.google.com/scholar?q=GLTR%3A+Statistical+Detection+and+Visualization+of+Generated+Text 6. The Curse of Recursion: Training on Generated Data Makes Models Forget — Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, Ross Anderson, 2024 (Nature, updated from 2023 preprint) https://scholar.google.com/scholar?q=The+Curse+of+Recursion%3A+Training+on+Generated+Data+Makes+Models+Forget 7. Is the Internet Dead? Evaluating Claims about Web Homogenization from AI-Generated Content (representative of the broader 2024-2025 measurement literature, e.g. Muzumdar et al. and related environmental-scanning studies of Dead Internet Theory discourse) — Various (this cluster of work is cited in the paper as Muzumdar et al., 2025 and similar), 2025 https://scholar.google.com/scholar?q=Is+the+Internet+Dead%3F+Evaluating+Claims+about+Web+Homogenization+from+AI-Generated+Content+%28representative+of+the+broader+2024-2025+measurement+literature%2C+e.g.+Muzumdar+et+al.+and+related+environmental-scanning+studies+of+Dead+Internet+Theory+discourse%29 8. Studies on bot/inauthentic-content prevalence on specific platforms (e.g., La Cava et al. 2025 on social media, Matatov et al. 2024 on platform-specific AI content) — Lucio La Cava et al.; Jonathan Matatov et al. (representative platform-specific studies cited in the paper's related-work section), 2024-2025 https://scholar.google.com/scholar?q=Studies+on+bot%2Finauthentic-content+prevalence+on+specific+platforms+%28e.g.%2C+La+Cava+et+al.+2025+on+social+media%2C+Matatov+et+al.+2024+on+platform-specific+AI+content%29 9. Public Trust and Perceptions of Artificial Intelligence (Ipsos / Reuters Institute Digital News Report and Edelman Trust Barometer AI-focused editions) — Ipsos (various); Reuters Institute for the Study of Journalism (Nic Newman et al.); Edelman Trust Barometer team, 2023-2025 (recurring annual) https://scholar.google.com/scholar?q=Public+Trust+and+Perceptions+of+Artificial+Intelligence+%28Ipsos+%2F+Reuters+Institute+Digital+News+Report+and+Edelman+Trust+Barometer+AI-focused+editions%29 10. Americans' Views of Artificial Intelligence (Pew Research Center recurring survey series) — Pew Research Center (Alec Tyson, Emma Kikuchi, and colleagues), 2023-2025 (recurring) https://scholar.google.com/scholar?q=Americans%27+Views+of+Artificial+Intelligence+%28Pew+Research+Center+recurring+survey+series%29 11. The Perception Gap: Comparing Public Beliefs about Misinformation to Empirical Prevalence Estimates (representative of the risk-perception vs. measured-prevalence literature this paper's framing descends from, e.g. work following Duffy et al. and general misperception-of-misinformation-prevalence studies) — Andrew Guess, colleagues in the misinformation-prevalence research cluster (representative of this line, distinct from the AI-specific surveys above), 2019-2023 (foundational misinformation-perception literature) https://scholar.google.com/scholar?q=The+Perception+Gap%3A+Comparing+Public+Beliefs+about+Misinformation+to+Empirical+Prevalence+Estimates+%28representative+of+the+risk-perception+vs.+measured-prevalence+literature+this+paper%27s+framing+descends+from%2C+e.g.+work+following+Duffy+et+al.+and+general+misperception-of-misinformation-prevalence+studies%29 12. Documenting the English Colossal Clean Crawled Corpus (C4) — Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner, 2021 https://scholar.google.com/scholar?q=Documenting+the+English+Colossal+Clean+Crawled+Corpus+%28C4%29 13. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only — Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, Julien Launay, 2023 https://scholar.google.com/scholar?q=The+RefinedWeb+Dataset+for+Falcon+LLM%3A+Outperforming+Curated+Corpora+with+Web+Data%2C+and+Web+Data+Only 14. Quantifying Memorization Across Neural Language Models / broader Common Crawl representativeness and bias studies (e.g., work on Common Crawl's domain and language skew) — Various (Common Crawl bias/representativeness literature, e.g. work by Luccioni & Viviano on Common Crawl content quality, and follow-on studies), 2021-2023 https://scholar.google.com/scholar?q=Quantifying+Memorization+Across+Neural+Language+Models+%2F+broader+Common+Crawl+representativeness+and+bias+studies+%28e.g.%2C+work+on+Common+Crawl%27s+domain+and+language+skew%29 15. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research — Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, et al. (Allen Institute for AI), 2024 https://scholar.google.com/scholar?q=Dolma%3A+an+Open+Corpus+of+Three+Trillion+Tokens+for+Language+Model+Pretraining+Research 16. AI models collapse when trained on recursively generated data — I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, Y. Gal, 2024 https://scholar.google.com/scholar?q=AI+models+collapse+when+trained+on+recursively+generated+data 17. Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks — V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, S. Feizi, 2025 https://scholar.google.com/scholar?q=Can+AI-generated+text+be+reliably+detected%3F+Stress+testing+AI+text+detectors+under+various+attacks 18. RAID: A shared benchmark for robust evaluation of machine-generated text detectors — L. Dugan, A. Hwang, F. Trhlik, A. Zhu, J. M. Ludan, H. Xu, D. Ippolito, C. Callison-Burch, 2024 https://scholar.google.com/scholar?q=RAID%3A+A+shared+benchmark+for+robust+evaluation+of+machine-generated+text+detectors 19. When incentives backfire, data stops being human — S. Santy, P. Bhattacharya, M. H. Ribeiro, K. Allen, S. Oh, 2025 https://scholar.google.com/scholar?q=When+incentives+backfire%2C+data+stops+being+human 20. Longitudinal sampling of URLs from the Wayback Machine — K. Garg, S. Alam, D. Ayala, M. Graham, M. C. Weigle, M. L. Nelson, 2025 https://scholar.google.com/scholar?q=Longitudinal+sampling+of+URLs+from+the+Wayback+Machine 21. Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity — J. Zhang, S. Yu, D. Chong, A. Sicilia, M. R. Tomz, C. D. Manning, W. Shi, 2025 https://scholar.google.com/scholar?q=Verbalized+sampling%3A+How+to+mitigate+mode+collapse+and+unlock+LLM+diversity Interactive Visualization: AI-Generated Text and the Death of the Open Web