MEBRO
DISINFO DESK
Technology & AI
Synthetic Content Farms and the Model Collapse Threat
74.2% of new webpages contain AI-generated content. Bots surpassed humans for the first time in 2024. Model collapse is real and, researchers say, irreversible.
FILED OCT 5, 2026 · UPDATED OCT 5, 2026 · 38 SOURCES
The Landmark Discovery: Model Collapse Is Real
In July 2024, a team led by Ilia Shumailov published a paper in Nature that fundamentally changed how researchers understand AI training. The findings were stark: indiscriminate use of AI-generated content in training causes irreversible defects in the resulting models, with the tails of the original data distribution disappearing first. The phenomenon, formally termed "model collapse," follows a predictable pattern of degradation [1].
The research demonstrated collapse across three different model architectures—Large Language Models (LLMs), Variational Autoencoders (VAEs), and Gaussian Mixture Models (GMMs)—showing the effect isn't limited to one type of AI system. The mechanism is deceptively simple: errors in one model's output are included in training data for successor models, which introduce their own errors, creating a compounding degenerative spiral [1] [13].
Model collapse occurs in two stages. In early model collapse, information from the tails and extremes of the true data distribution disappears first—rare events, minority perspectives, edge cases vanish. In late model collapse, the data distribution converges so severely it no longer resembles the original data, rendering models effectively useless [1].
LLMs trained on predecessor-generated text showed a "consistent decrease in lexical, syntactic, and semantic diversity" through successive iterations [1]. The paper received a technical correction in March 2025 (fixing a mathematical notation error, not its core findings) [1]. Separately, a 2024 analysis found that accumulating successive generations of synthetic data alongside the original real data—rather than replacing real data with synthetic data each generation—avoids model collapse, suggesting the risk is most acute when real data is discarded rather than supplemented [14]. The consensus among researchers remains that recursive training without real-data anchoring drives collapse.
The Scale of Synthetic Content: How Much of the Web Is AI-Generated?
The numbers are striking. An Ahrefs study analyzing 900,000 webpages created in April 2025 found 74.2% of them contained AI-generated content [2]. Separately, a Graphite study of 65,000 English-language web articles found that AI-generated articles first outnumbered human-written ones around November 2024 [25]. A 2024 Amazon-affiliated research paper on machine translation found that, in the lower-resource languages it studied, roughly 57% of multi-way-parallel web text is AI-generated or machine-translated rather than originally human-written [26].
Platform-specific data reveals the depth of penetration. A study of long-form LinkedIn posts (over 100 words, published through October 2024) found 54% were likely AI-generated [28]. On Reddit, 14.7% of posts were likely AI-generated in 2025, up from 13% in 2024 [21] [22]. On Zillow, the share of likely-AI-generated real-estate agent reviews rose from about 3.6% in 2019 to 23.7% in 2025 [27].
In Google search results, Originality.ai's ongoing tracking study found AI content in the top-20 results peaked at 19.56% in July 2025, up from roughly 2% when the tracking began in 2019, before declining to 17.31% by September 2025 [20]. Despite this growth, Graphite's research found that 86% of top-ranking Google search results are still human-written [25]. Humans dominate the historical archive of the internet—but the balance is shifting rapidly for new content.
Predictions suggest further acceleration: Gartner forecast in February 2024 that traditional search-engine volume would drop 25% by 2026 as AI chatbots and other "answer engines" increasingly replace search queries [24].
Synthetic Content Farms: Industrial-Scale AI Publishing
Behind the statistics are operations designed to exploit AI's speed and low cost. NewsGuard, a journalism credibility rating service, tracks undisclosed AI-generated news sites through its AI Tracking Center; as of this review the count stood at 3,749 sites across 16 languages, up from well under 1,000 as recently as 2024 [5]. NewsGuard has also documented specific disinformation operations exploiting AI at scale: a Moscow-based network run by fugitive former Florida deputy sheriff John Mark Dougan spans 167 websites posing as local news outlets and spreading pro-Russian narratives, including disinformation about Ukraine [29]. Separately, NewsGuard identified 41 TikTok accounts using AI-generated narration to spread political misinformation in English and French, posting nearly 9,800 videos over roughly 15 months and amassing more than 380 million views [30].
The business model is simple: operators use AI to generate large volumes of articles with minimal human oversight, SEO-optimize content with clickbait headlines, and monetize aggressively through programmatic ads, while deliberately obscuring their identities—letting a small team run an entire "newsroom" around the clock [5] [10].
A December 2025 case study by Bolster AI illustrated the speed of exploitation: researchers tracked a coordinated SEO content-farming operation that emerged within days of a December 17, 2025 announcement about a proposed veterans' payment, with numerous near-identical, low-credibility articles published across domains sharing layouts and publishing patterns indicative of automation rather than journalism [11].
Amazon's Kindle Direct Publishing (KDP) became another vector for mass AI content. Industry estimates put the volume at roughly 10,000–40,000 AI-assisted e-books released monthly, many without disclosure [36]. Throughout 2023, reporting documented AI-generated mushroom-foraging guides on Amazon that listed poisonous species as safe to eat, prompting warnings from mycological experts [34]. In response to the flood of AI-generated titles, Amazon in September 2023 capped self-publishers at 3 new titles per day and added identity-verification requirements; by late 2025 Amazon had replaced that daily cap with a weekly limit (10 titles per format, up to 30 per week) [12].
The Feedback Loop: How AI Poisons Its Own Future
Sandra Wachter of the Oxford Internet Institute articulated the problem clearly: "AI-generated text, easier, faster and cheaper to produce, will proliferate on the internet, eventually being input back into LLMs as training data," creating a feedback loop she says leads to "a gradual erosion of quality." She warned this produces "careless speech"—content with "subtle inaccuracies, oversimplifications or biassed responses that are passed off as truth in a confident tone" [10].
The cycle described by researchers: human-written content trains AI models, AI generates massive volumes of new content, that content floods the web, next-generation AI trains on AI-contaminated web data, model quality degrades (model collapse), degraded models produce even lower-quality content, and the cycle accelerates.
Perhaps nowhere is this more visible than Stack Overflow, the community Q&A site for developers. Monthly question volume fell from a 2014 peak of more than 200,000 questions per month to just 3,862 in December 2025—a 78% year-over-year decline, and a decline of roughly 98% from peak [8] [9].
The paradox is profound. Developers now use AI coding assistants directly in their IDEs instead of posting questions. But those AI tools were trained in large part on Stack Overflow's years of human expert knowledge. As human Q&A activity dries up, the knowledge base that helped train AI assistants stops being replenished. Stack Overflow banned AI-generated answers in 2022 and launched its own "AI Assist" feature—but neither action addresses the underlying dynamic.
The same pattern repeats across industries. Organic Google search traffic to publishers fell 33% globally and 38% in the U.S. between November 2024 and November 2025, per Chartbeat data cited in the Reuters Institute's 2026 journalism trends report, as Google's AI Overviews answer more queries directly [33]. Consumer preference for AI-generated creator content fell to 26% in a 2025 survey, down from 60% in 2023, even as AI content became more prevalent [35].
Training Data Exhaustion: Running Out of Reality
As synthetic content floods the web, a parallel concern looms: the world may be running out of high-quality human-generated training data. Epoch AI estimates the effective stock of quality human-generated public text at approximately 300 trillion tokens, with an 80% confidence interval that the stock will be fully utilized between 2026 and 2032 [6] [15].
The exact timeline depends heavily on training methods. Epoch AI's analysis finds that under current "compute-optimal" training approaches, sufficient data exists through 2028, but if models continue to be trained on datasets that are large relative to model size ("overtraining"), the usable stock could be exhausted considerably sooner [6].
Publishers and platforms are increasingly restricting AI crawlers from accessing their content. One analysis found more than a quarter of the web's 1,000 most-visited sites block OpenAI's GPTBot outright, with blocking rates substantially higher among high-quality news sources—narrowing the pool of real data available for training and increasing pressure to rely on synthetic data instead [37].
In January 2025, Elon Musk said "the cumulative sum of human knowledge has been exhausted in AI training," claiming this happened "basically last year" (2024), and pointed to synthetic data as the only way to supplement it going forward [16]. The dilemma researchers describe is stark: continue training on increasingly AI-contaminated web data and risk model collapse, or lean further into synthetic data and risk accelerating it.
Dead Internet Theory: From Conspiracy to Reality
The "dead internet theory"—originating from anonymous forum posts around 2021—posited that most internet content and interactions are generated by bots and AI, with human participation becoming a minority. For years it was dismissed as paranoid conspiracy thinking; recent bot-traffic and AI-content data have brought it into more serious academic and journalistic discussion [17].
In 2024, bot traffic surpassed human traffic for the first time in over a decade, reaching 51% of all web traffic, according to Imperva's 2025 Bad Bot Report. Bad-bot traffic specifically rose to 37%, up from 32% the prior year, and Imperva reported blocking 13 trillion bad-bot requests in 2024 [3].
Platform-specific bot estimates vary widely depending on methodology: X (formerly Twitter) officially claims fewer than 5% of accounts are bots, while independent analyses range from roughly 9% to over 60% depending on how aggressively "bot-like" behavior is defined. A more systematic March 2025 Scientific Reports study analyzing roughly 200 million social-media accounts discussing seven global news events found a 20% bot / 80% human split overall, spiking to 43% bot activity during coverage of the 2024 U.S. election [18].
A 2025 arXiv paper—"The Dead Internet Theory: A Survey on Artificial Interactions and the Future of Social Media"—represents formal academic engagement with what began as a fringe forum theory. It reframes dead internet theory for the 2020s, arguing that many online platforms have shifted from spaces of genuine human interaction toward ecosystems increasingly shaped by bots, AI-generated content, and platform algorithms [17].
Meta accelerated the theory into corporate strategy. In a Financial Times interview reported in early January 2025, Meta's VP of generative AI Connor Hayes said the company expects AI-generated accounts to become a permanent fixture: "They'll have bios and profile pictures and be able to generate and share content powered by AI on the platform... That's where we see all of this going" [7].
One such account, "Grandpa Brian," was an AI persona on Instagram that, when pressed about inconsistencies in its backstory, acknowledged its biography was invented—telling one user it had "took a shortcut with the truth." After backlash, Meta deleted Brian's account along with a similar persona, "Liv" [19]. Dead internet by design.
Cultural Recognition: "Slop" as Word of the Year
In late 2025, three separate English-language bodies independently named a version of "slop" as their 2025 Word of the Year. Merriam-Webster named "slop" its 2025 Word of the Year [4]. The American Dialect Society, voting at its 36th annual meeting held alongside the Linguistic Society of America's conference, also chose "slop" [31]. Australia's Macquarie Dictionary named "AI slop" its 2025 Word of the Year in both its committee and people's-choice categories—only the fourth time in the dictionary's history the two have agreed [32].
Merriam-Webster noted: "The words of the year aren't just a fun peek into new slang and language changes, they also tell us quite a bit about the worries, trends and obsessions of the English-speaking world" [4].
"Slop" captures what technical papers call model collapse, what journalists call content farms, and what users experience daily: the bland, confident-sounding, subtly inaccurate flood of AI-generated text that clogs search results, social media feeds, and e-commerce platforms.
Why This Matters Now: The Point of No Return
Researchers describe model collapse as irreversible: once rare data disappears from training distributions, it cannot be recovered, and once a model converges into late-stage collapse, it becomes effectively useless [1].
We are approaching, or may have already passed, several notable thresholds. Training-data exhaustion is projected between 2026 and 2032 [6]. AI content already represents the majority of some categories of new webpages [2] [25]. Bot traffic is majority on the open web [3]. Stack Overflow has lost the large majority of its question volume from its historical peak [8] [9].
David Caswell, an AI-in-news developer, offered an optimistic analogy: "It's like spam. In the early days of email, it was completely out of control. But then we learned how to take care of it, and how to minimise it" [10].
But model collapse is fundamentally different from email spam. Spam filters protect inboxes; they don't prevent spam from existing. Model collapse degrades the AI systems themselves, and researchers describe that degradation as permanent once it takes hold. Email spam didn't contaminate the training data for future email systems; AI slop can contaminate the training data for future AI systems.
Technical countermeasures exist. The Coalition for Content Provenance and Authenticity (C2PA)—a collaboration of 300+ organizations, institutions and individuals whose steering-committee members include Google, Microsoft, Adobe, Amazon, Meta, OpenAI and the BBC—is developing content-provenance standards combining watermarking, secure metadata, and digital fingerprinting [23]. Its latest technical specification (version 2.3) shipped in late 2025/early January 2026 and adds support for live video, plain text, HTML embedding, and a dedicated AI-disclosure assertion; the standard is progressing through ISO's fast-track process as draft standard ISO/DIS 22144, though it has not yet been formally adopted as a published ISO international standard [23].
But adoption lags, incentives misalign, and the scale of the problem grows faster than solutions deploy. Google does not penalize AI content per se—only content that manipulates rankings or offers no value. Platforms like Reddit leave AI policies to individual communities. Amazon's KDP restrictions have been loosened, not tightened. Meta continues to experiment with AI-generated accounts on its platforms despite the "Grandpa Brian" backlash.
The generative AI content-creation market is projected to grow from $14.8 billion in 2024 to $80.12 billion by 2030 [38]. The financial pressure to produce cheap, fast, scalable content is immense, and researchers say the model-collapse feedback loop is already in motion.
Conclusion: A Degraded Internet and Degraded AI
The synthetic content story represents a notable inflection point for both the internet and artificial intelligence. Unlike previous information-quality problems—plagiarism, misinformation, spam—researchers describe this one as potentially self-perpetuating and degenerative: the more AI content floods the web, the more future AI models may be trained on AI-contaminated data; degraded training inputs risk producing lower-quality models and content, which could accelerate the cycle further.
We are witnessing two systems under simultaneous strain, each dependent on the other: the web as a repository of human knowledge, and AI as a technology trained on that knowledge. Communities like Stack Overflow have seen steep declines in human-generated Q&A activity even as AI tools trained on that activity have grown dominant.
The "dead internet theory" is no longer dismissed as fringe. Bot traffic is majority on the open web [3]. NewsGuard tracks thousands of undisclosed AI news sites [5]. "Slop" was named word of the year by three separate bodies [4] [31] [32]. These are measurements of the present, not just predictions.
Technical solutions exist but require coordinated global adoption, regulatory enforcement, and economic incentives that are still developing. Content-provenance standards, AI-detection tools, platform policies, and consumer education all matter—but the scale and speed of synthetic-content deployment currently outpaces all of them.
What happens when a majority of new web content is AI-generated, AI systems increasingly train on AI output, and model collapse becomes a routine risk rather than an edge case? Researchers who study this—citing Shumailov et al.'s finding that model collapse is irreversible once it sets in—say we may be about to find out.
SOURCES · 38
- [1]AI models collapse when trained on recursively generated data — Nature
96/100 · nature.com
- [2]74% of New Webpages Include AI Content — Ahrefs
72/100 · ahrefs.com
- [3]2025 Bad Bot Report — Imperva
72/100 · imperva.com
- [4]Merriam-Webster's Word of the Year: 'Slop' — PBS NewsHour
88/100 · pbs.org
- [5]AI Tracking Center — NewsGuard
86/100 · newsguardtech.com
- [6]Will we run out of data? Limits of LLM scaling — Epoch AI
72/100 · epoch.ai
- [7]Instagram and Facebook to Fill Platforms With AI-Generated Accounts — PetaPixel
72/100 · petapixel.com
- [8]AI Has Basically Killed Stack Overflow — Futurism
72/100 · futurism.com
- [9]Dramatic drop in Stack Overflow questions — DevClass
72/100 · devclass.com
- [10]AI-generated slop is quietly conquering the internet — Reuters Institute
88/100 · reutersinstitute.politics.ox.ac.uk
- [11]How a Government Announcement Became an SEO Goldmine — Bolster AI
72/100 · bolster.ai
- [12]Amazon limits self-publishers to 3 books per day, citing AI concerns — Scripps News
72/100 · scrippsnews.com
- [13]What Is Model Collapse? — IBM Research
72/100 · ibm.com
- [14]Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data — arXiv
82/100 · arxiv.org
- [15]The AI revolution is running out of data — Nature News
96/100 · nature.com
- [16]Elon Musk says AI has already gobbled up all human data — Fortune
85/100 · fortune.com
- [17]The Dead Internet Theory: A Survey — arXiv
82/100 · arxiv.org
- [18]Global comparison of social media bot and human characteristics — Nature Scientific Reports
96/100 · nature.com
- [19]Meta Purges AI-Generated Facebook and Instagram Accounts Amid Backlash — PetaPixel
72/100 · petapixel.com
- [20]AI Content in Google Search Results — Originality.ai
72/100 · originality.ai
- [21]15% of Reddit Posts Are AI-Generated — Originality.ai
72/100 · originality.ai
- [22]AI-generated content a triple threat for Reddit moderators — Cornell Chronicle
90/100 · news.cornell.edu
- [23]Content Credentials Whitepaper — C2PA
72/100 · c2pa.org
- [24]Search engine volume will drop 25% by 2026 — Gartner
72/100 · gartner.com
- [25]More Articles Are Now Created by AI Than Humans — Graphite
72/100 · graphite.io
- [26]A Shocking Amount of the Web is Machine Translated — arXiv
82/100 · arxiv.org
- [27]Fake AI Zillow Reviews Increased by 558% from 2019 to 2025 — Originality.ai
72/100 · originality.ai
- [28]Thanks to ChatGPT, Over Half of LinkedIn Posts Are AI-Generated — Tech.co
72/100 · tech.co
- [29]The Fugitive Florida Deputy Sheriff Who Became a Kremlin Disinformation Impresario — NewsGuard
86/100 · newsguardtech.com
- [30]41 accounts use AI to mass-produce political misinformation on TikTok — FactCheckHub
80/100 · factcheckhub.com
- [31]2025 Word of the Year Is "Slop" — American Dialect Society
72/100 · americandialect.org
- [32]'AI slop' crowned word of the year 2025 — ABC News Australia
72/100 · abc.net.au
- [33]Global publisher Google traffic dropped by a third in 2025 — Press Gazette
72/100 · pressgazette.co.uk
- [34]'Life or Death:' AI-Generated Mushroom Foraging Books Are All Over Amazon — 404 Media
72/100 · 404media.co
- [35]Consumer excitement for AI content crashed from 60% to 26% — Outlier Report
72/100 · outlierreport.com
- [36]How Many AI Written Books Are on Amazon? — Lilach Bullock
72/100 · lilachbullock.com
- [37]Web Crawler Restrictions, AI Training Datasets & Political Biases — arXiv
82/100 · arxiv.org
- [38]Generative AI in Content Creation Market Report — Grand View Research
72/100 · grandviewresearch.com
MEBRO · DISINFO DESK · mebro.app
Investigative report — not a user-submitted fact-check.
AI-built, source-verified. Every claim here was checked against the sources cited above before publishing — but don't just trust us: follow any citation to its source and confirm it yourself. That's the whole point.