what AI-generated content does to the web: measured effects (authored by agents unless marked đ§)
- literature review of the harm AI-generated text, images and code suggestions have been shown to cause on the web, as of 7 Oct 2026
- about 60 sources
related notes
the humanâs notes in gen_ai, section âIssues from AI-generated textâ: Liang 2024 (peer reviews, papers), Brooks 2024 (Wikipedia), Latona 2024 (ICLR reviews), Shumailov 2024 (model collapse), Chen 2025 (X during the 2024 election), Hao 2025 (spam email), and the WIRED Substack and Medium stories
how much content is generated: generated_web_measurement
search spam and SEO: seo_search_quality
detectors: llm_text
crawler load on sites: crawling
each section separates observation, laboratory tests, and arguments
- âin the wildâ means observations of real sites, users, or traffic
- âlabâ means authors generated content and tested a system on it
models trained on generated data
- my take: every result that shows models getting worse comes from a lab loop the authors built
- this review found no deployed-model study isolating harm caused by generated web text
- the lab results also disagree on how bad it is, and the disagreement comes down to one assumption: whether old human data stays in the training set
- the first study that trains on generated text actually found on the web (Russell 2026) reports harm, but the detector vendor wrote it
Lab, shows harm:
- Shumailov 2024 is in the humanâs gen_ai notes.
- Self-Consuming Generative Models Go MAD, Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, Richard G. Baraniuk, ICLR, 2024 (the arXiv page lists no venue).
- image models retrained on their own output over generations
- âwithout enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decreaseâ
- Nepotistically Trained Generative-AI Models Collapse, Matyas Bohacek, Hany Farid, ICLR DATA-FM workshop, 2025.
- âwhen retrained on even small amounts of their own creation, these generative-AI models produce highly distorted imagesâ
- âonce affected, the models struggle to fully heal even after retraining on only real imagesâ
- The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text, Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, Chloé Clavel, NAACL Findings, 2024.
- measures variety of words, sentence shapes and meanings instead of accuracy
- âa consistent decrease in the diversity of the model outputs through successive iterationsâ
- Strong Model Collapse, Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia Kempe, ICLR, 2025 (arXiv page lists no venue).
- mostly theory on regression, checked on small language and image models
- âeven the smallest fraction of synthetic data (e.g., as little as 1% of the total training dataset) can still lead to model collapse: larger and larger training sets do not enhance performanceâ
- here âcollapseâ means more data stops helping, which is much weaker than Shumailovâs gibberish
Lab, shows the harm is avoidable:
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data, Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A. Roberts, Diyi Yang, David L. Donoho, Sanmi Koyejo, arXiv, 2024.
- earlier studies âlargely assumed that new data replace old data over time, where an arguably more realistic assumption is that data accumulate over timeâ
- âaccumulating the successive generations of synthetic data alongside the original real data avoids model collapseâ
- Collapse or Thrive? Perils and Promises of Synthetic Data in a Self-Generating World, Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, Sanmi Koyejo, NeurIPS workshops, 2024.
- adds the realistic case where data piles up but each model can only afford a fixed-size sample of it
- there, âwe observe slow and gradual rather than explosive degradation of test loss performance across generationsâ
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls, Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad, Mostafa Elhoushi, Shubhabrata Sengupta, Shang-Wen Li, Ramya Raghavendra, Ruoxi Jia, Carole-Jean Wu, EMNLP, 2025.
- one round of training, over 1000 models
- âtraining on rephrased synthetic data shows no degradation in performance in foreseeable scales whereas training on mixtures of textbook-style pure-generated synthetic data shows patterns predicted by âmodel collapseââ
- so it matters whether the generated text restates a human page or is made up from nothing
Argued:
- Position: Model Collapse Does Not Mean What You Think, Rylan Schaeffer, Joshua Kazdan, Alvan Caleb Arulandu, Sanmi Koyejo, arXiv, 2025.
- counts âeight distinct and at times conflicting definitions of model collapseâ
- âcertain predicted claims of model collapse rely on assumptions and conditions that poorly match real-world conditionsâ
Closest to the wild:
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text, Jenna Russell et al., arXiv, 2026; the prevalence half is in generated_web_measurement.
- trains on generated text collected from real crawls: âwe pretrain 800 language models, varying the ratio of added AI tokens to human tokensâ
- âFor data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly reverses into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediatelyâ
- trust: the âAIâ label comes from Pangram and its founders are authors. If the detector flags a certain kind of low-quality human page, the result would look the same
search and retrieval
- my take: that rankers prefer generated text is well shown in the lab, on benchmark collections where the authors rewrote human passages with an LLM
- What I could not find is anyone measuring this on a live search engine
- DeGenTWebâs Bing result (16.4% of how-to result sites) is prevalence in results, which is a different thing from the ranker favoring them
- Search spam in general is in seo_search_quality, including Bevendorff 2024
Lab:
- Neural Retrievers are Biased Towards LLM-Generated Content, Sunhao Dai, Yuqi Zhou, Liang Pang, Weihao Liu, Xiaolin Hu, Yong Liu, Xiao Zhang, Gang Wang, Jun Xu, KDD, 2024.
- âneural retrieval models tend to rank LLM-generated documents higher. We refer to this category of biases in neural retrievers towards the LLM-generated content as the source biasâ
- holds for the first-pass retriever and the reranker
- Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents, Haoyu Wang, Sunhao Dai, Haiyuan Zhao, Liang Pang, Xiao Zhang, Gang Wang, Zhenhua Dong, Jun Xu, Ji-Rong Wen, ICLR, 2025.
- gives the cause: the retrievers âlearn perplexity features for relevance estimation, causing source bias by ranking the documents with low perplexity higherâ. Perplexity is how surprising a language model finds the text; generated text is less surprising
- worth noticing: low perplexity is also what zero-shot detectors such as Binoculars key on, so the ranker and the detector look at the same signal with opposite signs
- Invisible Relevance Bias: Text-Image Retrieval Models Prefer AI-Generated Images, Shicheng Xu, Danyang Hou, Liang Pang, Jingcheng Deng, Jun Xu, Huawei Shen, Xueqi Cheng, SIGIR, 2024.
- same effect for image search: models âtend to rank the AI-generated images higher than the real images, even though the AI-generated images do not exhibit more visually relevant featuresâ
- Spiral of Silence: How is Large Language Model Killing Information Retrieval?, Xiaoyang Chen, Ben He, Hongyu Lin, Xianpei Han, Tianshu Wang, Boxi Cao, Le Sun, Yingfei Sun, ACL, 2024.
- a simulation: answers a chatbot writes get added back to the collection it searches, round after round
- âLLM-generated text consistently outperforming human-authored content in search rankings, thereby diminishing the presence and impact of human contributions onlineâ
- Retrieval Collapses When AI Pollutes the Web, Hongyeon Yu, Dongchan Kim, Young-Bum Kim (NAVER), WWW, 2026.
- âa 67% pool contamination led to over 80% exposure contamination, creating a homogenized yet deceptively healthy state where answer accuracy remains stable despite the reliance on synthetic sourcesâ
- with deliberately harmful pages mixed in, âbaselines like BM25 exposed ~19% of harmful content, whereas LLM-based rankers demonstrated stronger suppressionâ
- the humanâs notes already name this paper as an anchor
human knowledge sites: Stack Overflow, Wikipedia
my take: this is the best measured harm in the whole file
- Two independent studies with comparison groups found Stack Overflow lost activity right after ChatGPT, and the site has since nearly emptied
- Wikipedia is less clear: early studies found little, and the 8% drop in human views that Wikimedia reported in 2025 is a before and after number with a cause the foundation asserts but did not test
- Note that this harm comes from people asking a chatbot instead, not from generated content sitting on the web
Large language models reduce public knowledge sharing on online Q&A platforms, R. Maria del Rio-Chanona, Nadzeya Laurentsyeva, Johannes Wachs, PNAS Nexus, 2024.
- âWithin 6 months of ChatGPTâs release, activity on Stack Overflow decreased by 25% relative to its Russian and Chinese counterparts, where access to ChatGPT is limited, and to similar forums for mathematics, where ChatGPT is less capableâ
- âWe find no significant change in post quality, measured by peer feedback, and observe similar decreases in content creation by more and less experienced users alikeâ
- trust: good design. They say themselves they cannot rule out VPN use in Russia and China, which would make 25% an underestimate
The consequences of generative AI for online knowledge communities, Gordon Burtch, Dokyun Lee, Zhichen Chen, Scientific Reports, 2024.
- Stack Overflow lost about â1 million individuals per dayâ, roughly â12% of the siteâs daily web trafficâ
- âmarked declines in both website visits and question volumes at Stack Overflow, particularly around topics where ChatGPT excelsâ; Reddit developer communities showed no such drop
- disagrees with del Rio-Chanona on who left: here newer users left more
Stack Overflow today, from a PPC Land story of 8 Aug 2026 that ran a public Stack Exchange Data Explorer query: â1,442 questionsâ in July 2026 against â207,204 questionsâ in March 2014; yearly totals â1,336,266â (2022), â788,512â (2023), â398,648â (2024), â108,981â (2025).
- trust: counts only questions not deleted, and the slide started in 2014, long before ChatGPT. The raw curve alone proves nothing about cause; the two studies above do that for the first six months only
Is Stack Overflow Obsolete? An Empirical Study of the Characteristics of ChatGPT Answers to Stack Overflow Questions, Samia Kabir, David N. Udo-Imeh, Bonan Kou, Tianyi Zhang, CHI, 2024.
- why the swap matters: â52% of ChatGPT answers contain incorrect information and 77% are verboseâ, yet users âoverlooked the misinformation in the ChatGPT answers 39% of the timeâ
- 517 questions, ChatGPT of 2023; the error rate is surely lower now
Exploring the Impact of ChatGPT on Wikipedia Engagement, Neal Reeves, Wenjie Yin, Elena Simperl, ACM Collective Intelligence, 2024.
- 12 language editions; âWe find no evidence of a fall in engagement across any of the four metricsâ
- but âa lower increase in languages where ChatGPT was available than in languages where it was notâ
Wikipedia Contributions in the Wake of ChatGPT, Liang Lyu, James Siderius, Hannah Li, Daron Acemoglu, Daniel Huttenlocher, Asuman Ozdaglar, WWW, 2025.
- compares articles ChatGPT can reproduce well against ones it cannot
- ânewly created, popular articles whose content overlaps with ChatGPT 3.5 saw a greater decline in editing and viewership after the November 2022 launch of ChatGPT than dissimilar articles didâ
New User Trends on Wikipedia, Marshall Miller, Wikimedia Foundation blog, 17 Oct 2025.
- human page views fell: âa decrease of roughly 8% as compared to the same months in 2024â, found only after they reclassified traffic because âmuch of the unusually high traffic for the period of May and June was coming from bots that were built to evade detectionâ
- their explanation: âThese declines reflect the impact of generative AI and social media on how people seek informationâ. That part is a claim; the post has no comparison group
misinformation and fake news sites
- my take: generated misinformation exists and is counted, but the counts are small next to ordinary misinformation, and I found no study that measures harm to readers in the wild
- the one persuasive measured case is a single Russian-linked site that produced more after adopting an LLM
- Hanley 2024 also sits in seo_search_quality
In the wild:
- NewsGuard AI Tracking Center, NewsGuard, updated 23 June 2026.
- âNewsGuardâs team has identified 3,749 AI Content Farm news and information websitesâ in 16 languages
- found by hand, largely from leftover chatbot error messages, so it is a floor and biased toward careless operators. No traffic numbers
- Machine-Made Media: Monitoring the Mobilization of Machine-Generated Articles on Misinformation and Mainstream News Websites, Hans W. A. Hanley, Zakir Durumeric, ICWSM, 2024.
- âbetween January 1, 2022, and May 1, 2023, the relative number of synthetic news articles increased by 57.3% on mainstream websites while increasing by 474% on misinformation sitesâ
- a relative rise from a small base, with their own trained detector
- Generative propaganda: Evidence of AIâs impact from a state-backed disinformation campaign, Morgan Wack, Carl Ehrett, Darren Linvill, Patrick Warren, PNAS Nexus, 2025.
- one site with ties to Russia, before and after it started using an LLM
- âthe use of generative-AI tools facilitated the outletâs generation of larger quantities of disinformationâ and âthe AI-assisted articles maintained their persuasivenessâ
- Characterizing AI-Generated Misinformation on Social Media, Chiara Drolsbach, Emma Demirel, Nicolas Pröllochs, accepted at ICWSM 2027.
- â82,076 misleading postsâ flagged by X Community Notes
- AI-generated ones are âperceived as less believable and less harmful than conventional misinformationâ yet âsignificantly more likely to go viralâ
- only covers posts where note writers noticed the AI, mostly images
- AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild, Nicholas Dufour, Arkanath Pathak, Pouya Samangouei, et al., Christoph Bregler (Google), arXiv, 2024.
- human raters labeled images in fact checks over two years
- generated images rose fast in 2023, but ââsimpleâ methods dominated historically, particularly context manipulations, and continued to hold a majority as of the end of data collection in November 2023â
- How spammers and scammers leverage AI-generated images on Facebook for audience growth, Renée DiResta, Josh A. Goldstein, HKS Misinformation Review, 2024.
- 125 Facebook Pages that each posted 50 or more generated images; mean following 146,681; one post got 40 million views
- the motive is money and followers, not politics
Lab:
- AI âNewsâ Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian, Giovanni Puccetti, Anna Rogers, Chiara Alzetta, Felice DellâOrletta, Andrea Esuli, ACL, 2024.
- fine-tuning an old Llama âon as little as 40K Italian news articles, is sufficient for producing news-like texts that native speakers of Italian struggle to identify as syntheticâ
Argued:
- Misinformation reloaded? Fears about the impact of generative AI on misinformation are overblown, Felix M. Simon, Sacha Altay, Hugo Mercier, HKS Misinformation Review, 2023.
- âcurrent concerns about the effects of generative AI on the misinformation landscape are overblownâ
- the argument: people who read misinformation are limited by how much they want, not by how much exists, so making more changes little. The measurements above have not contradicted this so far
One live measurement I found only second hand: an Ahrefs study of 27 July 2026, as reported by Letâs Data Science (I did not open the Ahrefs post itself). About 150,000 pages from Googleâs top 10 for 100,000 searches: â5.3% of top-three pages scored as fully AI-generatedâ, and pages flagged as heavily AI ranked a bit lower and were indexed less often. The report says âThe findings do not show that Google detects and penalizes AI-written textâ. This points the opposite way from the lab results, which fits: Google ranks on much more than text matching. It is a vendor study with its own detector.
fake reviews
my take: thin evidence
- I found only detector vendor reports, each on a small or oddly chosen sample, with no false positive rate measured on reviews
- nobody has shown generated reviews change what people buy
- Fake reviews and their detection from before LLMs are a large literature that I did not cover
Pangram Amazon review study, Pangram Labs blog, dated 4 May 2026 on the page.
- â30,000 front-page product reviews across 500 of Amazonâs best-selling productsâ; â3% of the total reviews studied - 909 total reviews - were AI-generated with high confidenceâ
- â74% of AI-written reviews gave products a 5-star rating. This compares to 59% of legitimate human reviewsâ, and 93% of the AI ones carried the âVerified Purchaseâ badge
- front page of best sellers only, so not a share of all reviews
Originality.ai Amazon review study, Originality.ai blog, 17 Nov 2025.
- 2,000 reviews sampled from 26,000; says the share of reviews with â50% or more AI Contentâ grew about 400% since 2022, from a tiny base
- âVerified reviewers are roughly 1.4 times less likely to be AI generated than non-verified reviewersâ, which disagrees in spirit with Pangramâs 93%
- no false positive rate given
academic publishing and peer review
- my take: use is well measured; harm is partly measured
- the harm with the cleanest evidence is made-up references, because a reference either exists or it does not, so no detector is needed
- Claims that LLMs flood science with weak papers rest on detectors and on one Science paper whose main number has been challenged
- Liang 2024 and Latona 2024 are in the humanâs gen_ai notes; prevalence counts are in generated_web_measurement
Made-up references (measured, no detector needed):
- âFabricated citations: an audit across 2·5 million biomedical papersâ, Maxim Topaz et al., The Lancet, 7 May 2026; I read the Columbia Nursing summary, not the letter.
- PubMed Central open access papers from 1 Jan 2023 to 18 Feb 2026; among 97.1 million references, 4,046 fake ones in 2,810 papers
- about one paper in 2,828 in 2023, one in 277 in early 2026 (the second pair of numbers as reported by The Next Web)
- the link to LLMs is timing only: âsharpest increase beginning mid-2024, coinciding with the rise of AI writing toolsâ
- Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences, Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem (Microsoft), arXiv, 2026.
- checks camera-ready papers from ICLR, ICML, NeurIPS and USENIX Security with an open tool, RefChecker
- âin 2025, roughly one in twenty NeurIPS and USENIX Security papers contains at least two likely hallucinated academic-paper-like references under our strict definitionâ
- âauditing is tractable (about 0.04$ per paper in one venue-scale scan)â
- one in twenty is far above GPTZeroâs 1% below. I have not read the body to see why; âlikelyâ may be doing a lot of work
- GPTZero NeurIPS 2025 investigation, Nazar Shmatko, Alex Adam, Paul Esau, Alex Cui, Edward Tian, GPTZero, 21 Jan 2026.
- â4841 papers accepted by NeurIPS 2025â scanned; â100 confirmed hallucinations in the table below, spanning over 51 NeurIPS papersâ, each âverified by a human expertâ
- Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025, Samar Ansari, arXiv, 2026.
- sorts GPTZeroâs 100: âTotal Fabrication (66%), Partial Attribute Corruption (27%)â; â92% of contaminated papers contain 1-2 hallucinationsâ
Generated papers and reviews (measured with detectors):
- Pangramâs ICLR 2026 analysis, Pangram Labs blog, 18 Nov 2025.
- â21%, or 15,899 reviews, were fully AI-generatedâ; âover half of the reviews had some form of AI involvementâ; â9% of submissions had over 50% AI contentâ
- âthe more AI is present in a review, the higher the score isâ, same direction as Latona 2024
- vendor claims â1 in 10,000â false positives; no independent check on reviews
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary, Dmitry Kobak, Rita GonzĂĄlez-MĂĄrquez, EmĆke-Ăgnes HorvĂĄt, Jan Lause, Science Advances, 2025.
- no detector: counts words that suddenly got more common in 15 million PubMed abstracts
- âat least 13.5% of 2024 abstracts were processed with LLMsâ, âreaching 40% for some subcorporaâ
- shows use, not harm
- Scientific production in the era of Large Language Models, Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin, Science, 2025.
- âscientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3%â
- âLLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelmingâ
- Comment on Scientific production in the era of large language models, Thomas Renault, Antonin Bergeaud, Clément Bosquet, arXiv, 2026.
- says the production number is an artifact: an author counts as an adopter from the first month one of their abstracts gets flagged, and âdetected-adoption months are disproportionately high-output monthsâ
- three placebo tests, including âa pre-ChatGPT observation windowâ, âeach produce a similarly positive post-treatment patternâ
- a good warning for any study that dates âadoptionâ by first detection
- GPT-fabricated scientific papers on Google Scholar, Jutta Haider, Kristofer Rolf Söderström, Björn Ekström, Malte Rödl, HKS Misinformation Review, 2024.
- found 139 papers by searching Google Scholar for leftover chatbot phrases; 57% on âpolicy-relevant subjects (i.e., environment, health, computing)â
- a floor from careless authors, like NewsGuardâs method
- Explosion of formulaic research articles, including inappropriate study designs and false discoveries, based on the NHANES US national health database, Tulsi Suchak, Anietie E. Aliu, Charlie Harrison, Reyer Zwiggelaar, Nophar Geifman, Matt Spick, PLOS Biology, 2025.
- papers of one template on one public dataset went from about 4 a year (2014 to 2021) to 190 in 2024
- the paper infers AI and paper mills from the timing and the sameness; it does not detect generated text
What venues did about it (shows the cost was real to them):
- arXiv CS stopped taking review and position papers without prior peer review on 31 Oct 2025. From 404 Media: âWe now receive hundreds of review articles every monthâ, and âGenerative AI / large language models have added to this flood by making papersâespecially papers not introducing new research resultsâfast and easy to writeâ.
- In May 2026 arXivâs CS chair Thomas Dietterich announced a one-year ban for âincontrovertible evidenceâ of unchecked generated content such as references that do not exist, per The Next Web.
made-up package names, court citations, bug reports
my take: models make up names at a measured and still nonzero rate, and the downstream damage is counted in courts and in open source bug trackers
- For packages, the attack is shown to be possible, but I found no measured case of an attacker registering a made-up name and getting installs
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs, Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath, Murtuza Jadliwala, USENIX Security, 2025.
- 576,000 generated code samples, 16 models
- âthe average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models, including a staggering 205,474 unique examples of hallucinated package namesâ
The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort, Aleksandr Churilov, arXiv, 2026.
- reruns Spracklen on five 2025 to 2026 models: âoverall hallucination rates between 4.62% (Claude Haiku 4.5) and 6.10% (GPT-5.4-mini)â
- â127 package names (109 on PyPI, 18 on npm) that all five evaluated models invent identicallyâ; after telling the registries, 53 âremain registrable by an attackerâ
- per CSO Online, he found âno evidence that any of the remaining 53 names have been registered maliciously, nor used in an attackâ
- single independent author, not peer reviewed
Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Matthew Dahl, Varun Magesh, Mirac Suzgun, Daniel E. Ho, Journal of Legal Analysis, 2024.
- lab: models asked checkable questions about random federal cases are wrong âbetween 58% of the time with ChatGPT 4 and 88% with Llama 2â
AI Hallucination Cases database, Damien Charlotin, ongoing.
- in the wild: court decisions where a court âexplicitly found (or implied) that a party relied on hallucinated contentâ; 2,149 cases as of 5 Oct 2026
- a count of people who got caught
The end of the curl bug bounty, Daniel Stenberg, blog, 26 Jan 2026.
- share of reports that were real bugs used to be ânorth of 15% of the submissions ending up confirmed vulnerabilitiesâ and in 2025 âplummeted to below 5%â; he blames an âexplosion in AI slop reportsâ
- one project, the maintainerâs own count, but it is a direct cost: they ended the bounty
On Autopilot? An Empirical Study of Human-AI Teaming and Review Practices in Open Source, Haoyu Gao, Peerachai Banyongrakkul, Hao Guan, Mansooreh Zahedi, Christoph Treude, MSR, 2026.
- âover 67.5% of AI-co-authored PRs originate from contributors without prior code ownershipâ, and from such contributors âapproximately 80% merged without any explicit reviewâ
- so in this dataset the problem is too little checking, not maintainers drowning
writing style and language
my take: LLM word habits have clearly entered human writing and even speech
- That is measured
- Whether this is harm is opinion
- the stronger harm claim, that people who write with an LLM end up sounding and thinking alike, is shown only in small lab studies
Empirical evidence of Large Language Modelâs influence on human spoken communication, Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio de la Serna, Lara Kirfel, Prateek Gupta, Ivan Soraperra, Thomas F. Eisenmann, Dirk U. Wulff, Iyad Rahwan, arXiv, 2024 (revised 2026).
- âwords preferentially generated by ChatGPT, such as delve, showcase, boast, intricacies and meticulous, increased abruptly in spontaneous human speechâ, in â737,083 hours of conversation from 824,634 podcast episodes, screened for unscripted speechâ
- plus an experiment with 496 people: âa brief chatbot interaction led participants to adopt its words as their ownâ
- matters for detectors: human text is drifting toward what detectors call AI. The humanâs notes already list âhuman may learn word from AIâ as a criticism of Liang 2024
Why Does ChatGPT âDelveâ So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models, Tom S. Juzek, Zina B. Ward, COLING, 2025.
- â21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usageâ; could not pin down why models overuse them
Does Writing with Language Models Reduce Content Diversity?, Vishakh Padmakumar, He He, ICLR, 2024.
- lab: essays written with InstructGPT are more alike; âthe user-contributed text remains unaffectedâ, so the sameness comes from the modelâs own sentences
Generative artificial intelligence enhances creativity but reduces the diversity of novel content, Anil R. Doshi, Oliver P. Hauser, Science Advances, 2024 (I read the arXiv page).
- lab: stories written with AI ideas rated better, but âmore similar to each other than stories by humans aloneâ
AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances, Dhruv Agarwal, Mor Naaman, Aditya Vashistha, CHI, 2025.
- lab, 118 people: âAI suggestions led Indian participants to adopt Western writing stylesâ
money: publishers, freelancers, creators
- my take: the money harm to publishers is real and now has causal evidence, but it comes from AI answers shown in place of links, not from generated pages competing with human ones
- I found no study that measures human sites losing readers or ad money to generated sites
- That is the gap closest to DeGenTWeb
AI answers taking clicks:
- Google users are less likely to click on links when an AI summary appears in the results, Athena Chapekis, Anna Lieb, Pew Research Center, 22 July 2025.
- browsing data of 900 US adults, March 2025, 68,879 searches
- users clicked a result link on 8% of visits with an AI summary and 15% without; they clicked a source inside the summary on âjust 1% of all visitsâ
- observational: searches that trigger a summary differ from those that do not
- A randomized experiment by Saharsh Agarwal and Ananya Sen, SSRN, 2026, which I read only through PPC Land: a browser extension hid AI Overviews for a random half of 1,065 US desktop Chrome users. Outbound clicks per search were 0.37 with overviews and 0.62 without; searches with no click were 73% against 54%.
- this is the causal number, and it is bigger than Pewâs
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia, Mehrzad Khosravi, Hema Yoganarasimhan, arXiv, 2026.
- compares the same articles across languages as AI Overviews rolled out country by country
- âdefault AIO availability reduced English search traffic by 5.45% and 4.82%â against German and French
- much smaller than the 8% and the click numbers above; Wikipedia gets traffic from many places besides Google
- Chartbeat data in the Reuters Instituteâs 2026 trends report, via Press Gazette (Charlotte Tobitt, 12 Jan 2026): Google search referrals to âmore than 2,500 publisher websitesâ âdeclined globally by a third in the year to Novemberâ; ChatGPT referrals are âjust 0.02% of total referral trafficâ.
- a trend; the report blames AI summaries but does not isolate them
- The crawl-to-click gap: Cloudflare data on AI bots, training, and referrals, David Belson, Sam Rhea, Cloudflare blog, 1 July 2025.
- pages crawled per visitor sent back; Anthropic 70,900 to 1 in the week of 19 to 26 June 2025
- more on crawler load in crawling
Work and creators:
- âThe Short-Term Effects of Generative Artificial Intelligence on Employment: Evidence from an Online Labor Marketâ, Xiang Hui, Oren Reshef, Luofeng Zhou, Organization Science, 2024; I read the WashU summary.
- Upwork writers after ChatGPT: monthly jobs down 2%, earnings down 5.2%; image freelancers after DALL-E and Midjourney: jobs down 3.7%, income down 9.4%; better-paid freelancers lost more
- Deezer newsroom, 21 July 2026.
- â90,000 AI-generated tracks per day now represent over 50% of all new music uploadsâ, yet only 1 to 3% of listening, and âup to 85% of the streams generated by fully AI-generated tracks were in fact fraudulent in 2025â
- the clearest picture I found of supply against demand: a flood of uploads, almost nobody listening, and most of the listening faked to collect royalties. It backs the Medium CEOâs claim in the humanâs notes that generated posts were hardly read, and Simon 2023âs argument above. The platformâs own detector and its own numbers
what the evidence adds up to
- Solid, with comparison groups: people moved from Stack Overflow to chatbots; AI answers in search cut clicks to sites; freelance writers and illustrators lost some work.
- Solid, by direct count: made-up references in published papers and court filings; models inventing package names at around 5%; curlâs bug reports.
- Shown in the lab only: models getting worse from generated training data; rankers preferring generated text; people writing more alike.
- Counted but with no measured harm: generated news sites, generated misinformation, generated reviews, generated music uploads.
- Mostly argued: that generated pages crowd human pages out of search and ad money; that the webâs training value is falling.
Two things stand out to me. First, the best measured harms come from people using chatbots instead of visiting sites, not from generated content on the web. Second, where supply and demand were both measured (Deezer, Medium, Community Notes), generated content is a large share of what gets uploaded and a small share of what gets seen. A count of generated sites, which is what DeGenTWeb gives, says little about harm until it is joined with who visits them.
research we could do
- Weigh DeGenTWebâs site labels by traffic. Do LLM-dominant sites get visits, search impressions and ads, or do they sit unread like Deezerâs uploads? Use Tranco or CrUX rank, Common Crawlâs host link graph, ads.txt and ad tags on the page. Builds on DeGenTWeb, Deezerâs upload against stream numbers, Simon 2023, NewsGuard (which has no traffic data).
- Test source bias on a live engine. The lab papers (Dai 2024, Wang 2025) rewrote passages; nobody ran the matched test on real results. For queries where both LLM-dominant and human sites answer, compare their rank after controlling for site age and inbound links. DeGenTWeb already has Bing results labeled per site. Builds on Dai 2024, Wang 2025, Yu 2026, the Ahrefs 2026 correlation.
- Do answer engines cite generated sites? Send the DeGenTWeb how-to queries to AI Overviews, ChatGPT search and Perplexity, collect cited URLs, and label the sites. This is Retrieval Collapse (Yu 2026) measured in the wild instead of simulated, and Spiral of Silence (Chen 2024) with real data.
- Did human sites that compete with generated sites lose more? For topics where LLM-dominant sites appeared early, check whether human sites on those topics lost search visibility or stopped posting sooner than human sites on untouched topics. Same design idea as Lyu 2025 (similar against dissimilar Wikipedia articles) and del Rio-Chanona 2024. This is the missing study in the money section. It needs a traffic or posting-rate signal; posting rate can come from Common Crawl and the Internet Archive.
- Site-level training value. Russell 2026 labels tokens with a paid detector its authors sell. Redo a small version with DeGenTWebâs site labels: train small models on text from LLM-dominant sites against matched human sites. Split generated sites by whether they restate human pages or make things up, because Kang 2025 says that split decides harm. An independent check on a result people will quote a lot.
- Reference checking on the web. Made-up references are the one harm that needs no detector. Run a RefChecker-style check (Russinovich 2026) on outbound links and cited sources of web pages: do LLM-dominant sites cite pages, papers or packages that do not exist more often? If yes, it is both a harm measure and a detector-free signal to validate DeGenTWebâs labels.
- Made-up package names in the wild. Churilov 2026 found no attack on his 127 names. Scan crawled tutorial pages and public repositories for install commands naming packages that do not exist, and watch the registries for who registers them. Builds on Spracklen 2025, Churilov 2026. This fits a web measurement group better than a model evaluation.
- Where do the readers of dead Q&A sitesâ topics go? Stack Overflow fell 99%. Check whether generated how-to sites now rank for the queries Stack Overflow used to answer. Joins the knowledge site section with DeGenTWebâs how-to query set.
- I would start with 1 and 2
- both reuse data DeGenTWeb already has, and both answer the question a reviewer will ask of a prevalence paper: so what?
gaps in this review
- I read abstracts and summary pages for most papers, not full texts. The Lancet letter, the Agarwal and Sen experiment, Hui 2024 and the Ahrefs study were read only through secondary pages, as marked.
- Not covered: generated images and video beyond misinformation counts, deepfake fraud, education and student cheating, generated books on Amazon, social media bots (the humanâs notes list them), and the fake review literature from before LLMs.
consultation
- see the provenance index for the current Extra High consultation
Last edited: