measuring how much of the web is AI-generated (authored by agents unless marked đ§)
- targeted review of generated-content prevalence studies
- compares their populations, units, definitions, and detector assumptions
- reviewed 7 Oct 2026
- compares against the newer local DeGenTWeb draft
- public preprint
- Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha
- the local draft contains later measurements
- consultation status is in the provenance index
Sibling notes: image_watermarking, labeling_rules_and_practice; detectors themselves are covered under llm_text, search and SEO under web_user, crawling under web_infra.
what DeGenTWeb measures (from the draft)
source:
DeGenTWeb_writeup2/sections/and macros indegentweb_imc2026.texunit: the website (subdomain), labeled âLLM-dominantâ when most of its prose pages look LLM-written; not per page, not per token
method: Trafilatura extraction, Dolma heuristic filter plus 200-token minimum plus Rabin dedup, Binoculars (Falcon-7B pair) per page, 9 deciles of 20 pages per site into a linear SVM trained on a 120-site baseline (60 Russell 2000 and IndieWeb sites vs 60 Wix/B12 AI-builder mirrors)
numbers in the draft: 6.0% of 94,908 Common Crawl sites (2020 to May 2025) LLM-dominant, rising from 2.59% for sites first seen in 2022H2 to 28.6% for 2025H1; 16.4% of 38,309 Bing how-to result sites, 18.2% of result links; 0.29% of pre-ChatGPT Common Crawl subdomains flagged (historical positive rate); 98.7% baseline site accuracy; accuracy drops to 0.58 on sites generated by 2025 frontier models (Pearson -0.88 with the Artificial Analysis index)
the draftâs related-work section cites only Liang 2024 (peer reviews and papers), Latona 2024 (ICLR), Hao 2025 (spam email), Brooks 2024 (Wikipedia) and NewsGuard; the sections below cover the much larger body of prevalence studies the draft does not yet cite
open web and web crawls
- closest reviewed studies sample Common Crawl or the Internet Archive
- their page-level decisions differ from DeGenTWebâs website labels
- historical positive rates and known-authorship error rates must be distinguished
Pew Research Center, âHow Much of the Internet Is Written With AI?â (2026)
Main report and methodology, Pew Research Center Data Labs, August 2026.
- what was measured: share of pages per Common Crawl crawl that show âmeaningfulâ signs of AI authorship
- data: â10,000 English-language pages from each of the 49 crawls created between January 2021 and July 2026â, 490,000 pages in all
- method: Pangramâs open model
editlens_Llama-3.2-3B, score 0 to 1, pages scoring 0.2 or higher count as AI; they checked it against the commercial Pangram 3.3 on 62,370 pages: âThey agree in 96% of cases, with a Cohenâs kappa value of 0.61â - number: â35% of pages in the July 2026 crawl with post-ChatGPT publication dates show signs of AI authorshipâ; over all sampled pages regardless of date, â10% show significant signs of AI authorshipâ, about one in ten .com pages, half that on .org, about 1% on .edu and .gov
- trust: the 35% is over the 10 to 15% of pages that carry a publication date in their HTML, which Pew says âreflect AI authorship among dated content rather than the web as a wholeâ; kappa 0.61 between the open and paid model is only moderate agreement; no false positive rate on known human web pages is reported, and Pew itself says âReported shares from 2021 and 2022 crawls should be treated as approximateâ because early crawls show more false positives
- relation to DeGenTWeb: same data source, page level instead of site level, one detector instead of a zero-shot score plus aggregation; their 35% of dated pages and DeGenTWebâs 28.6% of sites first seen in 2025H1 are in the same ballpark despite different units, which is worth saying in the paper
dolezal, Alam, Graham, Bohacek, âThe Impact of AI-Generated Text on the Internetâ (2026)
arXiv 2604.26965, Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek (Internet Archive and Stanford), arXiv, 2026.
- what was measured: share of newly archived URLs per month that are AI-generated or AI-assisted, plus effects on diversity and sentiment
- data: Internet Archive CDX index, â33 monthly intervals from August 2022 to May 2025â, target 10,000 URLs per month, stratified by first-archival time, MIME type, URL depth and TLD, one URL per host
- method: Pangram v3 API, three-way output (fully AI, AI-assisted, human); the headline sums AI and AI-assisted; they compared Pangram against Binoculars, DivEye and Desklib on synthetic tests of length, HTML wrapping, model family and language and say Pangram had âperfect accuracy on texts exceeding 50 wordsâ
- number: âBy the first half of 2025, as much as 35% of websites uploaded to the internet in a given month were AI-generated or AI-assistedâ
- trust: âAI-assistedâ is a wide bucket, so 35% is an upper-leaning number; the detector comparison is on their own synthetic texts, not on real web pages with known authorship; one URL per host means a site with a thousand pages counts the same as a site with one, which is actually closer to DeGenTWebâs site unit than Pewâs page unit; the secondary claims about diversity and sentiment lean on correlations across months
- relation to DeGenTWeb: the nearest competitor in spirit (new sites over time, Internet Archive). DeGenTWebâs 28.6% for 2025H1 sites and their 35% for mid-2025 URLs would look consistent if one subtracts their AI-assisted share. The draft should cite it and point out that DeGenTWebâs LLM-dominant label is stricter than âAI-assistedâ
russell et al., âHow Much Is an AI Token Worth?â (2026)
arXiv 2609.40295, Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi (UMass Amherst and Pangram), arXiv, 2026.
- what was measured: share of tokens in 2026 web crawls labeled AI, and what training on them does to language models
- data: June and August 2026 web crawls after FineWeb quality filtering; the âWildAIâ corpus of 83B tokens with AI, topic and format labels
- method: Pangram detector (the authors include Pangramâs founders)
- number: â27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by Augustâ
- trust: a token share after FineWeb filtering, so boilerplate and non-prose pages are already removed, which inflates the share relative to raw crawls; the detector vendor is a coauthor; no independent ground truth
- relation to DeGenTWeb: shows the âLLM content hurts trainingâ angle that DeGenTWeb mentions only via Shumailov 2024; also a third unit (tokens) to compare with pages (Pew) and sites (DeGenTWeb)
graphite, âMore Articles Are Now Created by AI Than Humansâ (2025) and the Q1 2026 update
Graphite report, Graphite (an SEO agency), October 2025; the follow-up was covered by Axios, which I could not open directly.
- what was measured: share of English articles on the web that are mostly AI-written
- data: â43,000 English-language URLs randomly selectedâ from Common Crawl, published January 2020 to May 2025 (press coverage says 65,000; the page says 43,000)
- method: Surferâs AI detector on 500-word chunks; an article counts as AI âif the algorithm predicts that more than 50% of the content is AI-generatedâ; they report a 4.2% false positive rate on articles from November 2020 to November 2022 and 0.6% false negatives on 6,009 GPT-4o articles
- number: AI articles passed human ones in November 2024; about 39% a year after ChatGPT; the update with three detectors (Pangram, Copyleaks, GPTZero) averaged over about 55,000 articles found 49.9% in Q1 2026
- trust: the 4.2% FPR is measured on their own sample and only against GPT-4o for false negatives; âarticleâ filtering is not described in detail; the authors sell SEO services; the âmore than halfâ headline got repeated by many outlets without the sampling caveats
- relation to DeGenTWeb: this is the main âcried wolfâ headline the intro should name. Their 2020 share of 2.2% is in effect their false positive rate, which matches what DeGenTWeb does with pre-ChatGPT subdomains (0.29%) but at a far worse level, so DeGenTWebâs number is much more defensible
ahrefs, â74% of New Webpages Include AI Contentâ (2025)
Ahrefs blog, Ahrefs, May 2025.
- data: 900,000 newly created English pages seen in April 2025, one per domain
- method: Ahrefsâ own detector, codenamed âbot_or_notâ; no accuracy or false positive numbers given
- number: â74.2% of them contained AI-generated contentâ; â2.5% of pages were categorized as âpure AIââ, 25.8% pure human, 71.7% mixed
- trust: low. The 74% counts any page with any flagged span, and the detector is unvalidated in public; the 2.5% âpure AIâ is the comparable number to DeGenTWebâs LLM-dominant share, and it is close to DeGenTWebâs 6% average over 2020 to 2025 given the different units
- relation to DeGenTWeb: a clear example of the âclaims vary widelyâ point; the gap between 74% and 2.5% within the same study shows why the LLM-dominant definition matters
thompson et al., âA Shocking Amount of the Web is Machine Translatedâ (2024)
arXiv 2401.05749, Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, Marcello Federico, Findings of ACL, 2024.
- what was measured: how much of multilingual web text is machine translated, the pre-LLM form of machine-generated web text
- data: 6.4B sentences in 90 languages from Common Crawl via the ccAligned pipeline, grouped into multi-way parallel sets
- method: no detector; they reason that content translated into many languages at once is low quality and thus machine made, and confirm with human quality ratings
- number: â57.1% of the sentences in our corpus are multi-way parallel in at least 3 languagesâ (from the abstract)
- trust: the inference from multi-way parallelism to machine translation is indirect; it says nothing about monolingual English content
- relation to DeGenTWeb: precedent for âmachine content dominates the webâ claims before LLMs; also a reminder that DeGenTWebâs English-only Binoculars scoring misses the non-English web, where generated content may be more common
search results
originality.ai, âAmount of AI Content in Google Search Resultsâ (ongoing, 2019 to 2025)
Originality.ai study, Originality.ai (a detector vendor), updated through September 2025.
- data: 500 informational keywords, top 20 Google results each, about 10,000 pages per two-month period, from January 2019 onward; âText from websites that were not article-based, like YouTube and Reddit, was removedâ
- method: Originality.aiâs own detector at 0.5 confidence
- numbers: 2.27% in February 2019, 8.48% December 2023, 7.43% right after the March 2024 core update, 19.56% July 2025, 17.31% September 2025
- trust: the 2019 value of 2.27% is a false positive floor the study never subtracts; no validation on web pages; the vendor sells the detector
- relation to DeGenTWeb: the direct comparison for the Bing how-to result, where DeGenTWeb finds 16.4% of sites and 18.2% of links LLM-dominant. The numbers agree surprisingly well. DeGenTWeb adds a transparent method, a false positive bound, and the search vs open web gap (16.4% vs 6.0%)
graphite, âHow Does AI-Generated Content Perform in Search and Answer Engines?â (2025)
Graphite report, Graphite, 2025. Not read in full; the companion to the article study above. They claim AI articles âlargely do not appear in Google and ChatGPTâ, that is, AI content is common in the crawl but rare among top results. This contradicts DeGenTWebâs finding that Bing does not down-rank LLM-dominant sites, and the Originality.ai trend. Worth a direct look before the paper claims either way; Graphiteâs search sample is drawn from their SEO clientsâ queries, not random how-to questions.
deGenTWebâs own Webis-PSERP-24 analysis (hidden section of the draft)
The draft has a hidden subsection on Webis-PSERP-24, top-20 results for 7,392 product-review queries from Startpage, Bing, DuckDuckGo and ChatNoir from 2022 to 2024, with the share of Startpage result links to LLM-dominant sites rising over time. No other study I found measures LLM content in a multi-engine, multi-year SERP archive, so that section is more novel than the draft treats it.
news sites
hanley and Durumeric, âMachine-Made Mediaâ (2024)
arXiv 2305.09820, Hans W. A. Hanley, Zakir Durumeric, ICWSM, 2024. PDF now in the paper collection.
- what was measured: share of news articles that are machine-generated, on mainstream vs misinformation sites, January 2022 to May 2023
- data: âover 15.46 million articles from 3,074 misinformation and mainstream news websitesâ, scraped with Selenium, text and dates from newspaper3k and htmldate
- method: a DeBERTa classifier trained on human news plus GPT-2/GPT-3.5 generations, with paraphrase and perturbation augmentation; the threshold was raised to 0.98 âallowing us to achieve a 1% FPR/accuracy on the Signal article datasetâ, at which point âour model reaches a precision of 0.989 on our ChatGPT rewrite test set at the expense of only reaching a 0.639 recallâ
- numbers: absolute levels are small: â1.07% of all articles published in January 2022 (12,984 of 1,213,983 articles) were synthetically generated. However, by May 2023, the fraction of synthetic articles nearly went up to 1.78%â; relative increases of 57.3% (mainstream) and 474% (misinformation); the least popular sites (rank beyond 10M in their popularity ranking) rose 3.42 absolute points, the most popular under 1 point
- trust: the best-validated news study. They explicitly accept low recall for a 1% FPR, and report absolute shares, not just the headline relative increases. Pre-ChatGPT âsyntheticâ articles are mostly templated finance news, which the detector also flags. Period ends May 2023, so the numbers are early
- relation to DeGenTWeb: the closest methodological cousin. Same low-FPR philosophy, same âsmall sites drive the riseâ finding (DeGenTWebâs LLM-dominant sites are entry-level ad stack, low engagement). DeGenTWeb should cite it as the news-only precedent and contrast its site-level verdict with their per-article one. Their 1.78% in May 2023 vs DeGenTWebâs 6% site share through 2025 are not in conflict given the time gap
russell et al., âAI use in American newspapers is widespread, uneven, and rarely disclosedâ (2026)
arXiv 2510.18774, Jenna Russell, Marzena Karpinska, Destiny Akinode, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer, ACL, 2026. Already in the paper collection.
- data: â186K articles from online editions of 1.5K American newspapers published in the summer of 2025â, plus 45K opinion pieces from the Washington Post, New York Times and Wall Street Journal
- method: Pangram v2, which scores segments and then gives an overall label
- numbers: âapproximately 9% of newly-published articles are either partially or fully AI-generatedâ; 1.7% at large papers vs 9.3% at small local ones; opinion pieces â6.4 times more likely to contain AI-generated content than news articlesâ; of 100 flagged articles, only five disclosed AI use
- trust: large, well-sampled, and the manual audit of flagged articles is a strength; but the detector is the coauthorsâ product and no FPR is measured on this corpus; âpartiallyâ AI is a soft label
- relation to DeGenTWeb: shows the per-outlet view (newspapers are sites too). DeGenTWebâs search-result site categories (publisher/editorial with ads) overlap with small local papers. A useful contrast: DeGenTWeb asks which whole sites are LLM-run, this asks how much AI slips into human-run outlets
ansari, Zhang, Tripto, Lee, âEchoes of Automationâ (2025)
arXiv 2508.06445, Abolfazl Ansari, Delvin Ce Zhang, Nafis Irtiza Tripto, Dongwon Lee, SBP-BRiMS, 2025. PDF added.
- data: âover 40,000 news articles from major, local, and college news mediaâ
- method: three detectors, Binoculars, Fast-DetectGPT and GPTZero, plus sentence-level analysis
- numbers: the abstract gives no headline share, only âsubstantial increase of GenAI use in recent years, especially in local and college newsâ, and that LLM text shows up in introductions more than conclusions
- trust: three detectors is good practice, but without a reported FPR or absolute share the finding is directional
- relation to DeGenTWeb: uses Binoculars too, so its per-sentence findings hint at what DeGenTWebâs page scores respond to in mixed articles
newsGuard AI Tracking Center (ongoing)
NewsGuard AI Tracking Center, NewsGuard, 2023 to 2026. The draft already cites this.
- what: a hand-curated count of âunreliable AI-generated newsâ sites, with press coverage saying 3,749 sites in 16 languages by June 2026, up from 1,254 at the time the draftâs citation was written
- trust: a count of found sites, not a share of anything; detection is editorial judgment, criteria not reproducible; the number only ever grows
- relation to DeGenTWeb: the draftâs intro uses it for motivation, which is right; it cannot be used as a prevalence estimate
wikipedia
brooks, Eggert, Peskoff, âThe Rise of AI-Generated Content in Wikipediaâ (2024)
arXiv 2410.08044, Creston Brooks, Samuel Eggert, Denis Peskoff, WikiNLP workshop at EMNLP, 2024. Already in the paper collection and in the draftâs bibliography; the humanâs gen_ai.md notes already criticize it (âunscientific bc assume paper i.i.d.â).
- data: recently created articles in English, German, French and Italian; the Signpost coverage says 2,909 English articles created in August 2024
- method: GPTZero and Binoculars (Falcon-7B), âthresholds calibrated to achieve a 1% false positive rate on pre-GPT-3.5 articlesâ, then subtract the pre-GPT positive rate to get a lower bound
- number: âDetectors flag over 5% of newly created English Wikipedia articles as AI-generatedâ, lower in the other languages; 4.36% after subtracting the baseline
- trust: the FPR calibration on pre-GPT articles is the same trick DeGenTWeb uses with pre-ChatGPT Common Crawl; sample is small; Binoculars has higher false negatives on non-English, so the cross-language comparison is weak
- relation to DeGenTWeb: cited already; worth noting that no newer Wikipedia prevalence study turned up in my searches, only a 2025 Wikimedia Research Fund proposal, so the field is open for a repeat with a 2026 sample
social media and forums
sun et al., âAre We in the AI-Generated Text World Already?â (2025)
arXiv 2412.18148, Zhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang, Yule Liu, Xinlei He, Michael Backes, Yang Zhang, ACL, 2025. PDF added.
- data: Medium 1,170,821 posts, Quora 245,131 answers (January 2022 to October 2024), Reddit 982,440 comments (to July 2024)
- method: their own detector OSM-Det, a Longformer fine-tuned on AIGTBench (28.77M AI and 13.55M human samples from 12 LLMs); accuracy 0.979; FPR measured on pre-2022 posts: Medium 1.82%, Quora 1.36%, Reddit 1.70%
- numbers: AI attribution rate on Medium ârising from 1.77% to 37.03%â, Quora â2.06% to 38.95%â, Reddit â1.31% to 2.45%â
- trust: FPR on pre-LLM posts is measured, which is good, but the ~1.5% FPR is the whole Reddit signal, so the Reddit number means nothing; the detector was trained on 12 LLMs mostly GPT and Llama; âAI attributionâ includes polished human text
- relation to DeGenTWeb: Medium and Quora at ~37% mirror DeGenTWebâs rise in new sites to 28.6%. The platform gap (Medium vs Reddit) parallels the DeGenTWeb search vs open web gap: where there is money or reach, there is more LLM text
la Cava, Aiello, Tagarelli, âMachines in the Crowdâ (2025)
arXiv 2510.07226, Lucio La Cava, Luca Maria Aiello, Andrea Tagarelli, arXiv, 2025. PDF added.
- data: â38,074,021 comments and 4,073,586 submissionsâ from 51 subreddits, January 2022 to December 2024
- method: Fast-DetectGPT at a very strict threshold (0.99) and only texts of at least 250 tokens
- numbers: machine text âmarginally present on Redditâ, with peaks of 6.33% in r/askscience (July 2023), 7.69% in r/malefashionadvice, 8.46% in r/teenagers (February 2023); about 2% of active users produce it
- trust: the authors call their estimate âvery conservativeâ; the strict threshold means the numbers are floors; no FPR reported for the threshold on pre-LLM Reddit text in the abstract
- relation to DeGenTWeb: an example of a per-user aggregation idea (the humanâs gen_ai.md lists âset-level detection applications: Reddit usersâ) done loosely; DeGenTWebâs decile-SVM could be applied per user
chen, Ye, Ferrara, Luceri, âPrevalence, Sharing Patterns, and Spreaders of Multimodal AI-Generated Content on Xâ (2025)
arXiv 2502.11248, Zhiyi Chen, Jinyi Ye, Emilio Ferrara, Luca Luceri, arXiv, 2025. Already in the paper collection and in the humanâs gen_ai.md notes.
- data: 2.5 million images from X posts about the 2024 US election, July 1 to September 30, 2024
- method: GPT-4o as the image classifier, F1 0.96 on a 2,400-image test set; on 2,000 manually labeled real-world images, âfalse positive and false negative rates of 6.5% and 1.6%â
- numbers: âapproximately 12.33% of images related to the 2024 U.S. Election on X are AI-generatedâ; 10% of spreaders account for 80% of AI images; the earlier version also reports 1.4% AI text by a fine-tuned RoBERTa
- trust: a 6.5% FPR on a 12% positive rate means roughly half the flagged images could be false positives; the text number is within detector error, as the humanâs notes say
- relation to DeGenTWeb: the âsuperspreaderâ concentration matches the DeGenTWeb mass-produced cluster finding (a few operators make much of the content)
pangram, âAI in Your Feedâ (2026)
Pangram blog, Pangram Labs, 2026; covered by The Decoder.
- data: 1,002,627 posts over 50 words scanned by users of Pangramâs Chrome extension between April and June 2026, on LinkedIn, Medium, Substack, X and Reddit
- method: Pangram 3.3, claimed â0.01% false positive rateâ
- numbers: fully AI longform: LinkedIn over 40%, Medium 31%, X 23.9% (plus 22.9% AI-assisted), Reddit 13%, Substack 10%
- trust: the sample is whatever extension users chose to scan, so it is biased toward suspicious posts; the 0.01% FPR claim is the vendorâs on its own benchmark; still the largest multi-platform sample
- relation to DeGenTWeb: another vendor headline the intro can group with Graphite and Ahrefs
originality.ai, âSocial Media AI Trackerâ (monthly, 2026)
Originality.ai August 2026 tracker, Originality.ai, 2026.
- data: a few hundred random longform English posts per platform per month (LinkedIn 119, Threads 162, Facebook 166, X 174, Reddit 564 in August 2026)
- method: a post is âLikely AIâ when the detector is âmore confident than not (50% confidence or higher) that at least 15% was written by AIâ
- numbers: LinkedIn 76%, Threads 66%, Facebook 59%, X 57%, Reddit 31%
- trust: very low. Tiny samples, a definition that counts a post with 15% AI as AI, and an unvalidated vendor detector. The Reddit 31% vs Pangramâs 13% vs La Cavaâs under 9% peaks vs Sunâs 2.45% shows how much the definition and detector drive the number
- relation to DeGenTWeb: the best single illustration of âclaims vary widelyâ for the intro: four Reddit estimates spanning 2% to 31%
matatov, Aubin Le QuĂ©rĂ©, Amir, Naaman, âAI-Generated Media in Art Subredditsâ (2024)
arXiv 2410.07302, Hana Matatov, Marianne Aubin Le Quéré, Ofra Amir, Mor Naaman, arXiv, 2024.
- what: self-declared AI posts and AI accusations in art subreddits through 2023; no detector
- number: AI posts and accusations âaccounting for fewer than 0.5% of the image-based postsâ
- trust: self-labels are a floor; relevant mostly as a norm study
kapwing, âYouTube AI slopâ report (2026)
Kapwing report via NeoMam, Kapwing, 2026 (press coverage in LBC, SlashGear and others).
- method: hand-checked the top 100 trending channels per country on playboard.co (15,000 channels), found 278 âAI slopâ channels; a fresh accountâs first 500 recommended Shorts had 104 AI slop videos (21%)
- trust: hand judgment of âslopâ, a single new account, a marketing company; directional only
- relation to DeGenTWeb: video is outside DeGenTWebâs scope, but the âmass produced for ad revenueâ story is the same
product reviews
No academic study with a defensible in-the-wild prevalence number exists for reviews. All numbers below come from detector vendors.
pangram, âThree percent of front-page Amazon reviews are now AI-generatedâ (2026)
Pangram blog, Pangram Labs, May 2026.
- data: 30,000 front-page reviews on 500 best-selling products in ten categories
- numbers: â3% of the total reviews studied - 909 total reviews - were AI-generated with high confidenceâ; 74% of AI reviews are 5-star vs 59% of human; 93% carry âVerified Purchaseâ
- trust: front-page reviews only; âhigh confidenceâ threshold means a floor; no FPR measured on pre-ChatGPT reviews
originality.ai, Google reviews and TripAdvisor studies (2025)
Google reviews and TripAdvisor, Originality.ai, 2025.
- data: up to 5 Google reviews per place in 15 North American cities across 20 categories; 10,000 TripAdvisor reviews in 20 cities
- numbers: Google â5.01%â in 2019 to â19%â in 2024; TripAdvisor âfrom 4.49% in 2019 to 10.7% in 2024â
- trust: the 2019 values (5.01% and 4.49%) are pure false positives since no LLM existed, and the studies never subtract them; short reviews are the worst case for any detector; vendor-run
loke and GC, âDetecting AI-Generated Reviews for Corporate Reputation Managementâ (2025)
SciTePress, R. Loke, M. GC, 2025. Builds the ARED dataset of Amazon reviews and estimates âalmost 10%â AI reviews in a product sample, validated only by comparing with Fakespot and ReviewMeta, which detect fake reviews, not AI ones. Low trust.
schĂ€fer, Mofreh, Steinebach, âGhostwriters of the Marketplaceâ (2026)
ICWSM 2026, Karla SchÀfer, Mohammed Mofreh, Martin Steinebach, ICWSM, 2026. A 78k-review corpus of machine-generated Google reviews from six LLMs for training detectors; no prevalence estimate. Listed so nobody mistakes it for one.
academic papers and peer reviews
This is the best-studied domain, and the one where the methods differ most from DeGenTWeb: the strongest results use corpus-level word statistics, not per-document detectors.
liang et al., âMonitoring AI-Modified Content at Scaleâ (2024)
arXiv 2403.07183, Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, James Y. Zou, ICML, 2024. Already in the draft and in the paper collection.
- data: peer reviews at ICLR 2024, NeurIPS 2023, CoRL 2023, EMNLP 2023
- method: âa maximum likelihood modelâ over adjective frequencies, fit on expert-written and AI-generated reference reviews; estimates a corpus fraction, never labels a single review
- number: âbetween 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMsâ
- trust: robust to paraphrase in their tests, but the humanâs gen_ai.md objection stands: the AI reference set was generated the same way as the validation set, and humans adopt LLM words over time, which inflates later estimates
- relation to DeGenTWeb: cited. The important contrast is that a corpus fraction answers âhow muchâ without ever answering âwhichâ, while DeGenTWeb needs âwhich siteâ for its characterization findings
liang et al., âMapping the Increasing Use of LLMs in Scientific Papersâ (2024) and the Nature Human Behaviour version (2025)
arXiv 2404.01268, Weixin Liang et al., arXiv, 2024; published as âQuantifying large language model usage in scientific papersâ, Nature Human Behaviour, 2025. Already in the draft and the collection.
- data: 950,965 papers (arXiv version) and 1,121,912 in the journal version, January 2020 to September 2024, from arXiv, bioRxiv and Nature journals
- number: computer science âup to 17.5%â in the preprint and âup to 22%â in the journal version; mathematics and Nature portfolio âup to 6.3%â then âup to 9%â
- trust: same method and caveats as above; the drift from 17.5% to 22% between versions is partly more months of data
kobak, GonzĂĄlez-MĂĄrquez, HorvĂĄt, Lause, âDelving into LLM-assisted writing in biomedical publications through excess vocabularyâ (2025)
arXiv 2406.07016, Dmitry Kobak, Rita GonzĂĄlez-MĂĄrquez, EmĆke-Ăgnes HorvĂĄt, Jan Lause, Science Advances, 2025. PDF added.
- data: over 15 million PubMed abstracts, 2010 to 2024
- method: no detector and no reference LLM texts; they find words whose frequency jumped in 2024 beyond what the 2021 to 2022 trend predicts (âexcess vocabularyâ, same idea as excess mortality) and count abstracts with at least one such word
- number: âat least 13.5% of 2024 abstracts were processed with LLMsâ, âreaching 40% for some subcorporaâ
- trust: the cleanest method in the field because it needs no AI training data; it is a lower bound by construction and a one-word test, so it can only see LLM style words, not LLM content
- relation to DeGenTWeb: the excess-vocabulary idea could give DeGenTWeb a detector-free cross-check: do words like âdelveâ spike on DeGenTWeb-flagged sites but not on unflagged ones?
gray, âChatGPT âcontaminationââ (2024) and Geng and Trotta (2024)
arXiv 2403.16887, Andrew Gray, arXiv, 2024 (PDF added): keyword counts in Dimensions, âAt least 60,000 papers (slightly over 1% of all articles) were LLM-assistedâ in 2023. arXiv 2404.08627, Mingmeng Geng, Roberto Trotta, arXiv, 2024: arXiv abstract word frequencies, estimating about 35% of computer science abstracts revised by ChatGPT, against a ârevise the following sentencesâ baseline. Both are early word-count studies; the 1% vs 35% gap comes from whether one counts a few rare marker words (Gray) or fits a whole frequency shift (Geng).
wolfrath et al., âRising Prevalence of Detected AI-Generated Text in Medical Literatureâ (2026)
arXiv 2603.19316, Nathan Wolfrath, Simrin Patel, Madelyn Flitcroft, Anjishnu Banerjee, Melek Somai, Bradley H. Crotty, Anai N Kothari, arXiv, 2026. PDF added.
- data: 7,251 JAMA Network Open articles, January 2022 to March 2025
- method: Originality.ai, âsignificantâ AI content by its probability
- number: â195 articles (2.7%) were classified as containing significant AI-generated textâ, âincreased from 0.0% in January 2022 to 11.3% in March 2025â; âOnly 15 articles (0.2%) disclosed large language model useâ
- trust: the 0.0% in January 2022 is a useful implicit FPR check for this corpus; single commercial detector; one journal
latona et al., âThe AI Review Lotteryâ (2024)
arXiv 2405.02150, Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, Robert West, arXiv, 2024. In the draft and the collection.
- data: ICLR 2024 reviews; method: GPTZero
- number: âat least 15.8% of reviews were written with AI assistanceâ; AI-assisted reviews scored higher in 53.4% of paired cases, and borderline papers with one were 4.9 points more likely to be accepted
- trust: single proprietary detector; the effect sizes are small and could come from reviewer selection
ICLR 2026: Pangramâs 21% and follow-ups (2025 to 2026)
Pangramâs blog post on ICLR 2026 reviews could not be opened (404 on the URL I tried); the number is quoted in Abdulhai et al., arXiv 2603.18161 as âthe 21% of AI-generated scientific peer reviews at a recent top AI conferenceâ, and in press as 21% of 75,800 reviews fully AI with over half showing some AI use. Shen and Wang, arXiv 2602.00319, Siyuan Shen, Kai Wang, arXiv, 2026 (PDF added), train a detector on old reviews and find âapproximately 20% of ICLR reviews and 12% of Nature Communications reviews classified as AI-generated in 2025â. Duarte et al., âSem-Detectâ, arXiv 2605.21713, AndrĂ© V. Duarte, Brian Tufts, Aditya Oke, Fei Fang, Arlindo L. Oliveira, Lei Li, arXiv, 2026, report in their section 5.6 that on ICLR 2026 reviews EditLens âclassifies 24% of reviews as AI-generated, 32% as LLM-refined, and 44% as humanâ while Sem-Detect âpredicts 5%, 61%, and 34%, respectivelyâ, which is the clearest demonstration anywhere that the âfully AIâ number is a detector artifact once mixed writing is common. Kim et al., arXiv 2609.19420, a randomized experiment at ICML 2026 with 1,486 survey responses, gives the only self-reported ground truth: â22.5% of conservative-policy reviewers reported using an LLM despite the prohibitionâ and 36.5% under the permissive policy reported a disallowed use.
- relation to DeGenTWeb: the peer-review line now has what DeGenTWeb lacks, a self-report ground truth (ICML 2026 survey) to compare detectors against; and the EditLens vs Sem-Detect split is a ready-made citation for the draftâs limitations point that âHybrid human-LLM collaboration content may be common, blurring boundaryâ
other text in the wild
liang et al., âThe Widespread Adoption of Large Language Model-Assisted Writing Across Societyâ (2025)
arXiv 2502.09747, Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, James Zou, arXiv, 2025. PDF added.
- data: 687,241 CFPB consumer complaints, 537,413 corporate press releases, 304.3 million job postings, 15,919 UN press releases, January 2022 to September 2024
- method: the same population-level word-frequency framework as the peer review paper
- numbers: âroughly 18% of financial consumer complaint text appears to be LLM-assistedâ, corporate press releases âup to 24%â, job postings âjust below 10% in small firmsâ, UN press releases ânearly 14%â; growth âstabilized by 2024â
- trust: same method strengths and weaknesses as the other Liang papers; the plateau in 2024 is interesting and matches Graphiteâs plateau near 50% and Pewâs slowing curve, or it means detectors stopped seeing newer models
- relation to DeGenTWeb: the âplateau or blind spotâ question is exactly what DeGenTWebâs frontier-model experiment (accuracy 0.58 on 2025 models) speaks to; the draft can use it to argue that observed plateaus in other studies may be detector decay
hao et al., âDo Spammers Dream of Electric Sheep?â (2025)
In the draft and the humanâs notes: about 240k Barracuda-flagged spam and BEC emails, three detectors with very different results, a clear rise. Not re-read here.
puccetti et al., âAI âNewsâ Content Farms Are Easy to Make and Hard to Detectâ (2024)
In the draft (Italian news content farms). Not re-read here.
copyleaks, âExplosive Growth of AI Content Across the Webâ (2024)
Copyleaks press release, Copyleaks, 2024. Claims an â8,362%â rise from November 2022 to March 2024, ending at â1.57% of all web pages analyzedâ. Sample and method undisclosed. The 1.57% is interesting only because it is so far below Ahrefsâ 74% and Graphiteâs 39% for the same period: it shows vendor numbers are incomparable even with each other.
images online
I found no study that estimates what share of images on the open web are AI-generated. The literature on image detection is about benchmarks, not prevalence. What exists:
Chen et al. 2025 (above): 12.33% of election-related images on X, with a measured 6.5% FPR
DiResta and Goldstein, âHow spammers and scammers leverage AI-generated images on Facebook for audience growthâ, RenĂ©e DiResta, Josh A. Goldstein, Harvard Kennedy School Misinformation Review, 2024: a qualitative study of 125 Facebook Pages that each posted at least 50 AI images; it shows the feed recommends unlabeled AI images, gives no share
Matatov et al. 2024 (above): under 0.5% self-declared AI image posts in art subreddits through 2023
âGPT-Image-2 in the Wildâ, arXiv, 2026: a dataset of self-reported AI images on X from one modelâs first week; a data source, not a prevalence estimate
Pangram announced an image detector research preview in 2026 (blog); no measurement published that I found
the humanâs notes on image watermarking and C2PA cover why provenance labels rather than detectors are the plausible path to counting images
no comprehensive open-web image estimate was identified in this review
- detector performance after resizing and recompression is a relevant limitation
- image-detection study
cross-cutting lessons on trusting the numbers
these are agent interpretations of the reviewed measurements
these estimates answer different questions and use different samples and detectors
- Ahrefs: 2.5% pure-AI pages versus 74% containing some AI
- DeGenTWeb: 6% dominant sites in its historical sample
- 28.6% among sites first seen in 2025H1
- Pew: 35% of dated pages show AI signs
- Dolezal: 35% of new URLs are AI or assisted
- Russell: 27.5% of filtered tokens
- Graphite: 50% of sampled articles are mostly AI
- population, date, authorship definition, detector, and unit all vary
- these percentages cannot be compared as estimates of one quantity
historical positive rates reveal a useful negative-control problem
- they equal false-positive rates only if the sampled text is known human
- reviewed examples include Originality.aiâs historical search and review flags
- 2.27% in 2019 search results, 5.01% in Google reviews, 4.49% in TripAdvisor
- Graphite reports 2.2% in 2020
- DeGenTWebâs 0.29% is a historical positive rate
- it is not a general authorship-error bound
several reviewed 2026 measurements rely on Pangram
- includes Pew, Dolezal, Russell, and Graphite
- shared detector errors can produce agreement between studies
- agreement is not independent confirmation
- DeGenTWeb uses an open scoring method
including AI assistance changes the estimated quantity
- Dolezal includes assisted text; Ahrefs includes any AI span
- Originality includes posts with 15% AI
- Sem-Detect and EditLens estimate 5% and 24% fully AI on the same reviews
- detector choice also matters
- one MAGE check reports 98.9% of GPT-4-paraphrased human text still scoring human
- supports that tested setting
- does not establish stability across workflows or generators
corpus-level estimates do not directly supply individual-site labels
- Liang and Kobak answer mixture questions
- site characterization and ranking studies need a separate labeling method
apparent adoption plateaus can reflect changing detector recall
- DeGenTWeb reports weaker performance on newer generators
- that makes detector drift a competing explanation
- it does not establish which explanation produced any particular plateau
this review found no comprehensive open-web image prevalence estimate
- multilingual web measurement coverage is thin beyond Thompsonâs translation work
research we could do
- recommendations below build on DeGenTWeb and the reviewed studies
- novelty remains unconfirmed
- choose a bounded pilot before scaling
1: compare methods and units on one sample
- run DeGenTWeb, open Pangram, page-level Binoculars, and excess-word estimates on one Common Crawl sample
- report page, site, and text-volume estimates separately
- word-frequency mixture estimates have their own interpretation
- calibrate on documented human and generated writing
- pre-ChatGPT pages are historical controls rather than proven human authorship
- compare with Pew, Dolezal, Russell, Kobak, and Graphite
- isolates method and denominator differences
- cannot alone explain the published 2.5% to 74% spread
- populations, dates, and authorship definitions also differ
- stop if the differences follow directly from the chosen denominator
2: test whether detector decay could contribute to apparent plateaus
- rescore DeGenTWebâs known-generated frontier-model sites with frozen accessible detectors
- compare open Pangram, Fast-DetectGPT, and affordable commercial tools
- build on the draftâs limitations and Liang 2025
- declining recall would weaken an adoption-plateau interpretation
- it would not prove that adoption continued growing
- stop if the effect disappears with a stronger frozen baseline
3: cross-check with excess vocabulary
- apply Kobakâs method to Common Crawl prose by half-year
- test whether excess words concentrate on DeGenTWeb-flagged sites
- agreement is supporting evidence rather than independent ground truth
- both methods can respond to shared topic or style shifts
- compare historical periods and documented writing before interpreting generation
4: test aggregation for another unit
- possible units include Reddit users, outlets, Amazon reviewers, and peer reviewers
- reviewed leads include La Cava, Sun, Russell, Pangram, Shen, and Wang
- choose a unit with checkable authoring histories
- Kim et al.âs ICML self-reports are a lead
- self-reports need their own reliability assessment
- the humanâs notes already propose set-level applications
- this is an extension rather than a new aggregation idea
5: track search exposure across engines and years
- combine Webis-PSERP-24, Originality.ai, Graphite, and Bevendorffâs SEO-spam methods
- measure which generated-dominant sites appear and at what ranks
- distinguish ranking changes from detector drift and changing site samples
- search-engine updates are candidate interventions
- concurrent changes limit causal interpretation
6: evaluate multilingual website detection
- compare non-English Common Crawl prose with Thompsonâs machine-translation results
- test a multilingual scoring pair on documented human and generated sites
- translation and original generation are different writing processes
- calibrate them separately
- builds on Thompson and Brooks
- do not infer transfer from English detection alone
7: measure detectable provenance in web images
- sample images from a defined crawl or DeGenTWeb site set
- report C2PA, IPTC source labels, and accessible generator watermarks separately
- missing marks do not establish human authorship
- mark presence needs validation before becoming an origin label
- builds on Chen, DiResta, Brooks, Hanley, and image watermarking
- measures detectable evidence rather than total generated-image prevalence
- no comprehensive open-web estimate was identified in this review
8: study longitudinal site switching
- begin with DeGenTWebâs reported switched-site set
- compare Hanley and Russellâs findings about small publishers
- inspect when sites change and whether traffic or ads change afterward
- a detector transition is not itself an observed authoring transition
- rank and traffic changes alone do not establish effects caused by generation
- audit known histories before expanding a crawl
Last edited: