Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

measuring how much of the web is AI-generated (authored by agents unless marked 🧑)

  • targeted review of generated-content prevalence studies
    • compares their populations, units, definitions, and detector assumptions
    • reviewed 7 Oct 2026
  • compares against the newer local DeGenTWeb draft
    • public preprint
    • Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha
    • the local draft contains later measurements
  • consultation status is in the provenance index

Sibling notes: image_watermarking, labeling_rules_and_practice; detectors themselves are covered under llm_text, search and SEO under web_user, crawling under web_infra.

what DeGenTWeb measures (from the draft)

  • source: DeGenTWeb_writeup2/sections/ and macros in degentweb_imc2026.tex

  • unit: the website (subdomain), labeled “LLM-dominant” when most of its prose pages look LLM-written; not per page, not per token

  • method: Trafilatura extraction, Dolma heuristic filter plus 200-token minimum plus Rabin dedup, Binoculars (Falcon-7B pair) per page, 9 deciles of 20 pages per site into a linear SVM trained on a 120-site baseline (60 Russell 2000 and IndieWeb sites vs 60 Wix/B12 AI-builder mirrors)

  • numbers in the draft: 6.0% of 94,908 Common Crawl sites (2020 to May 2025) LLM-dominant, rising from 2.59% for sites first seen in 2022H2 to 28.6% for 2025H1; 16.4% of 38,309 Bing how-to result sites, 18.2% of result links; 0.29% of pre-ChatGPT Common Crawl subdomains flagged (historical positive rate); 98.7% baseline site accuracy; accuracy drops to 0.58 on sites generated by 2025 frontier models (Pearson -0.88 with the Artificial Analysis index)

  • the draft’s related-work section cites only Liang 2024 (peer reviews and papers), Latona 2024 (ICLR), Hao 2025 (spam email), Brooks 2024 (Wikipedia) and NewsGuard; the sections below cover the much larger body of prevalence studies the draft does not yet cite

open web and web crawls

  • closest reviewed studies sample Common Crawl or the Internet Archive
    • their page-level decisions differ from DeGenTWeb’s website labels
    • historical positive rates and known-authorship error rates must be distinguished

Pew Research Center, “How Much of the Internet Is Written With AI?” (2026)

Main report and methodology, Pew Research Center Data Labs, August 2026.

  • what was measured: share of pages per Common Crawl crawl that show “meaningful” signs of AI authorship
  • data: “10,000 English-language pages from each of the 49 crawls created between January 2021 and July 2026”, 490,000 pages in all
  • method: Pangram’s open model editlens_Llama-3.2-3B, score 0 to 1, pages scoring 0.2 or higher count as AI; they checked it against the commercial Pangram 3.3 on 62,370 pages: “They agree in 96% of cases, with a Cohen’s kappa value of 0.61”
  • number: “35% of pages in the July 2026 crawl with post-ChatGPT publication dates show signs of AI authorship”; over all sampled pages regardless of date, “10% show significant signs of AI authorship”, about one in ten .com pages, half that on .org, about 1% on .edu and .gov
  • trust: the 35% is over the 10 to 15% of pages that carry a publication date in their HTML, which Pew says “reflect AI authorship among dated content rather than the web as a whole”; kappa 0.61 between the open and paid model is only moderate agreement; no false positive rate on known human web pages is reported, and Pew itself says “Reported shares from 2021 and 2022 crawls should be treated as approximate” because early crawls show more false positives
  • relation to DeGenTWeb: same data source, page level instead of site level, one detector instead of a zero-shot score plus aggregation; their 35% of dated pages and DeGenTWeb’s 28.6% of sites first seen in 2025H1 are in the same ballpark despite different units, which is worth saying in the paper

dolezal, Alam, Graham, Bohacek, “The Impact of AI-Generated Text on the Internet” (2026)

arXiv 2604.26965, Jonas Dolezal, Sawood Alam, Mark Graham, Maty Bohacek (Internet Archive and Stanford), arXiv, 2026.

  • what was measured: share of newly archived URLs per month that are AI-generated or AI-assisted, plus effects on diversity and sentiment
  • data: Internet Archive CDX index, “33 monthly intervals from August 2022 to May 2025”, target 10,000 URLs per month, stratified by first-archival time, MIME type, URL depth and TLD, one URL per host
  • method: Pangram v3 API, three-way output (fully AI, AI-assisted, human); the headline sums AI and AI-assisted; they compared Pangram against Binoculars, DivEye and Desklib on synthetic tests of length, HTML wrapping, model family and language and say Pangram had “perfect accuracy on texts exceeding 50 words”
  • number: “By the first half of 2025, as much as 35% of websites uploaded to the internet in a given month were AI-generated or AI-assisted”
  • trust: “AI-assisted” is a wide bucket, so 35% is an upper-leaning number; the detector comparison is on their own synthetic texts, not on real web pages with known authorship; one URL per host means a site with a thousand pages counts the same as a site with one, which is actually closer to DeGenTWeb’s site unit than Pew’s page unit; the secondary claims about diversity and sentiment lean on correlations across months
  • relation to DeGenTWeb: the nearest competitor in spirit (new sites over time, Internet Archive). DeGenTWeb’s 28.6% for 2025H1 sites and their 35% for mid-2025 URLs would look consistent if one subtracts their AI-assisted share. The draft should cite it and point out that DeGenTWeb’s LLM-dominant label is stricter than “AI-assisted”

russell et al., “How Much Is an AI Token Worth?” (2026)

arXiv 2609.40295, Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi (UMass Amherst and Pangram), arXiv, 2026.

  • what was measured: share of tokens in 2026 web crawls labeled AI, and what training on them does to language models
  • data: June and August 2026 web crawls after FineWeb quality filtering; the “WildAI” corpus of 83B tokens with AI, topic and format labels
  • method: Pangram detector (the authors include Pangram’s founders)
  • number: “27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August”
  • trust: a token share after FineWeb filtering, so boilerplate and non-prose pages are already removed, which inflates the share relative to raw crawls; the detector vendor is a coauthor; no independent ground truth
  • relation to DeGenTWeb: shows the “LLM content hurts training” angle that DeGenTWeb mentions only via Shumailov 2024; also a third unit (tokens) to compare with pages (Pew) and sites (DeGenTWeb)

graphite, “More Articles Are Now Created by AI Than Humans” (2025) and the Q1 2026 update

Graphite report, Graphite (an SEO agency), October 2025; the follow-up was covered by Axios, which I could not open directly.

  • what was measured: share of English articles on the web that are mostly AI-written
  • data: “43,000 English-language URLs randomly selected” from Common Crawl, published January 2020 to May 2025 (press coverage says 65,000; the page says 43,000)
  • method: Surfer’s AI detector on 500-word chunks; an article counts as AI “if the algorithm predicts that more than 50% of the content is AI-generated”; they report a 4.2% false positive rate on articles from November 2020 to November 2022 and 0.6% false negatives on 6,009 GPT-4o articles
  • number: AI articles passed human ones in November 2024; about 39% a year after ChatGPT; the update with three detectors (Pangram, Copyleaks, GPTZero) averaged over about 55,000 articles found 49.9% in Q1 2026
  • trust: the 4.2% FPR is measured on their own sample and only against GPT-4o for false negatives; “article” filtering is not described in detail; the authors sell SEO services; the “more than half” headline got repeated by many outlets without the sampling caveats
  • relation to DeGenTWeb: this is the main “cried wolf” headline the intro should name. Their 2020 share of 2.2% is in effect their false positive rate, which matches what DeGenTWeb does with pre-ChatGPT subdomains (0.29%) but at a far worse level, so DeGenTWeb’s number is much more defensible

ahrefs, “74% of New Webpages Include AI Content” (2025)

Ahrefs blog, Ahrefs, May 2025.

  • data: 900,000 newly created English pages seen in April 2025, one per domain
  • method: Ahrefs’ own detector, codenamed “bot_or_not”; no accuracy or false positive numbers given
  • number: “74.2% of them contained AI-generated content”; “2.5% of pages were categorized as ‘pure AI’”, 25.8% pure human, 71.7% mixed
  • trust: low. The 74% counts any page with any flagged span, and the detector is unvalidated in public; the 2.5% “pure AI” is the comparable number to DeGenTWeb’s LLM-dominant share, and it is close to DeGenTWeb’s 6% average over 2020 to 2025 given the different units
  • relation to DeGenTWeb: a clear example of the “claims vary widely” point; the gap between 74% and 2.5% within the same study shows why the LLM-dominant definition matters

thompson et al., “A Shocking Amount of the Web is Machine Translated” (2024)

arXiv 2401.05749, Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, Marcello Federico, Findings of ACL, 2024.

  • what was measured: how much of multilingual web text is machine translated, the pre-LLM form of machine-generated web text
  • data: 6.4B sentences in 90 languages from Common Crawl via the ccAligned pipeline, grouped into multi-way parallel sets
  • method: no detector; they reason that content translated into many languages at once is low quality and thus machine made, and confirm with human quality ratings
  • number: “57.1% of the sentences in our corpus are multi-way parallel in at least 3 languages” (from the abstract)
  • trust: the inference from multi-way parallelism to machine translation is indirect; it says nothing about monolingual English content
  • relation to DeGenTWeb: precedent for “machine content dominates the web” claims before LLMs; also a reminder that DeGenTWeb’s English-only Binoculars scoring misses the non-English web, where generated content may be more common

search results

originality.ai, “Amount of AI Content in Google Search Results” (ongoing, 2019 to 2025)

Originality.ai study, Originality.ai (a detector vendor), updated through September 2025.

  • data: 500 informational keywords, top 20 Google results each, about 10,000 pages per two-month period, from January 2019 onward; “Text from websites that were not article-based, like YouTube and Reddit, was removed”
  • method: Originality.ai’s own detector at 0.5 confidence
  • numbers: 2.27% in February 2019, 8.48% December 2023, 7.43% right after the March 2024 core update, 19.56% July 2025, 17.31% September 2025
  • trust: the 2019 value of 2.27% is a false positive floor the study never subtracts; no validation on web pages; the vendor sells the detector
  • relation to DeGenTWeb: the direct comparison for the Bing how-to result, where DeGenTWeb finds 16.4% of sites and 18.2% of links LLM-dominant. The numbers agree surprisingly well. DeGenTWeb adds a transparent method, a false positive bound, and the search vs open web gap (16.4% vs 6.0%)

graphite, “How Does AI-Generated Content Perform in Search and Answer Engines?” (2025)

Graphite report, Graphite, 2025. Not read in full; the companion to the article study above. They claim AI articles “largely do not appear in Google and ChatGPT”, that is, AI content is common in the crawl but rare among top results. This contradicts DeGenTWeb’s finding that Bing does not down-rank LLM-dominant sites, and the Originality.ai trend. Worth a direct look before the paper claims either way; Graphite’s search sample is drawn from their SEO clients’ queries, not random how-to questions.

deGenTWeb’s own Webis-PSERP-24 analysis (hidden section of the draft)

The draft has a hidden subsection on Webis-PSERP-24, top-20 results for 7,392 product-review queries from Startpage, Bing, DuckDuckGo and ChatNoir from 2022 to 2024, with the share of Startpage result links to LLM-dominant sites rising over time. No other study I found measures LLM content in a multi-engine, multi-year SERP archive, so that section is more novel than the draft treats it.

news sites

hanley and Durumeric, “Machine-Made Media” (2024)

arXiv 2305.09820, Hans W. A. Hanley, Zakir Durumeric, ICWSM, 2024. PDF now in the paper collection.

  • what was measured: share of news articles that are machine-generated, on mainstream vs misinformation sites, January 2022 to May 2023
  • data: “over 15.46 million articles from 3,074 misinformation and mainstream news websites”, scraped with Selenium, text and dates from newspaper3k and htmldate
  • method: a DeBERTa classifier trained on human news plus GPT-2/GPT-3.5 generations, with paraphrase and perturbation augmentation; the threshold was raised to 0.98 “allowing us to achieve a 1% FPR/accuracy on the Signal article dataset”, at which point “our model reaches a precision of 0.989 on our ChatGPT rewrite test set at the expense of only reaching a 0.639 recall”
  • numbers: absolute levels are small: “1.07% of all articles published in January 2022 (12,984 of 1,213,983 articles) were synthetically generated. However, by May 2023, the fraction of synthetic articles nearly went up to 1.78%”; relative increases of 57.3% (mainstream) and 474% (misinformation); the least popular sites (rank beyond 10M in their popularity ranking) rose 3.42 absolute points, the most popular under 1 point
  • trust: the best-validated news study. They explicitly accept low recall for a 1% FPR, and report absolute shares, not just the headline relative increases. Pre-ChatGPT “synthetic” articles are mostly templated finance news, which the detector also flags. Period ends May 2023, so the numbers are early
  • relation to DeGenTWeb: the closest methodological cousin. Same low-FPR philosophy, same “small sites drive the rise” finding (DeGenTWeb’s LLM-dominant sites are entry-level ad stack, low engagement). DeGenTWeb should cite it as the news-only precedent and contrast its site-level verdict with their per-article one. Their 1.78% in May 2023 vs DeGenTWeb’s 6% site share through 2025 are not in conflict given the time gap

russell et al., “AI use in American newspapers is widespread, uneven, and rarely disclosed” (2026)

arXiv 2510.18774, Jenna Russell, Marzena Karpinska, Destiny Akinode, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer, ACL, 2026. Already in the paper collection.

  • data: “186K articles from online editions of 1.5K American newspapers published in the summer of 2025”, plus 45K opinion pieces from the Washington Post, New York Times and Wall Street Journal
  • method: Pangram v2, which scores segments and then gives an overall label
  • numbers: “approximately 9% of newly-published articles are either partially or fully AI-generated”; 1.7% at large papers vs 9.3% at small local ones; opinion pieces “6.4 times more likely to contain AI-generated content than news articles”; of 100 flagged articles, only five disclosed AI use
  • trust: large, well-sampled, and the manual audit of flagged articles is a strength; but the detector is the coauthors’ product and no FPR is measured on this corpus; “partially” AI is a soft label
  • relation to DeGenTWeb: shows the per-outlet view (newspapers are sites too). DeGenTWeb’s search-result site categories (publisher/editorial with ads) overlap with small local papers. A useful contrast: DeGenTWeb asks which whole sites are LLM-run, this asks how much AI slips into human-run outlets

ansari, Zhang, Tripto, Lee, “Echoes of Automation” (2025)

arXiv 2508.06445, Abolfazl Ansari, Delvin Ce Zhang, Nafis Irtiza Tripto, Dongwon Lee, SBP-BRiMS, 2025. PDF added.

  • data: “over 40,000 news articles from major, local, and college news media”
  • method: three detectors, Binoculars, Fast-DetectGPT and GPTZero, plus sentence-level analysis
  • numbers: the abstract gives no headline share, only “substantial increase of GenAI use in recent years, especially in local and college news”, and that LLM text shows up in introductions more than conclusions
  • trust: three detectors is good practice, but without a reported FPR or absolute share the finding is directional
  • relation to DeGenTWeb: uses Binoculars too, so its per-sentence findings hint at what DeGenTWeb’s page scores respond to in mixed articles

newsGuard AI Tracking Center (ongoing)

NewsGuard AI Tracking Center, NewsGuard, 2023 to 2026. The draft already cites this.

  • what: a hand-curated count of “unreliable AI-generated news” sites, with press coverage saying 3,749 sites in 16 languages by June 2026, up from 1,254 at the time the draft’s citation was written
  • trust: a count of found sites, not a share of anything; detection is editorial judgment, criteria not reproducible; the number only ever grows
  • relation to DeGenTWeb: the draft’s intro uses it for motivation, which is right; it cannot be used as a prevalence estimate

wikipedia

brooks, Eggert, Peskoff, “The Rise of AI-Generated Content in Wikipedia” (2024)

arXiv 2410.08044, Creston Brooks, Samuel Eggert, Denis Peskoff, WikiNLP workshop at EMNLP, 2024. Already in the paper collection and in the draft’s bibliography; the human’s gen_ai.md notes already criticize it (“unscientific bc assume paper i.i.d.”).

  • data: recently created articles in English, German, French and Italian; the Signpost coverage says 2,909 English articles created in August 2024
  • method: GPTZero and Binoculars (Falcon-7B), “thresholds calibrated to achieve a 1% false positive rate on pre-GPT-3.5 articles”, then subtract the pre-GPT positive rate to get a lower bound
  • number: “Detectors flag over 5% of newly created English Wikipedia articles as AI-generated”, lower in the other languages; 4.36% after subtracting the baseline
  • trust: the FPR calibration on pre-GPT articles is the same trick DeGenTWeb uses with pre-ChatGPT Common Crawl; sample is small; Binoculars has higher false negatives on non-English, so the cross-language comparison is weak
  • relation to DeGenTWeb: cited already; worth noting that no newer Wikipedia prevalence study turned up in my searches, only a 2025 Wikimedia Research Fund proposal, so the field is open for a repeat with a 2026 sample

social media and forums

sun et al., “Are We in the AI-Generated Text World Already?” (2025)

arXiv 2412.18148, Zhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang, Yule Liu, Xinlei He, Michael Backes, Yang Zhang, ACL, 2025. PDF added.

  • data: Medium 1,170,821 posts, Quora 245,131 answers (January 2022 to October 2024), Reddit 982,440 comments (to July 2024)
  • method: their own detector OSM-Det, a Longformer fine-tuned on AIGTBench (28.77M AI and 13.55M human samples from 12 LLMs); accuracy 0.979; FPR measured on pre-2022 posts: Medium 1.82%, Quora 1.36%, Reddit 1.70%
  • numbers: AI attribution rate on Medium “rising from 1.77% to 37.03%”, Quora “2.06% to 38.95%”, Reddit “1.31% to 2.45%”
  • trust: FPR on pre-LLM posts is measured, which is good, but the ~1.5% FPR is the whole Reddit signal, so the Reddit number means nothing; the detector was trained on 12 LLMs mostly GPT and Llama; “AI attribution” includes polished human text
  • relation to DeGenTWeb: Medium and Quora at ~37% mirror DeGenTWeb’s rise in new sites to 28.6%. The platform gap (Medium vs Reddit) parallels the DeGenTWeb search vs open web gap: where there is money or reach, there is more LLM text

la Cava, Aiello, Tagarelli, “Machines in the Crowd” (2025)

arXiv 2510.07226, Lucio La Cava, Luca Maria Aiello, Andrea Tagarelli, arXiv, 2025. PDF added.

  • data: “38,074,021 comments and 4,073,586 submissions” from 51 subreddits, January 2022 to December 2024
  • method: Fast-DetectGPT at a very strict threshold (0.99) and only texts of at least 250 tokens
  • numbers: machine text “marginally present on Reddit”, with peaks of 6.33% in r/askscience (July 2023), 7.69% in r/malefashionadvice, 8.46% in r/teenagers (February 2023); about 2% of active users produce it
  • trust: the authors call their estimate “very conservative”; the strict threshold means the numbers are floors; no FPR reported for the threshold on pre-LLM Reddit text in the abstract
  • relation to DeGenTWeb: an example of a per-user aggregation idea (the human’s gen_ai.md lists “set-level detection applications: Reddit users”) done loosely; DeGenTWeb’s decile-SVM could be applied per user

chen, Ye, Ferrara, Luceri, “Prevalence, Sharing Patterns, and Spreaders of Multimodal AI-Generated Content on X” (2025)

arXiv 2502.11248, Zhiyi Chen, Jinyi Ye, Emilio Ferrara, Luca Luceri, arXiv, 2025. Already in the paper collection and in the human’s gen_ai.md notes.

  • data: 2.5 million images from X posts about the 2024 US election, July 1 to September 30, 2024
  • method: GPT-4o as the image classifier, F1 0.96 on a 2,400-image test set; on 2,000 manually labeled real-world images, “false positive and false negative rates of 6.5% and 1.6%”
  • numbers: “approximately 12.33% of images related to the 2024 U.S. Election on X are AI-generated”; 10% of spreaders account for 80% of AI images; the earlier version also reports 1.4% AI text by a fine-tuned RoBERTa
  • trust: a 6.5% FPR on a 12% positive rate means roughly half the flagged images could be false positives; the text number is within detector error, as the human’s notes say
  • relation to DeGenTWeb: the “superspreader” concentration matches the DeGenTWeb mass-produced cluster finding (a few operators make much of the content)

pangram, “AI in Your Feed” (2026)

Pangram blog, Pangram Labs, 2026; covered by The Decoder.

  • data: 1,002,627 posts over 50 words scanned by users of Pangram’s Chrome extension between April and June 2026, on LinkedIn, Medium, Substack, X and Reddit
  • method: Pangram 3.3, claimed “0.01% false positive rate”
  • numbers: fully AI longform: LinkedIn over 40%, Medium 31%, X 23.9% (plus 22.9% AI-assisted), Reddit 13%, Substack 10%
  • trust: the sample is whatever extension users chose to scan, so it is biased toward suspicious posts; the 0.01% FPR claim is the vendor’s on its own benchmark; still the largest multi-platform sample
  • relation to DeGenTWeb: another vendor headline the intro can group with Graphite and Ahrefs

originality.ai, “Social Media AI Tracker” (monthly, 2026)

Originality.ai August 2026 tracker, Originality.ai, 2026.

  • data: a few hundred random longform English posts per platform per month (LinkedIn 119, Threads 162, Facebook 166, X 174, Reddit 564 in August 2026)
  • method: a post is “Likely AI” when the detector is “more confident than not (50% confidence or higher) that at least 15% was written by AI”
  • numbers: LinkedIn 76%, Threads 66%, Facebook 59%, X 57%, Reddit 31%
  • trust: very low. Tiny samples, a definition that counts a post with 15% AI as AI, and an unvalidated vendor detector. The Reddit 31% vs Pangram’s 13% vs La Cava’s under 9% peaks vs Sun’s 2.45% shows how much the definition and detector drive the number
  • relation to DeGenTWeb: the best single illustration of “claims vary widely” for the intro: four Reddit estimates spanning 2% to 31%

matatov, Aubin Le QuĂ©rĂ©, Amir, Naaman, “AI-Generated Media in Art Subreddits” (2024)

arXiv 2410.07302, Hana Matatov, Marianne Aubin Le Quéré, Ofra Amir, Mor Naaman, arXiv, 2024.

  • what: self-declared AI posts and AI accusations in art subreddits through 2023; no detector
  • number: AI posts and accusations “accounting for fewer than 0.5% of the image-based posts”
  • trust: self-labels are a floor; relevant mostly as a norm study

kapwing, “YouTube AI slop” report (2026)

Kapwing report via NeoMam, Kapwing, 2026 (press coverage in LBC, SlashGear and others).

  • method: hand-checked the top 100 trending channels per country on playboard.co (15,000 channels), found 278 “AI slop” channels; a fresh account’s first 500 recommended Shorts had 104 AI slop videos (21%)
  • trust: hand judgment of “slop”, a single new account, a marketing company; directional only
  • relation to DeGenTWeb: video is outside DeGenTWeb’s scope, but the “mass produced for ad revenue” story is the same

product reviews

No academic study with a defensible in-the-wild prevalence number exists for reviews. All numbers below come from detector vendors.

pangram, “Three percent of front-page Amazon reviews are now AI-generated” (2026)

Pangram blog, Pangram Labs, May 2026.

  • data: 30,000 front-page reviews on 500 best-selling products in ten categories
  • numbers: “3% of the total reviews studied - 909 total reviews - were AI-generated with high confidence”; 74% of AI reviews are 5-star vs 59% of human; 93% carry “Verified Purchase”
  • trust: front-page reviews only; “high confidence” threshold means a floor; no FPR measured on pre-ChatGPT reviews

originality.ai, Google reviews and TripAdvisor studies (2025)

Google reviews and TripAdvisor, Originality.ai, 2025.

  • data: up to 5 Google reviews per place in 15 North American cities across 20 categories; 10,000 TripAdvisor reviews in 20 cities
  • numbers: Google “5.01%” in 2019 to “19%” in 2024; TripAdvisor “from 4.49% in 2019 to 10.7% in 2024”
  • trust: the 2019 values (5.01% and 4.49%) are pure false positives since no LLM existed, and the studies never subtract them; short reviews are the worst case for any detector; vendor-run

loke and GC, “Detecting AI-Generated Reviews for Corporate Reputation Management” (2025)

SciTePress, R. Loke, M. GC, 2025. Builds the ARED dataset of Amazon reviews and estimates “almost 10%” AI reviews in a product sample, validated only by comparing with Fakespot and ReviewMeta, which detect fake reviews, not AI ones. Low trust.

schĂ€fer, Mofreh, Steinebach, “Ghostwriters of the Marketplace” (2026)

ICWSM 2026, Karla SchÀfer, Mohammed Mofreh, Martin Steinebach, ICWSM, 2026. A 78k-review corpus of machine-generated Google reviews from six LLMs for training detectors; no prevalence estimate. Listed so nobody mistakes it for one.

academic papers and peer reviews

This is the best-studied domain, and the one where the methods differ most from DeGenTWeb: the strongest results use corpus-level word statistics, not per-document detectors.

liang et al., “Monitoring AI-Modified Content at Scale” (2024)

arXiv 2403.07183, Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel A. McFarland, James Y. Zou, ICML, 2024. Already in the draft and in the paper collection.

  • data: peer reviews at ICLR 2024, NeurIPS 2023, CoRL 2023, EMNLP 2023
  • method: “a maximum likelihood model” over adjective frequencies, fit on expert-written and AI-generated reference reviews; estimates a corpus fraction, never labels a single review
  • number: “between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs”
  • trust: robust to paraphrase in their tests, but the human’s gen_ai.md objection stands: the AI reference set was generated the same way as the validation set, and humans adopt LLM words over time, which inflates later estimates
  • relation to DeGenTWeb: cited. The important contrast is that a corpus fraction answers “how much” without ever answering “which”, while DeGenTWeb needs “which site” for its characterization findings

liang et al., “Mapping the Increasing Use of LLMs in Scientific Papers” (2024) and the Nature Human Behaviour version (2025)

arXiv 2404.01268, Weixin Liang et al., arXiv, 2024; published as “Quantifying large language model usage in scientific papers”, Nature Human Behaviour, 2025. Already in the draft and the collection.

  • data: 950,965 papers (arXiv version) and 1,121,912 in the journal version, January 2020 to September 2024, from arXiv, bioRxiv and Nature journals
  • number: computer science “up to 17.5%” in the preprint and “up to 22%” in the journal version; mathematics and Nature portfolio “up to 6.3%” then “up to 9%”
  • trust: same method and caveats as above; the drift from 17.5% to 22% between versions is partly more months of data

kobak, González-Márquez, Horvát, Lause, “Delving into LLM-assisted writing in biomedical publications through excess vocabulary” (2025)

arXiv 2406.07016, Dmitry Kobak, Rita González-Márquez, EmƑke-Ágnes Horvát, Jan Lause, Science Advances, 2025. PDF added.

  • data: over 15 million PubMed abstracts, 2010 to 2024
  • method: no detector and no reference LLM texts; they find words whose frequency jumped in 2024 beyond what the 2021 to 2022 trend predicts (“excess vocabulary”, same idea as excess mortality) and count abstracts with at least one such word
  • number: “at least 13.5% of 2024 abstracts were processed with LLMs”, “reaching 40% for some subcorpora”
  • trust: the cleanest method in the field because it needs no AI training data; it is a lower bound by construction and a one-word test, so it can only see LLM style words, not LLM content
  • relation to DeGenTWeb: the excess-vocabulary idea could give DeGenTWeb a detector-free cross-check: do words like “delve” spike on DeGenTWeb-flagged sites but not on unflagged ones?

gray, “ChatGPT ‘contamination’” (2024) and Geng and Trotta (2024)

arXiv 2403.16887, Andrew Gray, arXiv, 2024 (PDF added): keyword counts in Dimensions, “At least 60,000 papers (slightly over 1% of all articles) were LLM-assisted” in 2023. arXiv 2404.08627, Mingmeng Geng, Roberto Trotta, arXiv, 2024: arXiv abstract word frequencies, estimating about 35% of computer science abstracts revised by ChatGPT, against a “revise the following sentences” baseline. Both are early word-count studies; the 1% vs 35% gap comes from whether one counts a few rare marker words (Gray) or fits a whole frequency shift (Geng).

wolfrath et al., “Rising Prevalence of Detected AI-Generated Text in Medical Literature” (2026)

arXiv 2603.19316, Nathan Wolfrath, Simrin Patel, Madelyn Flitcroft, Anjishnu Banerjee, Melek Somai, Bradley H. Crotty, Anai N Kothari, arXiv, 2026. PDF added.

  • data: 7,251 JAMA Network Open articles, January 2022 to March 2025
  • method: Originality.ai, “significant” AI content by its probability
  • number: “195 articles (2.7%) were classified as containing significant AI-generated text”, “increased from 0.0% in January 2022 to 11.3% in March 2025”; “Only 15 articles (0.2%) disclosed large language model use”
  • trust: the 0.0% in January 2022 is a useful implicit FPR check for this corpus; single commercial detector; one journal

latona et al., “The AI Review Lottery” (2024)

arXiv 2405.02150, Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, Robert West, arXiv, 2024. In the draft and the collection.

  • data: ICLR 2024 reviews; method: GPTZero
  • number: “at least 15.8% of reviews were written with AI assistance”; AI-assisted reviews scored higher in 53.4% of paired cases, and borderline papers with one were 4.9 points more likely to be accepted
  • trust: single proprietary detector; the effect sizes are small and could come from reviewer selection

ICLR 2026: Pangram’s 21% and follow-ups (2025 to 2026)

Pangram’s blog post on ICLR 2026 reviews could not be opened (404 on the URL I tried); the number is quoted in Abdulhai et al., arXiv 2603.18161 as “the 21% of AI-generated scientific peer reviews at a recent top AI conference”, and in press as 21% of 75,800 reviews fully AI with over half showing some AI use. Shen and Wang, arXiv 2602.00319, Siyuan Shen, Kai Wang, arXiv, 2026 (PDF added), train a detector on old reviews and find “approximately 20% of ICLR reviews and 12% of Nature Communications reviews classified as AI-generated in 2025”. Duarte et al., “Sem-Detect”, arXiv 2605.21713, AndrĂ© V. Duarte, Brian Tufts, Aditya Oke, Fei Fang, Arlindo L. Oliveira, Lei Li, arXiv, 2026, report in their section 5.6 that on ICLR 2026 reviews EditLens “classifies 24% of reviews as AI-generated, 32% as LLM-refined, and 44% as human” while Sem-Detect “predicts 5%, 61%, and 34%, respectively”, which is the clearest demonstration anywhere that the “fully AI” number is a detector artifact once mixed writing is common. Kim et al., arXiv 2609.19420, a randomized experiment at ICML 2026 with 1,486 survey responses, gives the only self-reported ground truth: “22.5% of conservative-policy reviewers reported using an LLM despite the prohibition” and 36.5% under the permissive policy reported a disallowed use.

  • relation to DeGenTWeb: the peer-review line now has what DeGenTWeb lacks, a self-report ground truth (ICML 2026 survey) to compare detectors against; and the EditLens vs Sem-Detect split is a ready-made citation for the draft’s limitations point that “Hybrid human-LLM collaboration content may be common, blurring boundary”

other text in the wild

liang et al., “The Widespread Adoption of Large Language Model-Assisted Writing Across Society” (2025)

arXiv 2502.09747, Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, James Zou, arXiv, 2025. PDF added.

  • data: 687,241 CFPB consumer complaints, 537,413 corporate press releases, 304.3 million job postings, 15,919 UN press releases, January 2022 to September 2024
  • method: the same population-level word-frequency framework as the peer review paper
  • numbers: “roughly 18% of financial consumer complaint text appears to be LLM-assisted”, corporate press releases “up to 24%”, job postings “just below 10% in small firms”, UN press releases “nearly 14%”; growth “stabilized by 2024”
  • trust: same method strengths and weaknesses as the other Liang papers; the plateau in 2024 is interesting and matches Graphite’s plateau near 50% and Pew’s slowing curve, or it means detectors stopped seeing newer models
  • relation to DeGenTWeb: the “plateau or blind spot” question is exactly what DeGenTWeb’s frontier-model experiment (accuracy 0.58 on 2025 models) speaks to; the draft can use it to argue that observed plateaus in other studies may be detector decay

hao et al., “Do Spammers Dream of Electric Sheep?” (2025)

In the draft and the human’s notes: about 240k Barracuda-flagged spam and BEC emails, three detectors with very different results, a clear rise. Not re-read here.

puccetti et al., “AI ‘News’ Content Farms Are Easy to Make and Hard to Detect” (2024)

In the draft (Italian news content farms). Not re-read here.

copyleaks, “Explosive Growth of AI Content Across the Web” (2024)

Copyleaks press release, Copyleaks, 2024. Claims an “8,362%” rise from November 2022 to March 2024, ending at “1.57% of all web pages analyzed”. Sample and method undisclosed. The 1.57% is interesting only because it is so far below Ahrefs’ 74% and Graphite’s 39% for the same period: it shows vendor numbers are incomparable even with each other.

images online

I found no study that estimates what share of images on the open web are AI-generated. The literature on image detection is about benchmarks, not prevalence. What exists:

  • Chen et al. 2025 (above): 12.33% of election-related images on X, with a measured 6.5% FPR

  • DiResta and Goldstein, “How spammers and scammers leverage AI-generated images on Facebook for audience growth”, RenĂ©e DiResta, Josh A. Goldstein, Harvard Kennedy School Misinformation Review, 2024: a qualitative study of 125 Facebook Pages that each posted at least 50 AI images; it shows the feed recommends unlabeled AI images, gives no share

  • Matatov et al. 2024 (above): under 0.5% self-declared AI image posts in art subreddits through 2023

  • “GPT-Image-2 in the Wild”, arXiv, 2026: a dataset of self-reported AI images on X from one model’s first week; a data source, not a prevalence estimate

  • Pangram announced an image detector research preview in 2026 (blog); no measurement published that I found

  • the human’s notes on image watermarking and C2PA cover why provenance labels rather than detectors are the plausible path to counting images

  • no comprehensive open-web image estimate was identified in this review

cross-cutting lessons on trusting the numbers

  • these are agent interpretations of the reviewed measurements

  • these estimates answer different questions and use different samples and detectors

    • Ahrefs: 2.5% pure-AI pages versus 74% containing some AI
    • DeGenTWeb: 6% dominant sites in its historical sample
      • 28.6% among sites first seen in 2025H1
    • Pew: 35% of dated pages show AI signs
    • Dolezal: 35% of new URLs are AI or assisted
    • Russell: 27.5% of filtered tokens
    • Graphite: 50% of sampled articles are mostly AI
    • population, date, authorship definition, detector, and unit all vary
    • these percentages cannot be compared as estimates of one quantity
  • historical positive rates reveal a useful negative-control problem

    • they equal false-positive rates only if the sampled text is known human
    • reviewed examples include Originality.ai’s historical search and review flags
      • 2.27% in 2019 search results, 5.01% in Google reviews, 4.49% in TripAdvisor
    • Graphite reports 2.2% in 2020
    • DeGenTWeb’s 0.29% is a historical positive rate
      • it is not a general authorship-error bound
  • several reviewed 2026 measurements rely on Pangram

    • includes Pew, Dolezal, Russell, and Graphite
    • shared detector errors can produce agreement between studies
    • agreement is not independent confirmation
    • DeGenTWeb uses an open scoring method
  • including AI assistance changes the estimated quantity

    • Dolezal includes assisted text; Ahrefs includes any AI span
    • Originality includes posts with 15% AI
    • Sem-Detect and EditLens estimate 5% and 24% fully AI on the same reviews
      • detector choice also matters
    • one MAGE check reports 98.9% of GPT-4-paraphrased human text still scoring human
      • supports that tested setting
      • does not establish stability across workflows or generators
  • corpus-level estimates do not directly supply individual-site labels

    • Liang and Kobak answer mixture questions
    • site characterization and ranking studies need a separate labeling method
  • apparent adoption plateaus can reflect changing detector recall

    • DeGenTWeb reports weaker performance on newer generators
    • that makes detector drift a competing explanation
    • it does not establish which explanation produced any particular plateau
  • this review found no comprehensive open-web image prevalence estimate

    • multilingual web measurement coverage is thin beyond Thompson’s translation work

research we could do

  • recommendations below build on DeGenTWeb and the reviewed studies
    • novelty remains unconfirmed
    • choose a bounded pilot before scaling

1: compare methods and units on one sample

  • run DeGenTWeb, open Pangram, page-level Binoculars, and excess-word estimates on one Common Crawl sample
  • report page, site, and text-volume estimates separately
    • word-frequency mixture estimates have their own interpretation
  • calibrate on documented human and generated writing
    • pre-ChatGPT pages are historical controls rather than proven human authorship
  • compare with Pew, Dolezal, Russell, Kobak, and Graphite
  • isolates method and denominator differences
    • cannot alone explain the published 2.5% to 74% spread
    • populations, dates, and authorship definitions also differ
  • stop if the differences follow directly from the chosen denominator

2: test whether detector decay could contribute to apparent plateaus

  • rescore DeGenTWeb’s known-generated frontier-model sites with frozen accessible detectors
    • compare open Pangram, Fast-DetectGPT, and affordable commercial tools
  • build on the draft’s limitations and Liang 2025
  • declining recall would weaken an adoption-plateau interpretation
    • it would not prove that adoption continued growing
  • stop if the effect disappears with a stronger frozen baseline

3: cross-check with excess vocabulary

  • apply Kobak’s method to Common Crawl prose by half-year
  • test whether excess words concentrate on DeGenTWeb-flagged sites
  • agreement is supporting evidence rather than independent ground truth
    • both methods can respond to shared topic or style shifts
  • compare historical periods and documented writing before interpreting generation

4: test aggregation for another unit

  • possible units include Reddit users, outlets, Amazon reviewers, and peer reviewers
    • reviewed leads include La Cava, Sun, Russell, Pangram, Shen, and Wang
  • choose a unit with checkable authoring histories
    • Kim et al.’s ICML self-reports are a lead
    • self-reports need their own reliability assessment
  • the human’s notes already propose set-level applications
    • this is an extension rather than a new aggregation idea

5: track search exposure across engines and years

  • combine Webis-PSERP-24, Originality.ai, Graphite, and Bevendorff’s SEO-spam methods
  • measure which generated-dominant sites appear and at what ranks
  • distinguish ranking changes from detector drift and changing site samples
  • search-engine updates are candidate interventions
    • concurrent changes limit causal interpretation

6: evaluate multilingual website detection

  • compare non-English Common Crawl prose with Thompson’s machine-translation results
  • test a multilingual scoring pair on documented human and generated sites
  • translation and original generation are different writing processes
    • calibrate them separately
  • builds on Thompson and Brooks
    • do not infer transfer from English detection alone

7: measure detectable provenance in web images

  • sample images from a defined crawl or DeGenTWeb site set
  • report C2PA, IPTC source labels, and accessible generator watermarks separately
    • missing marks do not establish human authorship
    • mark presence needs validation before becoming an origin label
  • builds on Chen, DiResta, Brooks, Hanley, and image watermarking
  • measures detectable evidence rather than total generated-image prevalence
  • no comprehensive open-web estimate was identified in this review

8: study longitudinal site switching

  • begin with DeGenTWeb’s reported switched-site set
  • compare Hanley and Russell’s findings about small publishers
  • inspect when sites change and whether traffic or ads change afterward
    • a detector transition is not itself an observed authoring transition
    • rank and traffic changes alone do not establish effects caused by generation
  • audit known histories before expanding a crawl

Last edited: