search spam: finding knowledge among spam on the web
(authored by agents unless marked đ§)
short version
- the humanâs questions: âhow to find knowledge among spam on the webâ, idea âsearch engine exclude spamâ
- what the literature settles
- search spam is a 20-year contest; the measured pattern is containment, not victory
- inference from Leontiadis 2014 and Bevendorff 2024: spam enters, an update pushes it out, it returns
- LLMs made mass production cheap; the share of generated text on the open web is now large
- how much reaches search results is disputed and depends on the detector
- this selected reading did not establish whether exclusion helps people find answers on live engines
- Cormack 2011 did it once on a frozen 2009 collection with TREC judges
- the recorded searches found no matched academic evaluation of uBlacklist, Kagi, Brave Goggles, and Marginalia
- a missed evaluation could change the proposed contribution
- search spam is a 20-year contest; the measured pattern is containment, not victory
- strongest research ideas, agent ranking
- does blocking spam fix search? run existing blocklists and quality classifiers as filters over real results and measure what users gain and lose (idea 1 below)
- do answer engines cite generated sites more than organic search does? a field version of the lab finding that rankers prefer LLM text, reusing the humanâs DeGenTWeb detector (idea 2)
- who runs the slop and who pays them? cluster generated search-result sites by ad and affiliate identifiers (idea 3)
- these are agent opinions; no prototype or novelty check beyond the searches recorded below
humanâs words and scope
- from the research notes index
- meta question: âhow to find knowledge among spam on the webâ
- idea: âsearch engine exclude spamâ
- wants: âsignificant & popular, easy sellâ, âeasy to implementâ
- from web user-facing notes, humanâs own ideas
- âranking based on user feedbackâ
- âmeasuring retrieval of Perplexity, ChatGPT, etc. in search modeâ
- this file covers web search spam and finding useful pages
- misinformation moved to misinformation.md
- bot campaigns moved to social_media_bots.md
- deeper notes on search quality, evidence checking, and 2026 RAG poisoning papers: sibling search quality study
- that study recommends a pilot on version-specific technical claims; I do not repeat it here
- how much of the web is AI text: sibling measurement review
- the humanâs own ongoing work: DeGenTWeb draft for IMC 2026
- decides per website whether most of its prose pages look LLM-written
- scores pages with Binoculars, a zero-shot LLM-text detector, then combines 20 pages per site
- false-positive rate measured on sites crawled entirely before ChatGPT
- draft abstract reports Common Crawl sites 6.0% LLM-dominant, Bing how-to result sites 16.4%, âa measurable subset of LLM-dominant search-result sites are mass-produced from shared templatesâ
- ideas 1, 2, 3, and 5 below reuse this pipeline
- decides per website whether most of its prose pages look LLM-written
what âspamâ means here
- Gyöngyi and Garcia-Molina, Web Spam Taxonomy, AIRWeb 2005
- âWeb spamming refers to actions intended to mislead search engines into ranking some pages higher than they deserveâ
- term spam, link spam, hiding (cloaking, redirection)
- Googleâs current categories, spam policies
- scaled content abuse: âproducing content at scale to boost search ranking â whether automation, humans or a combination are involvedâ
- site reputation abuse: âthird-party content is published on a host site mainly because of that hostâs already-established ranking signalsâ
- expired domain abuse
- âslopâ has no agreed definition
- Shaib, Chakrabarty, Garcia-Olano, Wallace, Measuring AI âSlopâ in Text, arXiv Sept 2025
- âthere is currently no agreed upon definition of this term nor a means to measure its occurrenceâ
- their dimensions: information utility, information quality, style quality
- inference: for research, label observable things separately
- affiliate links, ads, generated text, copied text, cloaking, usefulness to the query
- âspamâ as one label hides which failure is being measured
literature 1: the classic arms race, 2004 to 2013
- Fetterly, Manasse, Najork, Spam, Damn Spam, and Statistics, WebDB 2004
- âoutliers in the statistical distribution of these properties are highly likely to be caused by web spamâ
- finds machine-generated spam, not hand-built spam
- Gyöngyi, Garcia-Molina, Pedersen, Combating Web Spam with TrustRank, VLDB 2004
- âWe first select a small set of seed pages to be evaluated by an expertâ, then spread trust along links
- 178 seed sites after inspecting over 2,000; assumption: good sites rarely link to bad sites
- still the basis of Marginaliaâs ranking today (see tools below)
- Wu and Davison, Identifying Link Farm Spam Pages, WWW 2005
- âA link farm is a network of web sites which are densely connected with each otherâ
- Ntoulas, Najork, Manasse, Fetterly, Detecting Spam Web Pages through Content Analysis, WWW 2006
- âcorrectly identify 2,037 (86.2%) of the 2,364 spam pages (13.8%) in our judged collection of 17,168 pages, while misidentifying 526 spam and non-spam pages (3.1%)â
- features: word counts, compressibility, visible-text fraction, n-gram likelihood
- inference: 2004-era spam was statistically odd; LLM text is fluent, so these signals weaken
- Castillo, Donato, Gionis, Murdock, Silvestri, Know your Neighbors, SIGIR 2007
- âlinked hosts tend to belong to the same class: either both are spam or both are non-spamâ
- datasets WEBSPAM-UK2006/UK2007, about 6,479 labeled hosts of 114,529 (snippet only)
- Cormack, Smucker, Clarke, Efficient and Effective Spam Filtering and Re-ranking for Large Web Datasets, Information Retrieval 2011
- âa simple content-based classifier with minimal training is efficient enough to rank the âspamminessâ of every page in the dataset using a standard personal computer in 48 hoursâ
- âThe results of classical information retrieval methods are particularly enhanced by filtering â from among the worst to among the bestâ
- this is the closest existing âsearch engine exclude spamâ experiment: filter ClueWeb09, then rerun TREC
- limitation: one static collection, 2009 spam, the authorsâ own labels
- Spirin and Han, Survey on web spam detection, SIGKDD Explorations 2012
- groups methods into content, link, and user-behavior signals (snippet only)
- takeaway: content, link, and trust-seed methods all exist and all are static
- only Cormack measures retrieval quality after filtering, and only on a frozen collection with TREC judges
- none measures what a person finds on a live engine
literature 2: search poisoning measured at security venues
- Leontiadis, Moore, Christin, Measuring and Analyzing Search-Redirection Attacks, USENIX Security 2011
- âabout one third of all search results are one of over 7 000 infected hosts triggered to redirect to a few hundred pharmacy websitesâ
- âLegitimate pharmacies and health resources have been largely crowded out by search-redirection attacks and blog spamâ
- method: 218 drug queries scraped daily for nine months
- Leontiadis, Moore, Christin, A Nearly Four-Year Longitudinal Study of Search-Engine Poisoning, CCS 2014
- ârising from around 30% in late 2010 to a peak of nearly 60% in late 2012, despite efforts by search engines and browsersâ
- median clean-up time fell from about 30 days to about 15 days
- Wang, Savage, Voelker, Cloak and Dagger, CCS 2011
- âcloakers can expect to maintain their pages in search results for several days on popular search enginesâ
- cloaked results about 9.4% on Google, 7.7% on Yahoo for the studied terms
- Wang, Savage, Voelker, Juice: A Longitudinal Study of an SEO Botnet, NDSS 2013
- âthis botnet is both modest in size and has low churnâsuggesting little adversarial pressure from defendersâ
- one botnet produced 69% of poisoned trending-search results at its peak
- John, Yu, Xie, Krishnamurthy, Abadi, deSEO, USENIX Security 2011
- âThis attack employs over 5,000 compromised Web sites and poisons more than 20,000 popular search termsâ
- detects from URL patterns in search logs, no page content
- Lu, Perdisci, Lee, SURF, CCS 2011
- browser-side detector of search-then-redirect sessions, âdetection rate of 99.1% at a false positive rate of 0.9%â
- Chinese black-hat SEO line
- Yang et al., How to Learn Klingon without a Dictionary, S&P 2017: 478,879 âblack keywordsâ on Baidu
- Du et al., The Ever-changing Labyrinth, USENIX Security 2016: wildcard-DNS SEO
- Yang et al., Scalable Detection of Promotional Website Defacements, USENIX Security 2021: âfound defacements in 11% of these websitesâ across 7000+ commercial Chinese sites
- Zhang et al., Into the Dark: internal site search abused for SEO, USENIX Security 2024 (snippet only)
- Wu, Xue, Zhou, Mi, Reflected Search Poisoning for Illicit Promotion, arXiv 2024: âover 11 million distinctâ illicit promotion texts in â14 different illicit categoriesâ
- takeaway: the security community measures spam in verticals with clear harm (pharma, malware, gambling)
- the method is always the same: fixed query panel, daily scrape, classify, track over time
- we found no such study of the ordinary âlow quality but legalâ spam that now dominates complaints
- inference: that gap is where âhow to find knowledgeâ lives, and it fits IMC
literature 3: content farms and SEO in mainstream results
- McCreadie, Macdonald, Ounis, Giles, Jabr, An Examination of Content Farms in Web Search using Crowdsourcing, CIKM 2012
- âbetween the period of March and August 2011, the number of content farm articles observed on a number of indicative queries was reduced by up to 55% in the top ranksâ
- the Panda era; 4-page paper, few queries
- Lewandowski, SĂŒnkler, Yagci, The influence of search engine optimization on Googleâs results, WebSci 2021
- âa large fraction of pages found in Google is at least probably optimizedâ
- detects optimization, not harm; the humanâs notes already call its heuristics questionable
- Bevendorff, Wiegmann, Potthast, Stein, Is Google Getting Worse?, ECIR 2024, full text read
- âWe monitored Google, Bing and DuckDuckGo for a year on 7,392 product review queriesâ
- âonly a small portion of product reviews on the web uses affiliate marketing, but the majority of all search results doâ
- âhigher-ranked pages are on average more optimized, more monetized with affiliate marketing, and they show signs of lower text qualityâ
- âGoogleâs updates in particular are having a noticeable, yet mostly short-lived, effectâ
- three scenarios they test: engines losing, engines winning, or ârepeated breathing patternsâ; âit seems like (3) is the most likely scenarioâ
- âthe line between benign content and spam in the form of content and link farms becomes increasingly blurryâa situation that will surely worsen in the wake of generative AIâ
- future work: âevaluate how we can better build and evaluate truly robust web IR systems in competitive environmentsâ
- code and data; the humanâs DeGenTWeb draft already uses a âretrospective multi-engine SERP archiveâ (SERP: a search result page)
- limits: product reviews only, English only, no AI-text measurement, Bing and DuckDuckGo share an index
- Kollnig, The enshittification of online search?, arXiv Dec 2025
- 1,467 coding queries, Oct 2023; âthe quality of coding advice â as measured by the average rank of Stack Overflow â was highest on Bingâ
- crude proxy, but shows programming queries are an easy stratum with a natural ground truth
- Geraci, Identification of Web Spam through Clustering of Website Structures, WWW 2015 companion, full text read
- parked domains âhosted by the same service provider tend to have similar look-and-feelâ; clusters by page structure
- inference: template similarity is a cheap operator signal, matching the DeGenTWeb template finding
literature 4: the LLM era
4a. how much generated text is out there
- details and all studies: sibling review
- numbers disagree because detectors, thresholds, and populations differ
- Pew, How Much of the Internet Is Written With AI?, Aug 2026: â35% of pages in the July 2026 crawl with post-ChatGPT publication dates show signs of AI authorshipâ, about 1 in 10 of all pages
- Graphite, More Articles Are Now Created by AI Than Humans, Oct 2025: âIn November 2024, the quantity of AI-generated articles being published on the web surpassed the quantity of human-written articlesâ; 65,000 Common Crawl article URLs, Surfer detector
- Ahrefs, 900k new pages, May 2025: 74.2% contain some AI text, 2.5% pure AI
- DeGenTWeb draft: 6.0% of Common Crawl sites LLM-dominant, 28.6% of sites first seen in 2025H1
- content farms specifically
- NewsGuard AI Tracking Center: 3,749 âAI Content Farmâ news sites in 16 languages as of 23 June 2026; was 49 in May 2023, 840 in June 2024 (Puccetti et al. cite those)
- âthe revenue model for these websites is programmatic advertisingâ
- 2023: âNinety percent of the ads from major brands found on these AI-generated news sites were served by Googleâ (MIT Technology Review)
- Puccetti, Rogers, Alzetta, DellâOrletta, Esuli, AI âNewsâ Content Farms Are Easy to Make and Hard to Detect, ACL 2024, full text read
- fine-tune Llama on 40k Italian news articles; natives spot synthetic text âwith only 64% accuracy, vs 50% random guessâ
- âthere are currently no practical methods for detecting synthetic news-like texts âin the wildâ, while generating them is too easyâ
- Hanley and Durumeric, Machine-Made Media, ICWSM 2024
- 3,074 news sites, Jan 2022 to May 2023; synthetic articles rose 57.3% on mainstream and 474% on misinformation sites after ChatGPT
- NewsGuard AI Tracking Center: 3,749 âAI Content Farmâ news sites in 16 languages as of 23 June 2026; was 49 in May 2023, 840 in June 2024 (Puccetti et al. cite those)
4b. how much reaches search results
- Originality.ai, AI Content in Google Search Results, ongoing
- â500 Google Search keywords were chosenâ, top 20 results every two months from Internet Archive snapshots, own detector at 0.5
- âAs of September 2025, 17.31% of the top 20 search results are AI-generatedâ; peak 19.56% July 2025; 7.43% on 5 March 2024 right after Googleâs update
- vendor sells the detector; no false-positive figure on this population
- Graphite, AI content in search and answer engines, 2025
- â86% of articles ranking in Google Search are written by humans, and only 14% are generated using AIâ; âonly 7% of the articles ranking number oneâ
- ChatGPT and Perplexity citations: â82% of cited articles are written by humans, and only 18% ⊠generated using AIâ
- inference: if half of new articles are AI but 14% of ranking pages are, ranking already filters hard
- weak: the two samples differ in page age and detector, so this is a hint, not a measurement
- Ahrefs, AI Overviews cite AI content, July 2025: top-3 citations 3.6% pure AI, 8.6% pure human, 87.8% mixed
- DeGenTWeb draft: Bing how-to result sites 16.4% LLM-dominant, âa 10.4-percentage-point gap above the open-web baseline on the same classifierâ
- takeaway: result-side numbers come from detector vendors, except the humanâs draft and Allahamâs single snapshot (below)
- we found no peer-reviewed longitudinal measurement of generated text in results
- the humanâs draft is the closest; it is one engine, one query genre, and looks backward
4c. ranking systems prefer generated text in the lab
- Dai et al., Neural Retrievers are Biased Towards LLM-Generated Content, KDD 2024
- âneural retrieval models tend to rank LLM-generated documents higher. We refer to this category of biases ⊠as the source biasâ
- explanation: âLLM-generated texts exhibit more focused semantics with less noiseâ
- companion benchmark Cocktail, Findings of ACL 2024
- Chen et al., Spiral of Silence, ACL 2024
- simulated feedback loop; âLLM-generated text consistently outperforming human-authored content in search rankingsâ
- Yu, Kim, Kim (NAVER), Retrieval Collapses When AI Pollutes the Web, WWW 2026 short
- injected SEO-style AI pages into retrieval pools; 67% pool contamination gave over 80% exposure contamination
- BM25 exposed about 19% harmful content under adversarial injection; LLM rankers suppressed it better
- limits shared by all three: synthetic corpora, lab retrievers, no real engine
- inference: whether Google, Bing, or answer engines show source bias in the field is unmeasured; idea 2 below
4d. answer engines and what they cite
- Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines, 2023
- âonly 74.5% of citations support their associated sentenceâ
- Allaham and Diakopoulos, Synthetic Sources?, AIES 2026
- 712 queries, ChatGPT, Copilot, Gemini, Perplexity; âevidence of AI-generated sources being cited across all four generative search engines (~16% of cited sources)â
- single snapshot, three topics, detector error not checked in what we read
- Xu, Iqbal, Montgomery, Measuring Google AI Overviews, IMC 2026
- 55,393 trending queries, 19 categories, 40 days, March to April 2026
- activation 13.7% overall, 64.7% for question-form queries
- cited domains more credible than first page: mean 0.732 vs 0.645
- â29.8% of AIO-cited domains do not appear anywhere on the corresponding first pageâ
- 11.0% of 98,020 atomic claims unsupported by the cited pages
- âwe scope our findings to the U.S.-localized AIO experienceâ; measures credibility ratings, not generated text
- Grossman et al., How Generative AI Disrupts Search, SIGIR 2026
- 11,500 queries; âthe retrieved sources are substantially different for each search engine (<0.2 average Jaccard similarity)â; âsignificantly more likely to retrieve Google-owned contentâ
- AIO rate 51.5% vs Xuâs 13.7%: query sampling matters a lot
- Aral, Li, Zuo, The Rise of AI Search, arXiv Feb 2026
- 24,000 queries, 243 countries, 2024 and 2025; âAI search surfaces significantly fewer long tail information sources, lower response variety, and significantly more low credibility ⊠information sources, compared to traditional searchâ
- Pew, AI summary click study, July 2025
- 900 US adults, March 2025; clicked a result in 8% of visits with a summary vs 15% without
- takeaway: 2026 has a wave of answer-engine audits, all measuring credibility, overlap, or support
- only Allaham measured generated sources in peer review, once; Graphite and Ahrefs did it as vendors; Xuâs and Aralâs credibility results point opposite ways
- gap: generated-text share of citations vs organic results for the same queries, over time, with a detector whose error is known
4e. manipulation of answer engines
- covered in depth by the sibling recent work note
- Aggarwal et al., GEO, KDD 2024: âboost visibility by up to 40%â
- Nestaas, Debenedetti, TramĂšr, Adversarial SEO for LLMs, 2024: already in the humanâs notes
- Pfrommer et al., Ranking Manipulation for Conversational Search Engines, EMNLP 2024
- Zou et al., PoisonedRAG, USENIX Security 2025: â90% attack success rate when injecting five malicious textsâ
- Martinez, critical survey of GEO 2023 to 2026, arXiv July 2026, 45 studies
- no method shows âa stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behaviorâ
- inference: the GEO literature is attack demos without field prevalence; measurement, not another attack, is the open slot
literature 5: what the engines say they do
- Google, March 2024 update post
- âwe expect that the combination of this update and our previous efforts will collectively reduce low-quality, unoriginal content in search results by 40%â
- April update: âYouâll now see 45% less low-quality, unoriginal content in search resultsâ
- self-reported, no method; Originality.aiâs drop to 7.43% in March 2024 is the only outside number, and it rebounded within a year
- later spam updates on the status dashboard: August 2025, September 2026
- site reputation abuse manual actions limited outside the EEA from 30 Aug 2026 (SERoundtable)
- 2024 Content Warehouse API leak, Fishkin, SparkToro
- 14,014 attributes; click attributes such as âgoodClicksâ, âbadClicksâ, âlastLongestClicksâ (paraphrase by the fetch tool)
- Google did not confirm usage
- inference: user clicks are probably already a ranking signal, but this is a leak, not a confirmation
- Google Personal Blocklist extension, Chrome blog 2011
- Google said it would âstudy the resulting feedback and explore using it as a potential ranking signalâ (snippet only)
- precedent for the humanâs âranking based on user feedbackâ; no published evaluation
literature 6: deployed exclusion tools and their evidence limits
- uBlacklist, 6.7k stars
- âBlocks specific sites from appearing in Google search resultsâ; subscribes to public rule lists
- machine-translated Stack Exchange clone list, 958 stars
- HUGE AI image blocklist, 5.8k stars, images only
- no text content farm or âAI slopâ list found
- Kagi
- per-site block, lower, raise, pin
- public leaderboard of most blocked domains: Pinterest variants top, then Quora, Amazon, Stack Exchange mirrors
- Small Web: âa curated list of nearly 6,000 genuine websitesâ
- own indexes Teclis and TinyGem for non-commercial content; Marginalia among sources
- Brave Search Goggles
- âenable anyone, be it individuals or a community, to alter the ranking of Brave search by using a set of instructions (rules and filters)â
- rules
$boost,$downrank,$discard; Braveâs 2022 Goggles paper argues ranking secrecy is needed because openness âwould immediately result in a boost of those sites that rely on SEOâ
- Marginalia Search, one person, open source
- Trust in Ranking, Jan 2026: âA large set of trusted domains known to be high quality was selectedâ and trust flows through links
- my reading: this is the TrustRank idea; the post does not name it
- claims it âdrastically reduces the number of content farm results, as long as there are human results it usually finds them across all the usual test queriesâ
- admits âit becomes harder for new websites to establish a footholdâ, âworks poorly across language barriersâ, and âI will not share exactly which websites these are, as it paints a target on them for black-hat SEOâ
- no independent test
- Trust in Ranking, Jan 2026: âA large set of trusted domains known to be high quality was selectedâ and trust flows through links
- Stract: Goggles-like âopticsâ, archived April 2026
- OpenWebSearch.eu Open Web Index: federated crawl, about 28M hosts, research-only access; funded Bevendorffâs paper
- LLM pretraining quality filters supply web classifiers worth comparing with search-specific filters
- this reading did not establish whether that comparison is already published
- FineWeb-Edu classifier, Penedo et al. 2024: Llama-3-70B labeled 450k pages; card warns of âpotential bias toward academically-formatted materialâ
- DCLM, Li et al. 2024: âfastText OH-2.5 + ELI5 classifier score to keep the top 10% of documentsâ
- Klimaszewski and Andruszkiewicz, Is a Document Educational or Just Wikipedia-Style?, ACL 2026 short: âa straightforward Wikipedia-style reformatting operation can substantially alter a modelâs quality assessmentâ; FineWeb-Edu flips about 7% of documents
- takeaway: the âsearch engine that excludes spamâ exists in at least five forms
- the examples use lists, trusted seeds, and user preferences among their mechanisms
- this reading did not establish a matched comparison of coverage, agreement, discovery delay, and user outcomes
candidate questions and closest-work checks
- these are unresolved questions in this selected reading
- no field-wide absence or novelty claim follows
- check the nearest prior evaluations before building an experiment
- does exclusion improve correct answers as well as classifier accuracy or measured prevalence?
- do real search and answer engines reproduce the source preferences observed in laboratory retrievers?
- how do deployed blocklists and alternative engines compare under matched queries and human judgments?
- who operates the generated sites, and which revenue observations support attribution?
- can a prospective multi-engine study retain measured detector false positives as content changes?
- Bevendorff names ârobust web IR systems in competitive environmentsâ as future work
- a later paper may already have addressed it
idea 1: does blocking spam fix search?
- question: when you exclude what todayâs tools call spam, do results answer the query more often, and what do you lose?
- why this is the direct answer to âsearch engine exclude spamâ
- the tools exist, so the research is the measurement, not the engine
- popular: uBlacklist, Kagi, Brave users already do this; a quotable number sells itself
- easy to implement for the list and classifier arms: a SearXNG metasearch plug-in that drops or reranks results after the engine returns them
- not reproducible as a post-filter: Braveâs own Goggles ranking, Kagiâs per-user ranking, Marginaliaâs secret seeds
- treat those as whole-engine arms instead: query Brave with a Goggle, query Kagi, query Marginalia, compare their top 10 directly
- filters to compare
- crowd lists: uBlacklist subscription lists; Kagiâs most-blocked leaderboard as a list
- classifiers: FineWeb-Edu score, DCLM fastText score, Ntoulas-style content statistics as the 2006 baseline
- the humanâs DeGenTWeb verdict, applied to the result pageâs site
- DeGenTWeb judges sites, results are pages; a page inherits its siteâs verdict, and this must be said in the paper
- affiliate-link density from Bevendorffâs code
- Cormack 2011 style content classifier trained on a small labeled set, the closest prior art
- query panel, stratified
- how-to (reuse DeGenTWebâs Bing how-to panel), programming, product review (Bevendorffâs
best <category>form), health - size: 100 queries per stratum, 400 total
- freeze the panel; scrape Google, Bing, Brave weekly for 12 weeks, top 30 per query so filtered lists can be refilled to 10
- archive every result page as WARC (the web archive file format), with fetch time
- scraping risk: Google blocks scrapers and personalizes; use one fixed vantage, no account, and scrape each query twice per week to measure the noise floor
- how-to (reuse DeGenTWebâs Bing how-to panel), programming, product review (Bevendorffâs
- measurements
- share of top 10 removed per filter and per stratum
- agreement between filters (pairwise overlap); hypothesis: low, which itself is a finding
- lag: days between a domainâs first top-10 appearance and its entry into a crowd list
- uBlacklist lists live in git, so the entry date is the commit date; Kagiâs leaderboard has no history, so lag is measured on uBlacklist only
- hypothesis: Bevendorffâs campaign domains appear and vanish within months, so lists may arrive too late
- usefulness: do the top 5 answer the query
- 2 annotators judge 40 queries per stratum, blind to filter condition, report agreement
- programming stratum has checkable answers, the objective anchor
- an LLM judge is used only after it matches the humans on those labels, and its own bias toward fluent text is reported
- collateral: good sites removed
- ground truth is the annotatorâs âanswers the queryâ label on removed pages
- also count removed pages from domains outside the top 10,000 by Tranco, as a proxy for small sites
- hypothesis: FineWeb-Eduâs format bias removes personal blogs
- cost: latency of each filter at query time
- baselines
- unfiltered results
- random removal of the same share, to separate âremoving anythingâ from âremoving spamâ
- one page per domain; exact and near-duplicate removal
- whole alternative engines (Kagi, Marginalia, Brave with a Goggle) on the same panel
- expected quotable results, either way
- âcrowd blocklists remove X% of Googleâs top 10 for how-to queries but list a new spam domain only after Y daysâ
- âthe LLM-text filter and the affiliate filter disagree on Z% of removed pagesâ
- most dangerous objection: usefulness judgments are subjective
- pre-empt: human labels first, report agreement, objective programming stratum, random-removal baseline
- second objection: Google already filters hard, so there is no headroom
- pre-empt: that is a result, not a failure; report the removal share per stratum and the usefulness delta with confidence intervals
- third objection: using DeGenTWeb both as a filter and as the measure of harm is circular
- pre-empt: harm is the annotatorâs usefulness label only; DeGenTWeb is one filter among many
- stop rule: if after 4 weeks every filter removes under 5% of the top 10 in every stratum and usefulness does not move beyond the noise floor, write it up as a short null result and lead with idea 2
- the 5% is a guess at where the effect stops being worth a paper
idea 2: do answer engines cite generated sites more than organic search does?
- question: for the same query, is the LLM-dominant share among AI Overview, Perplexity, and ChatGPT search citations higher than among Googleâs and Bingâs organic top 10, and does the gap grow over months?
- why new
- the preference of rankers for LLM text is shown only on lab retrievers (Dai 2024, Chen 2024, Yu 2026)
- 2026 audits (Xu IMC, Grossman SIGIR, Aral) measure credibility and overlap, not generated text
- Allaham AIES 2026 measured generated citations once, no organic comparison, detector error not reported
- the human already has the detector and the Bing how-to result, so the new work is the answer-engine side plus time
- design
- the question-form subset of idea 1âs panel, because AI Overviews trigger on 64.7% of question-form queries and 13.7% overall (Xu)
- weekly, from a US vantage; answer engines are nondeterministic, so run each query 3 times and keep the union and the intersection of citations
- collect organic top 10 from Google and Bing, AI Overview citations, Perplexity citations, ChatGPT search citations
- run DeGenTWeb on every cited and ranked site; archive everything as WARC
- record whether each cited page is in the organic top 10 (Xu: 29.8% of cited domains are not)
- measurements
- LLM-dominant share: organic vs cited, per engine, per stratum, per week, computed at the page level with the site verdict, and again at the domain level so Wikipedia-size domains do not dominate
- the control is the organic top 10 for the same query, not the Common Crawl base rate, which is not conditioned on the query
- persistence: does a generated site that gets cited stay cited longer than a human one
- claim check on a small subset: are the generated citations the ones carrying unsupported claims (Xuâs claim-by-claim method, costly, 50 queries)
- quotable either way
- âanswer engines cite generated sites twice as often as organic searchâ is a headline
- âanswer engines cite generated sites less than organic searchâ contradicts the retrieval-collapse story and is equally a headline
- confounds to state
- engines have different indexes; a citation gap is a difference in exposure, not a proof of a ranker preference
- citations are few per query against 10 organic slots, so compare shares, not counts, and report per-query paired differences
- most dangerous objection: detector drift
- DeGenTWebâs false-positive rate is measured on pre-ChatGPT sites; its recall falls to 0.58 on sites generated by 2025 frontier models
- so misses, not false alarms, are the problem, and misses may differ between the organic and cited arms
- pre-empt: re-calibrate monthly on a fresh synthetic set, report the gap under the measured recall as a range, add template and monetization signals that do not depend on text statistics
- feasibility: Perplexity and ChatGPT search access terms may forbid automated querying; use their APIs where offered and record the terms
- stop rule: if cited and organic shares match within the noise floor after 8 weeks, fold the result into idea 1âs paper as one section
idea 3: follow the money
- question: who operates the generated sites in search results, how concentrated is it, and how much search exposure do they get?
- dollars are out of reach without ad-network data; exposure (how often they appear in top 10) is measurable
- why new
- Papadogiannakis et al. (WWW 2023, WWW 2022) cluster sites by AdSense and analytics IDs, but for fake news, not search spam
- Bevendorff found Amazon Associates dominates affiliate spam but did not cluster operators
- the humanâs draft already finds âa narrow entry-level monetization signatureâ and shared templates; this makes it an operator study
- NewsGuard and ANA numbers are vendor or industry; we found no academic estimate
- design
- start from the LLM-dominant sites found in ideas 1 and 2 plus the Common Crawl sample
- extract AdSense publisher IDs (the
ca-pub-strings in page source), analytics IDs, Amazon Associates tags, shared page templates (Geraci-style structure hashes), shared hosting and DNS - cluster into operators
- agencies, ad networks, and CMS themes can share IDs across operators
- one operator can also use several IDs
- clusters can therefore merge distinct operators or split one operator
- require independent corroborating signals and report uncertain attribution
- control: run the same clustering on a matched set of human-written sites from the same result panels, so concentration can be compared
- agencies, ad networks, and CMS themes can share IDs across operators
- estimate each clusterâs exposure from the weekly panels; public traffic estimates are unreliable for small sites, so use them only as a secondary check
- check whether Googleâs ad network serves ads on the pages Google ranks: NewsGuard found 90% of brand ads on AI news farms were served by Google
- quotable either way
- âN operators run M% of the generated sites in top-10 resultsâ with a concentration curve
- âGoogleâs ad network serves ads on X% of the generated pages Google ranksâ; inference: this is the number a policy argument would use
- most dangerous objection: IDs are visible for only part of the sites, so clusters undercount
- pre-empt: report the covered fraction; validate clusters on the template clusters DeGenTWeb already found and on the affiliate campaigns in the ECIR-24 data, which are a partial check, not ground truth
- ethics: name operators only in aggregate; publish cluster statistics, not a list of people
- feasibility: parsing IDs is cheap; the clustering validation and the human-site control are the real work; it depends on idea 1 or 2 having produced the site set
idea 4: crowd feedback as a ranking signal
- the humanâs idea: âranking based on user feedbackâ
- what exists: Googleâs 2011 blocklist experiment, Kagiâs votes, and the leaked click attributes
- this selected reading did not recover a matched public outcome evaluation
- candidate question: how many colluding voters flip a domain, and does weighting by voter history preserve useful results?
- effectiveness needs outcome evidence as well as resistance to collusion
- needs users or a simulation with planted colluders, so it is not easy to implement
- SybilGuard is already in the humanâs notes
- recommendation: use the Kagi leaderboard as one list in idea 1 now, and leave collusion resistance for later
idea 5: is Google getting slopier?
- forward-looking, multi-engine, multi-genre version of Bevendorff with the humanâs detector
- partly done: the humanâs draft covers Bing how-to plus the retrospective search-result archive; Originality.ai publishes a vendor version every two months
- inference: alone it is an incremental measurement; it is the backbone that ideas 1 and 2 share, so build it once and report it inside them
what I would do first
- build the weekly panel scraper and archive once; ideas 1, 2, 3, and 5 all read from it
- run DeGenTWeb and the cheap filters over the first four weeks
- decide after week 4 which of ideas 1 and 2 has the larger effect and lead with it
review limits
- sources read in full: Bevendorff ECIR 2024, Puccetti ACL 2024, McCreadie CIKM 2012, Geraci WWW 2015, the DeGenTWeb draft abstract
- other quotes come from abstracts and pages fetched on 7 Oct 2026 by three search subagents; items marked âsnippet onlyâ were not verified against the source
- not found despite searching: an academic study of users appending âredditâ to queries; an academic measurement of made-for-advertising sites; a peer-reviewed study of expired-domain SEO abuse; any evaluation of blocklists or pretraining filters on search usefulness
- ideas 1 to 3 remain candidate experiments
- unsuccessful searches do not establish novelty
- check the closest prior methods and evaluations before claiming a contribution
- numbers from Originality.ai, Graphite, Ahrefs, NewsGuard, and Google are vendor or self-reported and are not comparable with each other
- the 7 October topic-specific ChatGPT attempt failed with âaccount_ui_login_requiredâ
- a context-free Claude reviewer checked this file instead; its objections are folded into the ideas above
- the 8 October cross-topic consultation completed with Extra High selected
- its selected scope does not certify every search-spam proposal
Last edited: