web crawling in the age of language models (authored by agents unless marked đ§)
- written 2026-10-07
- quotes are verbatim from the abstract, the body, or the web page I link
- source claims and my inferences are distinguished
- blog posts and vendor reports are marked as such
- they are the only source for most traffic numbers, and their underlying logs are not independently available
- novelty checks are incomplete
- search coverage limits appear at the end
- this file starts where the robots.txt section of crawling.md stops
- papers already quoted there get one line here
the picture in plain words
- three kinds of machine now fetch web pages for language models, and sites treat them very differently
- training crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, CCBot): bulk download, no visitor comes back
- search crawlers (OAI-SearchBot, PerplexityBot): build an index that an assistant answers from
- user fetchers and agents (ChatGPT-User, Perplexity-User, Operator, browser agents): fetch one page because one person asked, sometimes through a real browser
- what people try to find out
- how much load these machines put on a site, and who pays
- whether they do what robots.txt says
- what site owners do about it, and whether it works
- what the machine gets back: the same page a person gets, a block page, a paywall price, a maze of junk, or text written to mislead a model
- what assistants fetch and cite at answer time, and how old it is
- how they do it
- read robots.txt files of many sites over time (Common Crawl, the Wayback Machine, own crawls)
- read server logs of sites they run, or set up bait sites and change robots.txt on purpose
- plant secret strings in pages and ask the chatbots about them
- fetch the same URL under different user agents and compare
- ask assistants thousands of questions and collect the links they cite
- my take
- at least seven reviewed studies measure robots.txt blocking
- those rules express intent rather than actual access
- the effect side is thin: one universityâs logs for 40 days is the best public compliance data, and the large traffic estimates reviewed come from vendors selling bot control
- the defenses that spread fastest in 2025 and 2026 (proof of work, junk mazes, HTTP 402 pricing) have no measurement paper that I could find, only blog posts arguing both ways
- the reviewed literature does not measure these page differences across many sites
- the attack papers show it can be done, the measurement papers only count blocks
- at least seven reviewed studies measure robots.txt blocking
the rules: robots.txt and what is being added to it
- A Standard for Robot Exclusion, Koster, 1994 (web page)
- âIt is not an official standard backed by a standards body, or owned by any commercial organisation. It is not enforced by anybody, and there no guarantee that all current and future robots will use it.â
- RFC 9309: Robots Exclusion Protocol, Koster, Illyes, Zeller, Sassman, IETF, 2022
- âThese rules are not a form of access authorization.â
- fact, the parts a crawler can get wrong: âIf the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallowâ
- crawlers âSHOULD NOT use the cached version for more than 24 hours, unless the robots.txt file is unreachableâ
- open: the file can say which paths, per crawler name
- it cannot say what the content may be used for, and it cannot check that the crawler is who it says
- Determining Bias to Search Engines from Robots.txt, Sun, Zhuang, Councill, Giles, WI 2007 (short version at WWW 2007)
- âinvestigated 7,593 websites covering education, government, news, and business domains, and collected 2,925 distinct robots.txt filesâ
- âthe robots of popular search engines and information portals, such as Google, Yahoo, and MSN, are generally favored by most of the websitesâ
- why it matters now: the same ârich get richerâ worry, with Googlebot allowed and GPTBot blocked
- âinvestigated 7,593 websites covering education, government, news, and business domains, and collected 2,925 distinct robots.txt filesâ
- A Larger Scale Study of Robots.txt, Kolay, DâAlberto, Dasdan, Bhattacharjee, WWW 2008 (poster)
- âabout 2.2M non-empty robots.txt files from 6M sitesâ
- âsome sites may have bias towards specific crawlers but overall the top two crawlers seem to have access to the same amount of content. This is contrary to the point made by the previous study.â
- lesson: counting rules per crawler name and counting content actually reachable give different answers
- most 2024 to 2026 studies only do the first
- âabout 2.2M non-empty robots.txt files from 6M sitesâ
- A Survey of Web Content Control for Generative AI, Dinzinger, HeĂ, Granitzer, arXiv 2024
- site owners âare overwhelmed by the multitude of recent ad hoc standards to considerâ
- a catalog of the opt-out formats (robots.txt extensions, meta tags, TDM reservation, ai.txt and others) with the legal background
- IETF AI Preferences working group (aipref), drafts, not yet RFCs when I looked
- A Vocabulary For Expressing AI Usage Preferences: âdefines a vocabulary for expressing preferences regarding how digital assets are used by automated processing systemsâ
- Associating AI Usage Preferences with Content in HTTP: âdefines attachment methods using the Robots Exclusion Protocol and HTTP header fieldsâ
- âThis document updates RFC 9309â
- example in the draft: âContent-Usage: train-ai=nâ
- I did not verify the current revision numbers or milestone dates
- inference: this adds âused for whatâ to robots.txt
- the groupâs charter leaves enforcement and crawler identity out, so the honor system stays
- Web Bot Authentication working group (webbotauth), IETF charter
- âwill standardize methods for cryptographically authenticating automated clients and providing additional information about their operators to Web sitesâ
- the starting draft is HTTP Message Signatures for automated traffic Architecture: a bot signs its requests, and the site checks the key
- inference: this fixes âis this really GPTBotâ
- it does nothing about a crawler that prefers to look like Chrome
- Content Signals Policy, Cloudflare, 2025 (vendor blog)
- âdefines three content signals - search, ai-input, and ai-trainâ
- written as comments plus a line in robots.txt
- competes with the IETF vocabulary
- I found no count of how many sites use either
- âdefines three content signals - search, ai-input, and ai-trainâ
- The /llms.txt file, Howard, 2024 (proposal page)
- âA proposal to standardise on using an /llms.txt file to provide information to help agents use a website.â
- this one invites machines in (a short markdown guide to the site) instead of keeping them out
- adoption numbers exist only in marketing reports that I read second hand (ppc.land summary of Originality.ai and Ahrefs: files grew several fold in a year, and almost none get requested by AI crawlers)
- not verified
- proposals from researchers, none deployed as far as I know
- ai.txt: A Domain-Specific Language for Guiding AI Interactions with the Internet, Li et al., arXiv 2025: âenabling precise element-level regulations and incorporating natural language instructions interpretable by AI systemsâ
- Permission Manifests for Web Agents, Marro et al., arXiv 2025: âagent-permissions.json, a robots.txt-style lightweight manifest where websites specify allowed interactionsâ
- motive: âwebsite owners increasingly rely on blanket blocking and CAPTCHAsâ
- terms.txt: A Consent and Compensation Protocol for Agentic Web Access, Chowdhury, arXiv 2026: robots.txt âcannot express identity, purpose, terms, or priceâ
- a reference implementation âadds 0.20 to 0.65 ms per request on one vCPUâ
- Tag Your Fish in the Broken Net, Zhang et al., arXiv 2023: per-item consent tags in HTTP and HTML plus a ledger for withdrawal
- Will the Agent Recuse, and Will It Stop?, Munirathinam, arXiv 2026: the same honor-system idea for SSH and databases
- ârecusal to deny ranges from 100% to 55-75% among agents that received itâ
- âreliably stopping a running agent needs enforcement, not a requestâ
- law, briefly: The Liabilities of Robots.txt, Chang, He, 2025 argues the file âcan give rise to a unilateral contract or serve as a form of notice sufficient to establish tortious liabilityâ
- In the Mood to Exclude, Atkinson, 2025 argues for trespass to chattels
who blocks AI crawlers in robots.txt
- already in crawling.md: Longpre 2024 (â28%+ of the most actively maintained, critical sources in C4, fully restrictedâ), Bouchaud 2025 (a quarter of the top thousand, â9.5% disallow CCBotâ), Steinacker-Olsztyn 2025 (60.0% of reputable news versus 9.1% of misinformation sites)
- How many news websites block AI crawlers?, Fletcher, Reuters Institute, 2024 (factsheet)
- âBy the end of 2023, 48% of the most widely used news websites across ten countries were blocking OpenAIâs crawlers. A smaller number, 24%, were blocking Googleâs AI crawler.â
- blocking of OpenAI ranged âfrom 79% in the USA to just 20% in Mexico and Polandâ
- Somesite I Used To Crawl, Liu, Luo, Shan, Voelker, Zhao, Savage, IMC 2025
- beyond what crawling.md quotes: the people who want to block often cannot
- of 203 artists, â59% have never heard about robots.txtâ
- hosted site builders often do not let them edit it
- âA small but growing number of websites also explicitly invite AI crawlers to crawl their content.â
- How Do Data Owners Say No?, Lee et al., arXiv 2025
- fact, for an image and text training set: â60% of the samples in the top 50 domains come from websites with ToS that prohibit scrapingâ
- lesson: owners say no in terms of service, copyright notices and watermarks too, and crawlers read none of those
- Cloudflare Radar 2025 Year in Review, Cloudflare, 2025 (vendor report)
- âThe user agents with the highest number of fully disallowed directives are those associated with AI crawlers, including GPTBot, ClaudeBot, and CCBot.â
- does blocking cost the site anything
- Strategic Response of News Publishers to Generative AI, Zhao, Berman, arXiv 2025: âlarge publishers who block GenAI bots experience reduced website traffic compared to not blockingâ
- body: âa 7% post-blocking decline in weekly visits measured by SimilarWeb or Semrush within the 6 weeks after blockingâ
- How Generative AI Disrupts Search, Grossman et al., SIGIR 2026: âwebsites that block Googleâs AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the contentâ
- open: both are correlations around a choice the site made
- neither study randomizes blocking
- Strategic Response of News Publishers to Generative AI, Zhao, Berman, arXiv 2025: âlarge publishers who block GenAI bots experience reduced website traffic compared to not blockingâ
- what this line of work leaves open
- robots.txt rules are a wish
- the papers below show the wish and the outcome differ
- sites that block at the firewall and leave robots.txt alone are invisible to all of these counts (Liu: âmany sites indeed use active blocking as their soleâ mechanism)
- robots.txt rules are a wish
do crawlers obey
- Scrapers selectively respect robots.txt directives, Kim, Bock, Luo, Liswood, Poroslay, Wenger, IMC 2025
- what they did: logs of 36 sites at one university, three robots.txt versions of rising strictness, â130 self-declared bots (and many anonymous ones) over 40 daysâ
- âBots are less likely to respect robots.txt that employ strict directivesâ
- âSEO bots are most respectful of robots.txt, while search engine crawlers are among the least. AI-specific bots like AI assistants and AI data scrapers, fall in between.â
- some non-compliance âcan sometimes be attributed to spoofing, in which malicious bots present a false user agentâ
- open: one institution, 40 days
- nothing on bots that hide their name
- nothing on the newer signals (402, Content-Signal, signed requests)
- Liu et al. (above), on their own test sites: âmost large AI companies currently do respect robots.txt. However, a number of AI-powered apps and crawlers do not respect it (including crawlers from ByteDance)â
- Do Generative AI Assistants Respect robots.txt?, Lopez-Fonseca, Rodriguez, Bechtold, Del Alamo, arXiv 2026
- what they did: ten assistants, four robots.txt conditions, âserver-side logs and secret codes embedded in target pagesâ over 200 trials
- some âaccessed restricted resources without requesting robots.txt or used generic user-agents that complicated attributionâ
- âassistants may access pages without surfacing the retrieved content, or fail to access even allowed resourcesâ
- fact from the body, a trap for anyone repeating this: an assistant âmay instead answer from an intermediate layer such as a search index, cached copy, or other preprocessedâ version, so no request reaches the test server at all
- Identifying AI Web Scrapers Using Canary Tokens, Seiden, Ren, Zhang, Kim, Liu, Wenger, arXiv 2026
- what they did: âhost dynamic websites that serve unique canary tokens to each visiting scraper, then prompt LLMs for information about our sitesâ
- across â22 production LLM systemsâ the method âcan reliably identify which scrapers feed which LLM, including several that are not publicly known or disclosed by the companiesâ
- why I like it: it links a log line to a modelâs answer without any help from the company
- it is the one new measurement trick in this area
- open: tokens were plain page text
- nothing about content that needs JavaScript, a click, or a login
- AI Search Has A Citation Problem, JaĆșwiĆska, Chandrasekar, Tow Center, 2025 (journalism study)
- 1,600 queries over eight chatbots
- âPlatforms retrieved information from publishers that had intentionally blocked their crawlersâ
- Perplexityâs free version âcorrectly identified all ten excerpts from paywalled articles we shared from National Geographic, even though the publisher has disallowed Perplexityâs crawlersâ
- their own caveat: there are âother means through which the chatbots could obtain information about restricted contentâ
- 1,600 queries over eight chatbots
- Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, Cloudflare, 2025 (vendor blog, one-sided)
- claim: âwhen they are presented with a network block, they appear to obscure their crawling identityâ
- âobserved across tens of thousands of domains and millions of requests per dayâ
- claim: âwhen they are presented with a network block, they appear to obscure their crawling identityâ
- older baseline: Good Bot, Bad Bot: Characterizing Automated Browsing Activity, Li, Azad, Rahmati, Nikiforakis, S&P 2021
- â100 dedicated honeysitesâ for seven months, â26.4 million requests sent by more than 287K unique IP addressesâ
- comparing claimed identity with TLS and HTTP fingerprints exposes bots that lie about who they are
- the honeysite design is what Kim, Seiden and Lopez-Fonseca reuse at small scale
- â100 dedicated honeysitesâ for seven months, â26.4 million requests sent by more than 287K unique IP addressesâ
how much traffic, and what it costs the site
- vendor numbers (Cloudflare sees its own customers only; âAI botâ means bots it could name)
- Radar 2025 Year in Review: âtraffic from AI bots accounted for an average of 4.2% of HTML requestsâ
- âGooglebot alone accounted for 4.5%â
- âCrawling for model training is responsible for the overwhelming majority of AI crawler traffic, reaching as much as 7-8x search crawling and 32x user action crawling at peakâ
- user action crawling was âup over 21x from January through early Decemberâ
- From Googlebot to GPTBot: whoâs crawling your site in 2025: GPTBot âsurging from 5% to 30% shareâ of AI crawling in a year, Bytespider âplummeted from 42% to 7%â
- The crawl before the fall⊠of referrals: for one week in June 2025 âthe ratios range from Anthropicâs 70,900:1 down to Mistralâs 0.1:1â (pages crawled per visitor sent back)
- caveat in the post: âtraffic referred by Claudeâs native app does not include a Referer: headerâ
- The rise of the AI crawler, Vercel, 2024: GPTBot â569 million requests across Vercelâs network in the past monthâ, ClaudeBot 370 million, together âabout 20% of Googlebotâs 4.5 billionâ
- inference: 4% of page requests does not sound like a crisis
- the operator reports below explain why the average hides the damage
- Radar 2025 Year in Review: âtraffic from AI bots accounted for an average of 4.2% of HTML requestsâ
- operator reports: the cost sits in the long tail of uncached, expensive pages
- How crawlers impact the operations of the Wikimedia projects, Wikimedia Foundation, 2025: âSince January 2024, we have seen the bandwidth used for downloading multimedia content grow by 50%â
- âat least 65% of this resource-consuming traffic we get for the website is coming from bots, a disproportionate amount given the overall pageviews from bots are about 35% of the totalâ
- their reason: âcrawler bots tend to âbulk readâ larger numbers of pages and visit also the less popular pagesâ, which miss the cache and hit the core datacenter
- AI crawlers need to be more respectful, Read the Docs, 2024: âOne crawler downloaded 73 TB of zipped HTML files in May 2024, with almost 10 TB in a single day. This cost us over 0â
- Who does Anubis actually stop?, Zakaria, 2026 (blog): âFor a scraper, solving the Anubis challenge is a one-time, amortized-to-zero cost since the cookie can be cached and reusedâ
- text browsers, screen readers and feed readers âare completely left outâ
- inference: operators say it works, two engineers show the arithmetic says it should not
- the likely answer is that it stops crawlers that do not run JavaScript at all, which is a different claim from âmakes crawling expensiveâ, the reviewed sources do not measure that effect
- How crawlers impact the operations of the Wikimedia projects, Wikimedia Foundation, 2025: âSince January 2024, we have seen the bandwidth used for downloading multimedia content grow by 50%â
- junk mazes
- AI Labyrinth, Cloudflare, 2025 (vendor blog): âuses AI-generated content to slow down, confuse, and waste the resources of AI Crawlers and other bots that donât respect âno crawlâ directivesâ
- doubles as a detector: âNo real human would go four links deep into a maze of AI-generated nonsenseâ
- open source tarpits (Nepenthes, iocaine) do the same without the detector
- I have no primary quote for them
- inference for the humanâs DeGenTWeb line: defenders now publish generated pages on purpose, on real domains, aimed at crawlers
- a corpus built by a crawler that looks like a bot will contain some
- AI Labyrinth, Cloudflare, 2025 (vendor blog): âuses AI-generated content to slow down, confuse, and waste the resources of AI Crawlers and other bots that donât respect âno crawlâ directivesâ
- text aimed at the model instead of the crawler
- Indirect Prompt Injection in the Wild, Khodayari, Zhang, Acharya, Pellegrino, arXiv 2026: âAnalyzing 1.2B URLs from 24.8M hosts, we identify 15.3K validated instances across 11.7K pagesâ
- uses include âcontent-protection directives, and AI-bot detectionâ
- âabout 70% appear in non-rendered HTMLâ
- models obey rarely, âup to 8% for smaller models on plain-text inputsâ
- mostly Common Crawl data, so it counts what a polite named crawler was shown
- Indirect Prompt Injection in the Wild, Khodayari, Zhang, Acharya, Pellegrino, arXiv 2026: âAnalyzing 1.2B URLs from 24.8M hosts, we identify 15.3K validated instances across 11.7K pagesâ
- charging
- Introducing pay per crawl, Cloudflare, 2025 (vendor blog): crawlers âeither present payment intent via request headers for successful access (HTTP response code 200), or receive a 402 Payment Required response with pricingâ
- one âflat, per-request price across their entire siteâ
- Pay-Per-Crawl Pricing for AI: The LM-Tree Agent, Archer, Ghili, Haghpanah, arXiv 2026: learns per-article prices
- â8,939 articles and 80,451 buyer queries with willingness-to-pay calibrated from actual AI crawler trafficâ
- â65% revenue gain over a single static priceâ
- open: no public data on how many crawlers pay, or what a 402 does to crawl behavior
- Introducing pay per crawl, Cloudflare, 2025 (vendor blog): crawlers âeither present payment intent via request headers for successful access (HTTP response code 200), or receive a 402 Payment Required response with pricingâ
research and archive crawlers caught in the same net
- in crawling.md: Gundelach 2026 (âChromium headless encounters a 15% soft block rateâ), and the bias argument from Bouchaud and Steinacker-Olsztyn
- Web Crawl Refusals: Insights From Common Crawl, Ansar, Sperotto, Holz, PAM 2025
- what they did: regular expressions over page bodies in a Common Crawl snapshot to find refusal pages that a status code alone would miss
- âat least 1.68% of sites in a CC snapshot exhibit a form of explicit refusalâ
- âan inconsistent and even incorrect use of HTTP status codes to indicate refusalsâ
- âmost blocks resolve within one hour, but also that 80% of refusing domains block every request by CCâ
- open: âearly-stage workâ, one snapshot
- no trend, and it predates Anubis and the 2025 Cloudflare defaults
- the Internet Archive
- News publishers limit Internet Archive access due to AI scraping concerns, Nieman Lab, 2026 (journalism): â241 news sites from nine countries explicitly disallow at least one out of the four Internet Archive crawling botsâ
- the Times: âthe Wayback Machine provides unfettered access to Times content â including by AI companies â without authorizationâ
- follow-up, May 2026: âmore than 340 local news sites across the United States are now limiting the Internet Archiveâs ability to access and preserve their storiesâ
- âNo news publisher has confirmed to Nieman Lab that an AI company has already scraped their content from the Wayback Machineâ
- their caveat: âThis data is not comprehensive, but exploratoryâ
- News publishers limit Internet Archive access due to AI scraping concerns, Nieman Lab, 2026 (journalism): â241 news sites from nine countries explicitly disallow at least one out of the four Internet Archive crawling botsâ
- COAR (above): repositories that block bots are âalso inadvertently blocking other desired network services such as scholarly aggregators, indexing services, and directoriesâ (this sentence is from the search summary of the COAR page; I did not re-read it in the page text)
- review observation: two source pages served Anubis challenges to a plain HTTP client
- inference: a web archive or research corpus collected in 2026 is missing a different, larger and less random slice of the web than one from 2022, the reviewed studies do not estimate that change
the web that machines get versus the web people get
- classic cloaking, where the machine was the search crawler
- Cloak and Dagger: Dynamics of Web Search Cloaking, Wang, Savage, Voelker, CCS 2011: crawler âDaggerâ fetched each result as a crawler and as a browser âfor over five months, identifying when distinct results were provided to crawlers and browsersâ
- Cloak of Visibility: Detecting When Machines Browse A Different Web, Invernizzi et al., S&P 2016: bought âten cloaking packages that range in price from 13,188â
- they range from PHP plugins âthat check the User-Agent of incoming clientsâ to web servers that blacklist âbased on IP addresses, reverse DNS, User-Agentsâ
- cloaking toward agents, so far only shown as an attack
- A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See, Zychlinski, arXiv 2025: âA malicious website can identify an incoming request as originating from an AI agent and dynamically serve a different, âcloakedâ version of its contentâ
- a concept paper, no measurement
- A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See, Zychlinski, arXiv 2025: âA malicious website can identify an incoming request as originating from an AI agent and dynamically serve a different, âcloakedâ version of its contentâ
- legitimate âdifferent page for machinesâ is now a product
- Building an open Agentic Internet, Cloudflare, 2026 (vendor blog): âMarkdown for Agents lets agents read websites with fewer tokens and less bandwidthâ
- the agent asks with an Accept header and gets markdown made from the HTML
- together with 402 responses, challenge pages and mazes, a site can now return five or more different things for one URL depending on who seems to ask
- Building an open Agentic Internet, Cloudflare, 2026 (vendor blog): âMarkdown for Agents lets agents read websites with fewer tokens and less bandwidthâ
- what the crawlerâs software can see at all
- Vercel (above, vendor log analysis): ânone of the major AI crawlers currently render JavaScriptâ
- they âdo fetch JavaScript files (ChatGPT: 11.50%, Claude: 23.84% of requests), they donât execute themâ
- âCommon Crawl (CCBot) ⊠does not render pagesâ
- Gemini and AppleBot do render
- this is late 2024 and one hosting network
- I found no newer or independent check
- relevant to the JSphere line: text that only exists after scripts run is missing from training crawls, while browser agents do see it
- Vercel (above, vendor log analysis): ânone of the major AI crawlers currently render JavaScriptâ
- what agents do with the human web (lab studies, mostly cloned or instrumented sites)
- Machine-Readable Ads, Nitu, MĂŒhle, Stöckl, arXiv 2025: agents ânever scroll beyond two viewports and ignore purely visual calls to actionâ
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents, Ersoy et al., S&P 2026: âwhen there is a single dark pattern present, agents are susceptible to it an average of 41% of the timeâ
- SusBench, Guo et al., IUI 2026: dark patterns injected into 55 live sites
- âboth human participants and agents are particularly susceptible to the dark patterns of Preselection, Trick Wording, and Hidden Informationâ
- Build the web for agents, not agents for the web, LĂč, Kamath, Mosbach, Reddy, arXiv 2025 (position): proposes âan Agentic Web Interface (AWI), an interface specifically designed for agents to navigate a websiteâ
- if this happens, the machine web and the human web split by design
crawlers that use a language model
- writing the scraper: AutoScraper, Huang et al., EMNLP 2024: âthe paradigm of generating web scrapers with LLMsâ
- the model writes extraction rules once per site instead of reading every page
- reading the page
- ReaderLM-v2, Wang et al., arXiv 2025: a 1.5B parameter model âtransforming messy HTML into clean Markdown or JSONâ
- HtmlRAG, Tan et al., WWW 2025: keep pruned HTML instead of plain text, because âmuch of the structural and semantic information inherent in HTML, such as headings and table structures, is lostâ
- choosing what to crawl: Craw4LLM (above)
- using an agent as the measurement crawler
- On the Suitability of LLM-Driven Agents for Dark Pattern Audits, Sun, Vekaria, Nithyanand, arXiv 2026: an agent walks data-rights request forms on â456 data broker websitesâ
- they report âthe reliability and reproducibility of its dark pattern classificationsâ and where it fails
- crawling.md idea 4 covers the general version (agent that clicks consent and logs in), and cites the security scanner work
- On the Suitability of LLM-Driven Agents for Dark Pattern Audits, Sun, Vekaria, Nithyanand, arXiv 2026: an agent walks data-rights request forms on â456 data broker websitesâ
- inference: the tools exist, but I found no paper that compares what an LLM crawler collects against Heritrix or a scripted browser on the same sites, with cost and repeatability
what assistants fetch and cite at answer time
- the user-facing side (are citations correct, can they be gamed) is in the SEO folder
- here I keep what bears on fetching
- which sources
- Search Arena: Analyzing Search-Augmented LLMs, Miroyan et al., ICLR 2026: âover 24,000 paired multi-turn user interactions with search-augmented LLMsâ with full traces, open data
- âuser preferences are influenced by the number of citations, even when the cited content does not directly support the attributed claimsâ
- News Source Citing Patterns in AI Search Systems, Yang, arXiv 2025, same data: âAmong the over 366,000 citations embedded in these responses, 9% reference news sourcesâ
- âNews citations concentrate heavily among a small number of outletsâ
- Grossman et al. (above): for â51.5% of representative, real-user queries, AIOs are generatedâ
- sources differ a lot between Google search, AI Overviews and Gemini (â<0.2 average Jaccard similarityâ)
- generative search is âsignificantly more likely to retrieve Google-owned contentâ
- Navigating the Shift, Chen, Wang, Chen, Koudas, EDBT/ICDT workshops 2026: AI answers and Google results âdiverge significantly in their consulted source domains ⊠and the freshness of the information providedâ
- they also study how âpre-training ⊠interacts with and influences real-time web search when enabledâ
- From Citation Selection to Citation Absorption, Kai, Xinyue, Jingang, arXiv 2026: â21,143 valid search-layer citationsâ
- âPerplexity and Google cite more sources on average, while ChatGPT cites fewer sources but shows substantially higher average citation influenceâ
- Synthetic Sources?, Allaham, Diakopoulos, arXiv 2026: âevidence of AI-generated sources being cited across all four generative search engines (~16% of cited sources)â
- Generative AI Search Engines as Arbiters of Public Knowledge, Li, Sinnamon, arXiv 2024: early audit, âcommercial and geographic bias in sourcesâ
- Search Arena: Analyzing Search-Augmented LLMs, Miroyan et al., ICLR 2026: âover 24,000 paired multi-turn user interactions with search-augmented LLMsâ with full traces, open data
- how fresh
- Dated Data: Tracing Knowledge Cutoffs in Large Language Models, Cheng et al., arXiv 2024: âeffective cutoffs often differ from reported cutoffsâ, partly from âtemporal biases of CommonCrawl data due to non-trivial amounts of old data in new dumpsâ
- Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs, Mansoor, Ahmad, Yoon, arXiv 2026: a benchmark with âstaleness labels drawn from real web snapshots at 1, 12, 24 hours, and 7 daysâ
- âonly 34.3% of detected content changes actually affect answer correctnessâ
- that last number matters for the web atoms idea: most page changes do not change any answer, so âchangedâ needs a definition tied to use
- inference: every study here looks at the answer and its links
- only Lopez-Fonseca and Seiden look at the requests
- the reviewed studies do not join the two at scale to say when the page behind a citation was last fetched
known pitfalls and unsolved problems
- a user agent string is a claim
- Kim et al. had to filter spoofers by network
- Cloudflare accuses Perplexity of dropping its name when blocked
- Lopez-Fonseca found assistants on âgeneric user-agentsâ
- so âGPTBot trafficâ means ârequests that say GPTBot and come from the right addressesâ
- these counts omit crawlers that successfully hide their identity
- Kim et al. had to filter spoofers by network
- the list of AI crawler names is crowd-sourced and changes monthly
- two studies with different lists get different block rates for the same file
- robots.txt says what the owner wants
- firewall rules, CDN defaults and challenge pages decide what happens
- most studies read only the first
- every large traffic number comes from a company that sells bot control, covers only its own customers, and classifies with a private method
- origin logs are private
- the one academic log study covers one university
- an assistant that answers from an index never touches the test server, so âno request seenâ does not mean ârespected the blockâ
- the crawl-to-refer ratio misses visitors that arrive without a Referer header (apps), by Cloudflareâs own note
- lab studies of agents use cloned or injected pages
- how real sites treat real agents is unmeasured
- ethics of the measuring itself: testing compliance means running bait sites and prompting commercial chatbots at volume
- testing defenses means sending fake crawler identities at other peopleâs servers
- things move fast: Bytespider went from 42% to 7% share in a year
- a snapshot paper is stale before it is published
research ideas
- one URL, many answers: how sites change the response by who seems to ask
- question: for the same URL, how do the responses differ between a browser, a named AI training crawler, a named AI search crawler, a user fetcher, and a browser agent, and how common is each kind of difference (block, challenge, 402 with a price, markdown version, junk maze, text aimed at the model, a quietly different article)
- why not answered
- Wang 2011 and Invernizzi 2016 did this for search crawlers
- Steinacker-Olsztyn 2025 and Liu 2025 swap the user agent but only record blocked or not, on news sites and the top 10k
- Zychlinski 2025 shows agent cloaking as an attack without measuring it
- Khodayari 2026 counts injected text but in what Common Crawl was served
- Ansar 2025 classifies refusal pages for one crawler in one snapshot
- what we would build: a differential fetcher that requests each URL under several identities close together in time, repeats the browser fetch to learn the pageâs normal churn, then classifies the differences
- start with a Tranco sample plus internal pages
- data and tools: Tranco, the block-page patterns from Ansar and Gundelach, the Dagger design, the humanâs Common Crawl and content-comparison experience from DeGenTWeb
- main risk: we can copy a crawlerâs name but not its addresses or signatures, and sites that verify (Cloudflare verified bots, Web Bot Auth) will treat us as an impostor
- so we measure âwhat a claimed GPTBot getsâ, which is still what researchers and small crawlers get
- also sending false identities needs an ethics argument
- confidence the gap is real: medium to high
- I could not finish the search for a 2026 paper doing this, and it is an obvious idea, so check IMC 2026 and USENIX Security 2026 programs first
- a census of the new defenses, and who they lock out
- question: how many sites deploy proof-of-work challenges, junk mazes, 402 pricing and AI-specific blocking, how fast is that growing, and which clients besides AI crawlers lose access (archives, research crawlers, text browsers, feed readers, screen readers)
- why not answered: Liu 2025 inferred one Cloudflare switch on the top 10k in 2024
- Ansar 2025 found 1.68% refusing Common Crawl in one snapshot
- Ormandy counted Anubis deployments once in a blog post
- Nieman Lab counted Internet Archive blocks on a news list by reading robots.txt
- nothing tracks these over time or across client types
- what we would measure: fingerprint each defense by its challenge page and headers
- scan a site list monthly with several client types
- mine past Common Crawl snapshots for challenge pages to get the history for free
- compare Wayback Machine capture success before and after a site adopts a defense
- data and tools: Common Crawl indexes record status and bodies per snapshot
- Anubis and similar tools are open source so their pages are easy to recognize
- the refusal patterns from Ansar
- main risk: Common Crawl only shows what CCBot was served
- mazes are built to be hard to tell from real pages
- confidence the gap is real: medium
- this overlaps crawling.md idea 1 (effect of blocking on published numbers), so run them as one project with this as the âwho deploys whatâ half
- capability canaries: what the whole pipeline can reach
- question: which production assistants know content that is only reachable by running JavaScript, clicking a consent banner, scrolling, solving a proof-of-work challenge, paying a 402, or ignoring robots.txt
- why not answered: Seiden 2026 plants one token per visiting scraper in plain page text
- Lopez-Fonseca 2026 varies only robots.txt
- Vercel 2024 inferred âno JavaScriptâ from which files crawlers fetch, on one network, two years ago
- what we would build: bait sites where each token sits behind exactly one barrier
- then ask the assistants, as Seiden does
- the answer shows which barriers each companyâs pipeline crosses, with no need to trust user agents
- why it fits us: the barrier list is the JSphere question turned around (which page content depends on which browser features), and the result says how much of the script-built web is absent from model training
- data and tools: Seidenâs method
- their two-month wait for crawlers to arrive sets the timeline
- main risk: tokens on new, unlinked sites may never be crawled or may be dropped in filtering, so a missing token proves little
- needs many sites and positive controls
- confidence the gap is real: medium
- the Seiden authors are the obvious people to do this next
- do AI crawlers revalidate, and how stale are answers
- question: when AI crawlers come back to a page, do they send conditional requests (If-None-Match, If-Modified-Since) and honor 304 and sitemap dates, how much of their load is refetching unchanged pages, and how long after a page changes does an assistantâs answer change
- why not answered: Kim 2025 measures crawl-delay and disallow, not revalidation
- the Read the Docs complaint about missing ETag support is one anecdote
- Craw4LLM cuts waste by picking pages, not by skipping unchanged ones
- the freshness cache paper works on the assistantâs side
- Lopez-Fonseca notes that assistants answer from indexes but does not time them
- what we would measure: on sites we control, log conditional headers and recrawl gaps per crawler against known change times
- plant dated changes and poll assistants until the answer flips
- estimate bytes that a change-group hint (the web atoms idea) would save
- data and tools: honeysite designs from Li 2021 and Kim 2025
- the atom grouping from web_change_and_atoms.md
- main risk: bait sites get little crawler attention
- partnering with a real mid-size site (a university, Read the Docs, a library) fixes that but needs their logs
- confidence the gap is real: medium to high for the revalidation part, medium for answer latency, since marketing firms publish rough versions
- a conformance test for crawler rules, old and new
- question: beyond âdo they obey Disallowâ, do AI crawlers implement the rest of RFC 9309 (treat 5xx on robots.txt as full disallow, refresh within 24 hours, match the most specific group and path) and the newer signals (Content-Signal lines, Content-Usage headers, 402 with a price, valid Web Bot Auth signatures)
- why not answered: Kim 2025 tests three directive types at one institution and puts caching out of scope
- Lopez-Fonseca 2026 tests four allow and disallow conditions
- the newer signals are too new to have studies that I found
- what we would build: a public test suite of hostnames, each exercising one rule, with logs published on a schedule
- a scoreboard per crawler
- main risk: scoreboards attract gaming and lawyers
- crawlers that hide their name are untestable by design
- confidence the gap is real: medium
- cheap to start, and it reuses the bait sites from ideas 3 and 4
- an open, multi-site log panel for crawler load
- question: what do AI crawlers cost origin servers, measured the same way across many sites: bytes, share of uncached and expensive requests, and how that differs by site type
- why not answered: Wikimedia, Read the Docs, GLAM-E and COAR each report their own pain in their own units
- Kim 2025 has one university
- Cloudflare reports shares, not costs, and only for customers
- what we would build: a small log-sharing agreement with sites that already complain (libraries, repositories, open source forges), a common anonymization and bot-labeling pipeline, and a cost model per request type
- main risk: privacy review and getting anyone to share logs
- this is mostly organizing work, less a technical problem
- confidence the gap is real: high that the data does not exist publicly, low that we are the right people to collect it
opinions from ChatGPT
- no separate ChatGPT consultation for this crawler subreview
- group consultation records the completed LLM-text consultation
gaps in this review
- search coverage is incomplete, so 2026 conference papers that are not on arXiv are likely missing
- check the IMC 2026, USENIX Security 2026 and WWW 2026 programs for AI crawler measurement before starting any idea above
- known but not read or quoted: Dinzinger and Granitzer, âA Longitudinal Study of Content Control Mechanismsâ (WWW Companion 2024)
- Fastly, TollBit, Imperva and Akamai bot reports
- the Code4Lib 2025 article on crawler floods at UNC Libraries (unread)
- Nepenthes and iocaine project pages
- OpenAI, Anthropic and Perplexity crawler documentation
- thin: older server-log studies of crawler behavior from 2000 to 2015
- residential proxy networks used by scrapers
- licensing deals between publishers and AI companies
- the EU AI Act code of practice and its robots.txt clause
- MCP, WebMCP and other agent-facing site interfaces
- read as abstract only: the ai.txt, dark pattern, ad, ReaderLM, HtmlRAG, citation and freshness-cache papers
- AIPREF and Web Bot Auth draft status changes monthly
- I quoted the draftsâ text but did not track revisions
Last edited: