Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

web crawling in the age of language models (authored by agents unless marked 🧑)

  • written 2026-10-07
    • quotes are verbatim from the abstract, the body, or the web page I link
    • source claims and my inferences are distinguished
    • blog posts and vendor reports are marked as such
      • they are the only source for most traffic numbers, and their underlying logs are not independently available
    • novelty checks are incomplete
      • search coverage limits appear at the end
    • this file starts where the robots.txt section of crawling.md stops
      • papers already quoted there get one line here

the picture in plain words

  • three kinds of machine now fetch web pages for language models, and sites treat them very differently
    • training crawlers (GPTBot, ClaudeBot, Meta-ExternalAgent, Bytespider, CCBot): bulk download, no visitor comes back
    • search crawlers (OAI-SearchBot, PerplexityBot): build an index that an assistant answers from
    • user fetchers and agents (ChatGPT-User, Perplexity-User, Operator, browser agents): fetch one page because one person asked, sometimes through a real browser
  • what people try to find out
    • how much load these machines put on a site, and who pays
    • whether they do what robots.txt says
    • what site owners do about it, and whether it works
    • what the machine gets back: the same page a person gets, a block page, a paywall price, a maze of junk, or text written to mislead a model
    • what assistants fetch and cite at answer time, and how old it is
  • how they do it
    • read robots.txt files of many sites over time (Common Crawl, the Wayback Machine, own crawls)
    • read server logs of sites they run, or set up bait sites and change robots.txt on purpose
    • plant secret strings in pages and ask the chatbots about them
    • fetch the same URL under different user agents and compare
    • ask assistants thousands of questions and collect the links they cite
  • my take
    • at least seven reviewed studies measure robots.txt blocking
      • those rules express intent rather than actual access
    • the effect side is thin: one university’s logs for 40 days is the best public compliance data, and the large traffic estimates reviewed come from vendors selling bot control
    • the defenses that spread fastest in 2025 and 2026 (proof of work, junk mazes, HTTP 402 pricing) have no measurement paper that I could find, only blog posts arguing both ways
    • the reviewed literature does not measure these page differences across many sites
      • the attack papers show it can be done, the measurement papers only count blocks

the rules: robots.txt and what is being added to it

  • A Standard for Robot Exclusion, Koster, 1994 (web page)
    • “It is not an official standard backed by a standards body, or owned by any commercial organisation. It is not enforced by anybody, and there no guarantee that all current and future robots will use it.”
  • RFC 9309: Robots Exclusion Protocol, Koster, Illyes, Zeller, Sassman, IETF, 2022
    • “These rules are not a form of access authorization.”
    • fact, the parts a crawler can get wrong: “If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow”
      • crawlers “SHOULD NOT use the cached version for more than 24 hours, unless the robots.txt file is unreachable”
    • open: the file can say which paths, per crawler name
      • it cannot say what the content may be used for, and it cannot check that the crawler is who it says
  • Determining Bias to Search Engines from Robots.txt, Sun, Zhuang, Councill, Giles, WI 2007 (short version at WWW 2007)
    • “investigated 7,593 websites covering education, government, news, and business domains, and collected 2,925 distinct robots.txt files”
      • “the robots of popular search engines and information portals, such as Google, Yahoo, and MSN, are generally favored by most of the websites”
    • why it matters now: the same “rich get richer” worry, with Googlebot allowed and GPTBot blocked
  • A Larger Scale Study of Robots.txt, Kolay, D’Alberto, Dasdan, Bhattacharjee, WWW 2008 (poster)
    • “about 2.2M non-empty robots.txt files from 6M sites”
      • “some sites may have bias towards specific crawlers but overall the top two crawlers seem to have access to the same amount of content. This is contrary to the point made by the previous study.”
    • lesson: counting rules per crawler name and counting content actually reachable give different answers
      • most 2024 to 2026 studies only do the first
  • A Survey of Web Content Control for Generative AI, Dinzinger, Heß, Granitzer, arXiv 2024
    • site owners “are overwhelmed by the multitude of recent ad hoc standards to consider”
    • a catalog of the opt-out formats (robots.txt extensions, meta tags, TDM reservation, ai.txt and others) with the legal background
  • IETF AI Preferences working group (aipref), drafts, not yet RFCs when I looked
    • A Vocabulary For Expressing AI Usage Preferences: “defines a vocabulary for expressing preferences regarding how digital assets are used by automated processing systems”
    • Associating AI Usage Preferences with Content in HTTP: “defines attachment methods using the Robots Exclusion Protocol and HTTP header fields”
      • “This document updates RFC 9309”
      • example in the draft: “Content-Usage: train-ai=n”
    • I did not verify the current revision numbers or milestone dates
    • inference: this adds “used for what” to robots.txt
      • the group’s charter leaves enforcement and crawler identity out, so the honor system stays
  • Web Bot Authentication working group (webbotauth), IETF charter
    • “will standardize methods for cryptographically authenticating automated clients and providing additional information about their operators to Web sites”
    • the starting draft is HTTP Message Signatures for automated traffic Architecture: a bot signs its requests, and the site checks the key
    • inference: this fixes “is this really GPTBot”
      • it does nothing about a crawler that prefers to look like Chrome
  • Content Signals Policy, Cloudflare, 2025 (vendor blog)
    • “defines three content signals - search, ai-input, and ai-train”
      • written as comments plus a line in robots.txt
    • competes with the IETF vocabulary
      • I found no count of how many sites use either
  • The /llms.txt file, Howard, 2024 (proposal page)
    • “A proposal to standardise on using an /llms.txt file to provide information to help agents use a website.”
    • this one invites machines in (a short markdown guide to the site) instead of keeping them out
    • adoption numbers exist only in marketing reports that I read second hand (ppc.land summary of Originality.ai and Ahrefs: files grew several fold in a year, and almost none get requested by AI crawlers)
      • not verified
  • proposals from researchers, none deployed as far as I know
  • law, briefly: The Liabilities of Robots.txt, Chang, He, 2025 argues the file “can give rise to a unilateral contract or serve as a form of notice sufficient to establish tortious liability”

who blocks AI crawlers in robots.txt

  • already in crawling.md: Longpre 2024 (“28%+ of the most actively maintained, critical sources in C4, fully restricted”), Bouchaud 2025 (a quarter of the top thousand, “9.5% disallow CCBot”), Steinacker-Olsztyn 2025 (60.0% of reputable news versus 9.1% of misinformation sites)
  • How many news websites block AI crawlers?, Fletcher, Reuters Institute, 2024 (factsheet)
    • “By the end of 2023, 48% of the most widely used news websites across ten countries were blocking OpenAI’s crawlers. A smaller number, 24%, were blocking Google’s AI crawler.”
    • blocking of OpenAI ranged “from 79% in the USA to just 20% in Mexico and Poland”
  • Somesite I Used To Crawl, Liu, Luo, Shan, Voelker, Zhao, Savage, IMC 2025
    • beyond what crawling.md quotes: the people who want to block often cannot
    • of 203 artists, “59% have never heard about robots.txt”
      • hosted site builders often do not let them edit it
    • “A small but growing number of websites also explicitly invite AI crawlers to crawl their content.”
  • How Do Data Owners Say No?, Lee et al., arXiv 2025
    • fact, for an image and text training set: “60% of the samples in the top 50 domains come from websites with ToS that prohibit scraping”
    • lesson: owners say no in terms of service, copyright notices and watermarks too, and crawlers read none of those
  • Cloudflare Radar 2025 Year in Review, Cloudflare, 2025 (vendor report)
    • “The user agents with the highest number of fully disallowed directives are those associated with AI crawlers, including GPTBot, ClaudeBot, and CCBot.”
  • does blocking cost the site anything
    • Strategic Response of News Publishers to Generative AI, Zhao, Berman, arXiv 2025: “large publishers who block GenAI bots experience reduced website traffic compared to not blocking”
      • body: “a 7% post-blocking decline in weekly visits measured by SimilarWeb or Semrush within the 6 weeks after blocking”
    • How Generative AI Disrupts Search, Grossman et al., SIGIR 2026: “websites that block Google’s AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the content”
    • open: both are correlations around a choice the site made
      • neither study randomizes blocking
  • what this line of work leaves open
    • robots.txt rules are a wish
      • the papers below show the wish and the outcome differ
    • sites that block at the firewall and leave robots.txt alone are invisible to all of these counts (Liu: “many sites indeed use active blocking as their sole” mechanism)

do crawlers obey

  • Scrapers selectively respect robots.txt directives, Kim, Bock, Luo, Liswood, Poroslay, Wenger, IMC 2025
    • what they did: logs of 36 sites at one university, three robots.txt versions of rising strictness, “130 self-declared bots (and many anonymous ones) over 40 days”
    • “Bots are less likely to respect robots.txt that employ strict directives”
      • “SEO bots are most respectful of robots.txt, while search engine crawlers are among the least. AI-specific bots like AI assistants and AI data scrapers, fall in between.”
    • some non-compliance “can sometimes be attributed to spoofing, in which malicious bots present a false user agent”
    • open: one institution, 40 days
      • nothing on bots that hide their name
      • nothing on the newer signals (402, Content-Signal, signed requests)
  • Liu et al. (above), on their own test sites: “most large AI companies currently do respect robots.txt. However, a number of AI-powered apps and crawlers do not respect it (including crawlers from ByteDance)”
  • Do Generative AI Assistants Respect robots.txt?, Lopez-Fonseca, Rodriguez, Bechtold, Del Alamo, arXiv 2026
    • what they did: ten assistants, four robots.txt conditions, “server-side logs and secret codes embedded in target pages” over 200 trials
    • some “accessed restricted resources without requesting robots.txt or used generic user-agents that complicated attribution”
      • “assistants may access pages without surfacing the retrieved content, or fail to access even allowed resources”
    • fact from the body, a trap for anyone repeating this: an assistant “may instead answer from an intermediate layer such as a search index, cached copy, or other preprocessed” version, so no request reaches the test server at all
  • Identifying AI Web Scrapers Using Canary Tokens, Seiden, Ren, Zhang, Kim, Liu, Wenger, arXiv 2026
    • what they did: “host dynamic websites that serve unique canary tokens to each visiting scraper, then prompt LLMs for information about our sites”
    • across “22 production LLM systems” the method “can reliably identify which scrapers feed which LLM, including several that are not publicly known or disclosed by the companies”
    • why I like it: it links a log line to a model’s answer without any help from the company
      • it is the one new measurement trick in this area
    • open: tokens were plain page text
      • nothing about content that needs JavaScript, a click, or a login
  • AI Search Has A Citation Problem, JaĆșwiƄska, Chandrasekar, Tow Center, 2025 (journalism study)
    • 1,600 queries over eight chatbots
      • “Platforms retrieved information from publishers that had intentionally blocked their crawlers”
      • Perplexity’s free version “correctly identified all ten excerpts from paywalled articles we shared from National Geographic, even though the publisher has disallowed Perplexity’s crawlers”
    • their own caveat: there are “other means through which the chatbots could obtain information about restricted content”
  • Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, Cloudflare, 2025 (vendor blog, one-sided)
    • claim: “when they are presented with a network block, they appear to obscure their crawling identity”
      • “observed across tens of thousands of domains and millions of requests per day”
  • older baseline: Good Bot, Bad Bot: Characterizing Automated Browsing Activity, Li, Azad, Rahmati, Nikiforakis, S&P 2021
    • “100 dedicated honeysites” for seven months, “26.4 million requests sent by more than 287K unique IP addresses”
      • comparing claimed identity with TLS and HTTP fingerprints exposes bots that lie about who they are
    • the honeysite design is what Kim, Seiden and Lopez-Fonseca reuse at small scale

how much traffic, and what it costs the site

  • vendor numbers (Cloudflare sees its own customers only; “AI bot” means bots it could name)
    • Radar 2025 Year in Review: “traffic from AI bots accounted for an average of 4.2% of HTML requests”
      • “Googlebot alone accounted for 4.5%”
      • “Crawling for model training is responsible for the overwhelming majority of AI crawler traffic, reaching as much as 7-8x search crawling and 32x user action crawling at peak”
      • user action crawling was “up over 21x from January through early December”
    • From Googlebot to GPTBot: who’s crawling your site in 2025: GPTBot “surging from 5% to 30% share” of AI crawling in a year, Bytespider “plummeted from 42% to 7%”
    • The crawl before the fall
 of referrals: for one week in June 2025 “the ratios range from Anthropic’s 70,900:1 down to Mistral’s 0.1:1” (pages crawled per visitor sent back)
      • caveat in the post: “traffic referred by Claude’s native app does not include a Referer: header”
    • The rise of the AI crawler, Vercel, 2024: GPTBot “569 million requests across Vercel’s network in the past month”, ClaudeBot 370 million, together “about 20% of Googlebot’s 4.5 billion”
    • inference: 4% of page requests does not sound like a crisis
      • the operator reports below explain why the average hides the damage
  • operator reports: the cost sits in the long tail of uncached, expensive pages
    • How crawlers impact the operations of the Wikimedia projects, Wikimedia Foundation, 2025: “Since January 2024, we have seen the bandwidth used for downloading multimedia content grow by 50%”
      • “at least 65% of this resource-consuming traffic we get for the website is coming from bots, a disproportionate amount given the overall pageviews from bots are about 35% of the total”
      • their reason: “crawler bots tend to ‘bulk read’ larger numbers of pages and visit also the less popular pages”, which miss the cache and hit the core datacenter
    • AI crawlers need to be more respectful, Read the Docs, 2024: “One crawler downloaded 73 TB of zipped HTML files in May 2024, with almost 10 TB in a single day. This cost us over 0”
    • Who does Anubis actually stop?, Zakaria, 2026 (blog): “For a scraper, solving the Anubis challenge is a one-time, amortized-to-zero cost since the cookie can be cached and reused”
      • text browsers, screen readers and feed readers “are completely left out”
    • inference: operators say it works, two engineers show the arithmetic says it should not
      • the likely answer is that it stops crawlers that do not run JavaScript at all, which is a different claim from “makes crawling expensive”, the reviewed sources do not measure that effect
  • junk mazes
    • AI Labyrinth, Cloudflare, 2025 (vendor blog): “uses AI-generated content to slow down, confuse, and waste the resources of AI Crawlers and other bots that don’t respect ‘no crawl’ directives”
      • doubles as a detector: “No real human would go four links deep into a maze of AI-generated nonsense”
    • open source tarpits (Nepenthes, iocaine) do the same without the detector
      • I have no primary quote for them
    • inference for the human’s DeGenTWeb line: defenders now publish generated pages on purpose, on real domains, aimed at crawlers
      • a corpus built by a crawler that looks like a bot will contain some
  • text aimed at the model instead of the crawler
    • Indirect Prompt Injection in the Wild, Khodayari, Zhang, Acharya, Pellegrino, arXiv 2026: “Analyzing 1.2B URLs from 24.8M hosts, we identify 15.3K validated instances across 11.7K pages”
      • uses include “content-protection directives, and AI-bot detection”
      • “about 70% appear in non-rendered HTML”
      • models obey rarely, “up to 8% for smaller models on plain-text inputs”
    • mostly Common Crawl data, so it counts what a polite named crawler was shown
  • charging
    • Introducing pay per crawl, Cloudflare, 2025 (vendor blog): crawlers “either present payment intent via request headers for successful access (HTTP response code 200), or receive a 402 Payment Required response with pricing”
      • one “flat, per-request price across their entire site”
    • Pay-Per-Crawl Pricing for AI: The LM-Tree Agent, Archer, Ghili, Haghpanah, arXiv 2026: learns per-article prices
      • “8,939 articles and 80,451 buyer queries with willingness-to-pay calibrated from actual AI crawler traffic”
      • “65% revenue gain over a single static price”
    • open: no public data on how many crawlers pay, or what a 402 does to crawl behavior

research and archive crawlers caught in the same net

  • in crawling.md: Gundelach 2026 (“Chromium headless encounters a 15% soft block rate”), and the bias argument from Bouchaud and Steinacker-Olsztyn
  • Web Crawl Refusals: Insights From Common Crawl, Ansar, Sperotto, Holz, PAM 2025
    • what they did: regular expressions over page bodies in a Common Crawl snapshot to find refusal pages that a status code alone would miss
    • “at least 1.68% of sites in a CC snapshot exhibit a form of explicit refusal”
      • “an inconsistent and even incorrect use of HTTP status codes to indicate refusals”
      • “most blocks resolve within one hour, but also that 80% of refusing domains block every request by CC”
    • open: “early-stage work”, one snapshot
      • no trend, and it predates Anubis and the 2025 Cloudflare defaults
  • the Internet Archive
    • News publishers limit Internet Archive access due to AI scraping concerns, Nieman Lab, 2026 (journalism): “241 news sites from nine countries explicitly disallow at least one out of the four Internet Archive crawling bots”
      • the Times: “the Wayback Machine provides unfettered access to Times content — including by AI companies — without authorization”
    • follow-up, May 2026: “more than 340 local news sites across the United States are now limiting the Internet Archive’s ability to access and preserve their stories”
      • “No news publisher has confirmed to Nieman Lab that an AI company has already scraped their content from the Wayback Machine”
    • their caveat: “This data is not comprehensive, but exploratory”
  • COAR (above): repositories that block bots are “also inadvertently blocking other desired network services such as scholarly aggregators, indexing services, and directories” (this sentence is from the search summary of the COAR page; I did not re-read it in the page text)
  • review observation: two source pages served Anubis challenges to a plain HTTP client
  • inference: a web archive or research corpus collected in 2026 is missing a different, larger and less random slice of the web than one from 2022, the reviewed studies do not estimate that change

the web that machines get versus the web people get

  • classic cloaking, where the machine was the search crawler
    • Cloak and Dagger: Dynamics of Web Search Cloaking, Wang, Savage, Voelker, CCS 2011: crawler “Dagger” fetched each result as a crawler and as a browser “for over five months, identifying when distinct results were provided to crawlers and browsers”
    • Cloak of Visibility: Detecting When Machines Browse A Different Web, Invernizzi et al., S&P 2016: bought “ten cloaking packages that range in price from 13,188”
      • they range from PHP plugins “that check the User-Agent of incoming clients” to web servers that blacklist “based on IP addresses, reverse DNS, User-Agents”
  • cloaking toward agents, so far only shown as an attack
  • legitimate “different page for machines” is now a product
    • Building an open Agentic Internet, Cloudflare, 2026 (vendor blog): “Markdown for Agents lets agents read websites with fewer tokens and less bandwidth”
      • the agent asks with an Accept header and gets markdown made from the HTML
    • together with 402 responses, challenge pages and mazes, a site can now return five or more different things for one URL depending on who seems to ask
  • what the crawler’s software can see at all
    • Vercel (above, vendor log analysis): “none of the major AI crawlers currently render JavaScript”
      • they “do fetch JavaScript files (ChatGPT: 11.50%, Claude: 23.84% of requests), they don’t execute them”
      • “Common Crawl (CCBot) 
 does not render pages”
      • Gemini and AppleBot do render
    • this is late 2024 and one hosting network
      • I found no newer or independent check
    • relevant to the JSphere line: text that only exists after scripts run is missing from training crawls, while browser agents do see it
  • what agents do with the human web (lab studies, mostly cloned or instrumented sites)
    • Machine-Readable Ads, Nitu, MĂŒhle, Stöckl, arXiv 2025: agents “never scroll beyond two viewports and ignore purely visual calls to action”
    • Investigating the Impact of Dark Patterns on LLM-Based Web Agents, Ersoy et al., S&P 2026: “when there is a single dark pattern present, agents are susceptible to it an average of 41% of the time”
    • SusBench, Guo et al., IUI 2026: dark patterns injected into 55 live sites
      • “both human participants and agents are particularly susceptible to the dark patterns of Preselection, Trick Wording, and Hidden Information”
  • Build the web for agents, not agents for the web, LĂč, Kamath, Mosbach, Reddy, arXiv 2025 (position): proposes “an Agentic Web Interface (AWI), an interface specifically designed for agents to navigate a website”
    • if this happens, the machine web and the human web split by design

crawlers that use a language model

  • writing the scraper: AutoScraper, Huang et al., EMNLP 2024: “the paradigm of generating web scrapers with LLMs”
    • the model writes extraction rules once per site instead of reading every page
  • reading the page
    • ReaderLM-v2, Wang et al., arXiv 2025: a 1.5B parameter model “transforming messy HTML into clean Markdown or JSON”
    • HtmlRAG, Tan et al., WWW 2025: keep pruned HTML instead of plain text, because “much of the structural and semantic information inherent in HTML, such as headings and table structures, is lost”
  • choosing what to crawl: Craw4LLM (above)
  • using an agent as the measurement crawler
    • On the Suitability of LLM-Driven Agents for Dark Pattern Audits, Sun, Vekaria, Nithyanand, arXiv 2026: an agent walks data-rights request forms on “456 data broker websites”
      • they report “the reliability and reproducibility of its dark pattern classifications” and where it fails
    • crawling.md idea 4 covers the general version (agent that clicks consent and logs in), and cites the security scanner work
  • inference: the tools exist, but I found no paper that compares what an LLM crawler collects against Heritrix or a scripted browser on the same sites, with cost and repeatability

what assistants fetch and cite at answer time

  • the user-facing side (are citations correct, can they be gamed) is in the SEO folder
    • here I keep what bears on fetching
  • which sources
    • Search Arena: Analyzing Search-Augmented LLMs, Miroyan et al., ICLR 2026: “over 24,000 paired multi-turn user interactions with search-augmented LLMs” with full traces, open data
      • “user preferences are influenced by the number of citations, even when the cited content does not directly support the attributed claims”
    • News Source Citing Patterns in AI Search Systems, Yang, arXiv 2025, same data: “Among the over 366,000 citations embedded in these responses, 9% reference news sources”
      • “News citations concentrate heavily among a small number of outlets”
    • Grossman et al. (above): for “51.5% of representative, real-user queries, AIOs are generated”
      • sources differ a lot between Google search, AI Overviews and Gemini (“<0.2 average Jaccard similarity”)
      • generative search is “significantly more likely to retrieve Google-owned content”
    • Navigating the Shift, Chen, Wang, Chen, Koudas, EDBT/ICDT workshops 2026: AI answers and Google results “diverge significantly in their consulted source domains 
 and the freshness of the information provided”
      • they also study how “pre-training 
 interacts with and influences real-time web search when enabled”
    • From Citation Selection to Citation Absorption, Kai, Xinyue, Jingang, arXiv 2026: “21,143 valid search-layer citations”
      • “Perplexity and Google cite more sources on average, while ChatGPT cites fewer sources but shows substantially higher average citation influence”
    • Synthetic Sources?, Allaham, Diakopoulos, arXiv 2026: “evidence of AI-generated sources being cited across all four generative search engines (~16% of cited sources)”
    • Generative AI Search Engines as Arbiters of Public Knowledge, Li, Sinnamon, arXiv 2024: early audit, “commercial and geographic bias in sources”
  • how fresh
  • inference: every study here looks at the answer and its links
    • only Lopez-Fonseca and Seiden look at the requests
    • the reviewed studies do not join the two at scale to say when the page behind a citation was last fetched

known pitfalls and unsolved problems

  • a user agent string is a claim
    • Kim et al. had to filter spoofers by network
      • Cloudflare accuses Perplexity of dropping its name when blocked
      • Lopez-Fonseca found assistants on “generic user-agents”
    • so “GPTBot traffic” means “requests that say GPTBot and come from the right addresses”
      • these counts omit crawlers that successfully hide their identity
  • the list of AI crawler names is crowd-sourced and changes monthly
    • two studies with different lists get different block rates for the same file
  • robots.txt says what the owner wants
    • firewall rules, CDN defaults and challenge pages decide what happens
    • most studies read only the first
  • every large traffic number comes from a company that sells bot control, covers only its own customers, and classifies with a private method
  • origin logs are private
    • the one academic log study covers one university
  • an assistant that answers from an index never touches the test server, so “no request seen” does not mean “respected the block”
  • the crawl-to-refer ratio misses visitors that arrive without a Referer header (apps), by Cloudflare’s own note
  • lab studies of agents use cloned or injected pages
    • how real sites treat real agents is unmeasured
  • ethics of the measuring itself: testing compliance means running bait sites and prompting commercial chatbots at volume
    • testing defenses means sending fake crawler identities at other people’s servers
  • things move fast: Bytespider went from 42% to 7% share in a year
    • a snapshot paper is stale before it is published

research ideas

  1. one URL, many answers: how sites change the response by who seems to ask
  • question: for the same URL, how do the responses differ between a browser, a named AI training crawler, a named AI search crawler, a user fetcher, and a browser agent, and how common is each kind of difference (block, challenge, 402 with a price, markdown version, junk maze, text aimed at the model, a quietly different article)
  • why not answered
    • Wang 2011 and Invernizzi 2016 did this for search crawlers
    • Steinacker-Olsztyn 2025 and Liu 2025 swap the user agent but only record blocked or not, on news sites and the top 10k
    • Zychlinski 2025 shows agent cloaking as an attack without measuring it
      • Khodayari 2026 counts injected text but in what Common Crawl was served
    • Ansar 2025 classifies refusal pages for one crawler in one snapshot
  • what we would build: a differential fetcher that requests each URL under several identities close together in time, repeats the browser fetch to learn the page’s normal churn, then classifies the differences
    • start with a Tranco sample plus internal pages
  • data and tools: Tranco, the block-page patterns from Ansar and Gundelach, the Dagger design, the human’s Common Crawl and content-comparison experience from DeGenTWeb
  • main risk: we can copy a crawler’s name but not its addresses or signatures, and sites that verify (Cloudflare verified bots, Web Bot Auth) will treat us as an impostor
    • so we measure “what a claimed GPTBot gets”, which is still what researchers and small crawlers get
    • also sending false identities needs an ethics argument
  • confidence the gap is real: medium to high
    • I could not finish the search for a 2026 paper doing this, and it is an obvious idea, so check IMC 2026 and USENIX Security 2026 programs first
  1. a census of the new defenses, and who they lock out
  • question: how many sites deploy proof-of-work challenges, junk mazes, 402 pricing and AI-specific blocking, how fast is that growing, and which clients besides AI crawlers lose access (archives, research crawlers, text browsers, feed readers, screen readers)
  • why not answered: Liu 2025 inferred one Cloudflare switch on the top 10k in 2024
    • Ansar 2025 found 1.68% refusing Common Crawl in one snapshot
    • Ormandy counted Anubis deployments once in a blog post
    • Nieman Lab counted Internet Archive blocks on a news list by reading robots.txt
    • nothing tracks these over time or across client types
  • what we would measure: fingerprint each defense by its challenge page and headers
    • scan a site list monthly with several client types
    • mine past Common Crawl snapshots for challenge pages to get the history for free
    • compare Wayback Machine capture success before and after a site adopts a defense
  • data and tools: Common Crawl indexes record status and bodies per snapshot
    • Anubis and similar tools are open source so their pages are easy to recognize
    • the refusal patterns from Ansar
  • main risk: Common Crawl only shows what CCBot was served
    • mazes are built to be hard to tell from real pages
  • confidence the gap is real: medium
    • this overlaps crawling.md idea 1 (effect of blocking on published numbers), so run them as one project with this as the “who deploys what” half
  1. capability canaries: what the whole pipeline can reach
  • question: which production assistants know content that is only reachable by running JavaScript, clicking a consent banner, scrolling, solving a proof-of-work challenge, paying a 402, or ignoring robots.txt
  • why not answered: Seiden 2026 plants one token per visiting scraper in plain page text
    • Lopez-Fonseca 2026 varies only robots.txt
    • Vercel 2024 inferred “no JavaScript” from which files crawlers fetch, on one network, two years ago
  • what we would build: bait sites where each token sits behind exactly one barrier
    • then ask the assistants, as Seiden does
    • the answer shows which barriers each company’s pipeline crosses, with no need to trust user agents
  • why it fits us: the barrier list is the JSphere question turned around (which page content depends on which browser features), and the result says how much of the script-built web is absent from model training
  • data and tools: Seiden’s method
    • their two-month wait for crawlers to arrive sets the timeline
  • main risk: tokens on new, unlinked sites may never be crawled or may be dropped in filtering, so a missing token proves little
    • needs many sites and positive controls
  • confidence the gap is real: medium
    • the Seiden authors are the obvious people to do this next
  1. do AI crawlers revalidate, and how stale are answers
  • question: when AI crawlers come back to a page, do they send conditional requests (If-None-Match, If-Modified-Since) and honor 304 and sitemap dates, how much of their load is refetching unchanged pages, and how long after a page changes does an assistant’s answer change
  • why not answered: Kim 2025 measures crawl-delay and disallow, not revalidation
    • the Read the Docs complaint about missing ETag support is one anecdote
    • Craw4LLM cuts waste by picking pages, not by skipping unchanged ones
    • the freshness cache paper works on the assistant’s side
    • Lopez-Fonseca notes that assistants answer from indexes but does not time them
  • what we would measure: on sites we control, log conditional headers and recrawl gaps per crawler against known change times
    • plant dated changes and poll assistants until the answer flips
    • estimate bytes that a change-group hint (the web atoms idea) would save
  • data and tools: honeysite designs from Li 2021 and Kim 2025
  • main risk: bait sites get little crawler attention
    • partnering with a real mid-size site (a university, Read the Docs, a library) fixes that but needs their logs
  • confidence the gap is real: medium to high for the revalidation part, medium for answer latency, since marketing firms publish rough versions
  1. a conformance test for crawler rules, old and new
  • question: beyond “do they obey Disallow”, do AI crawlers implement the rest of RFC 9309 (treat 5xx on robots.txt as full disallow, refresh within 24 hours, match the most specific group and path) and the newer signals (Content-Signal lines, Content-Usage headers, 402 with a price, valid Web Bot Auth signatures)
  • why not answered: Kim 2025 tests three directive types at one institution and puts caching out of scope
    • Lopez-Fonseca 2026 tests four allow and disallow conditions
    • the newer signals are too new to have studies that I found
  • what we would build: a public test suite of hostnames, each exercising one rule, with logs published on a schedule
    • a scoreboard per crawler
  • main risk: scoreboards attract gaming and lawyers
    • crawlers that hide their name are untestable by design
  • confidence the gap is real: medium
    • cheap to start, and it reuses the bait sites from ideas 3 and 4
  1. an open, multi-site log panel for crawler load
  • question: what do AI crawlers cost origin servers, measured the same way across many sites: bytes, share of uncached and expensive requests, and how that differs by site type
  • why not answered: Wikimedia, Read the Docs, GLAM-E and COAR each report their own pain in their own units
    • Kim 2025 has one university
    • Cloudflare reports shares, not costs, and only for customers
  • what we would build: a small log-sharing agreement with sites that already complain (libraries, repositories, open source forges), a common anonymization and bot-labeling pipeline, and a cost model per request type
  • main risk: privacy review and getting anyone to share logs
    • this is mostly organizing work, less a technical problem
  • confidence the gap is real: high that the data does not exist publicly, low that we are the right people to collect it

opinions from ChatGPT

  • no separate ChatGPT consultation for this crawler subreview

gaps in this review

  • search coverage is incomplete, so 2026 conference papers that are not on arXiv are likely missing
    • check the IMC 2026, USENIX Security 2026 and WWW 2026 programs for AI crawler measurement before starting any idea above
  • known but not read or quoted: Dinzinger and Granitzer, “A Longitudinal Study of Content Control Mechanisms” (WWW Companion 2024)
    • Fastly, TollBit, Imperva and Akamai bot reports
    • the Code4Lib 2025 article on crawler floods at UNC Libraries (unread)
    • Nepenthes and iocaine project pages
    • OpenAI, Anthropic and Perplexity crawler documentation
  • thin: older server-log studies of crawler behavior from 2000 to 2015
    • residential proxy networks used by scrapers
    • licensing deals between publishers and AI companies
    • the EU AI Act code of practice and its robots.txt clause
    • MCP, WebMCP and other agent-facing site interfaces
  • read as abstract only: the ai.txt, dark pattern, ad, ReaderLM, HtmlRAG, citation and freshness-cache papers
  • AIPREF and Web Bot Auth draft status changes monthly
    • I quoted the drafts’ text but did not track revisions

Last edited: