agents and the web: who visits, who blocks, who adapts, who gets cited (authored by agents unless marked đ§)
short version
- the supply side of the âweb for agentsâ is measured; the demand side is not
- papers count who publishes robots.txt rules, llms.txt, MCP servers, skills, and on-chain agent payments
- I found no peer-reviewed paper that checks which agents actually read or use these things on a site the researcher controls
- robots.txt is the only widely used control, and it is weak for the newest visitors
- assistants that fetch a page because a user asked often skip robots.txt or use an ordinary browser name
- 10 of 18 chatbots tested got page content through a search engineâs crawler (Googlebot, Bingbot, Bravebot), so blocking the chatbotâs own crawler does not keep content out
- agents can be told apart from people in the lab, but nobody has published how much real traffic they are
- 4 preprints fingerprint 6 to 7 agents each on test sites
- every traffic share number I found comes from a vendor that sells bot blocking
- the counts that look like adoption are mostly hollow
- more than half of MCP market listings are invalid or low value
- 21% of agent payment settlements on one chain are fictitious and 64% stay inside one linked cluster
- security scanners flag up to 47% of skills as malicious; 0.52% stay suspicious after a closer look
- AI search cites different pages than normal search, and the citation lists are noisy
- overlap between Google results and AI Overview sources is under 0.2 (Jaccard)
- sites that block Googleâs AI crawler are cited less by Gemini; blocking is also linked to a 7% traffic drop for large news publishers
- best research ideas, in my judgment
- 1: put up honey sites that offer every agent-facing door, then see which real agents use which door
- 2: visit top sites as a person and as an agent and compare what comes back
- 3: track robots.txt changes and AI search citations together over time to see whether blocking causes the citation loss
what the topic is
- AI companies send three kinds of visitors to websites
- training crawlers collect pages to train models (GPTBot, ClaudeBot)
- search crawlers build an index for AI answers (OAI-SearchBot, PerplexityBot)
- user-triggered visitors fetch a page, or drive a whole browser, because one person asked (ChatGPT-User, Comet, Operator)
- websites react in three ways
- block: robots.txt rules, bot blocking by a CDN, CAPTCHAs
- adapt: publish files and endpoints made for agents
- llms.txt: a markdown file that lists the pages a model should read
- MCP (Model Context Protocol): a standard way for a site or program to offer tools to an agent
- A2A agent card: a JSON file that says what an agent on this site can do
- x402: a site answers HTTP 402 âPayment Requiredâ and the agent pays with a stablecoin
- Web Bot Auth: the agent signs its HTTP requests so the site can verify who sent them
- charge: per-crawl pricing or licensing deals
- a web measurement researcher asks: how common is each thing, does it work, who gains, who loses
- reading depth
- full text skimmed (methods, results, limits): Kim 2025, Liu 2025, Lopez-Fonseca 2026, Seiden 2026, Grossman 2026, Stein 2026, Yang 2025, Zhou 2026, Zhao 2025, Ling 2026, Kang 2026, Holzbauer 2026
- abstract only: every other paper below
- numbers are the authorsâ reports for their setup and date; products change monthly
- already covered elsewhere in this tree, so only summarized here
- robots.txt blocking and corpus bias: crawling
- bot detection as a problem for research crawlers: same file
- manipulating AI search answers: SEO and search quality
- browser agent design and prompt injection: browser agents
what existing work shows
- do AI visitors obey robots.txt, and can a site even tell who they are?
- Scrapers selectively respect robots.txt directives, Kim, Bock, Luo, Liswood, Poroslay, Wenger, arXiv 2025, preprint on arXiv
- did: changed robots.txt on a universityâs sites and watched the logs
- fact: âapproximately 3.9 million external web requests made to a set of 36 websitesâ at âa large private US universityâ; â130 self-declared bots ⊠over 40 daysâ
- fact: âbots are less likely to comply with stricter robots.txt directives, and ⊠certain categories of bots, including AI search crawlers, rarely check robots.txt at allâ
- limit: one institution, 40 days; bots that hide their name are hard to attribute
- Do Generative AI Assistants Respect robots.txt?, Lopez-Fonseca, Rodriguez, Bechtold, Del Alamo, arXiv 2026, preprint
- did: asked 10 assistants about pages on the authorsâ server under 4 robots.txt settings, â200 trialsâ, with secret codes in the pages
- fact: âSome assistants exposed identifiable user-agents (Claude, Gemini, ChatGPT, Perplexity, and Mistral), whereas others used generic user-agents (Grok, Diffy Chat, Qwen, Copilot, and DeepSeek)â
- fact: some âaccessed restricted resources without requesting robots.txtâ
- fact: âassistants may access pages without surfacing the retrieved content, or fail to access even allowed resourcesâ
- limit: small trial count; one site; only robots.txt was tested
- Identifying AI Web Scrapers Using Canary Tokens, Seiden, Ren, Zhang, Kim, Liu, Wenger, arXiv 2026, preprint
- did: served a different made-up fact to each visiting scraper, then asked chatbots about the sites; the fact in the answer reveals which scraper fed it
- fact: âwe elicited User-Agent information for 18 of the 22 AI systemsâ
- fact: generic browser user-agents âfor six of the 18 AI systemsâ
- fact: âWe observed content associated with Googlebot, Bingbot, and Bravebot in responses from 10 of the 18 systemsâ
- inference: a site that wants to stay in Google but out of chatbots has no robots.txt rule that does it
- limit: needs the chatbot to answer about an unknown site; 4 systems never did
- Somesite I Used To Crawl, Liu, Luo, Shan, Voelker, Zhao, Savage, IMC 2025, peer reviewed
- did: large crawl plus âa targeted user study of 203 professional artistsâ; tested blocking by reverse proxies
- fact: âstrong demand for tools like robots.txt, but significantly constrained by critical hurdles in technical awareness, agency in deploying them, and limited efficacy against unresponsive crawlersâ
- limit: tests crawler names, not agents that drive a real browser
- From robots.txt to ai.txt, Hoffmann, Goergens, Khosla, Bajpai, ACM SIGCOMM CCR 2026, peer reviewed; I read only the abstract, on CiteDrive
- did: âa large-scale measurement of emerging Artificial Intelligence (AI)-oriented web permission and descriptor mechanisms across approximately 4M domainsâ
- fact: âadoption of AI-specific permission files is emerging, especially among technology-focused domainsâ
- fact: âsome llms.txt link targets point to content blocked by access-control filesâ
- limit: counts files; does not test whether any AI system reads them
- who blocks: Bouchaud and Ramaciotti, arXiv 2025, preprint; Steinacker-Olsztyn, Gosain, Dao, WWW 2026, peer reviewed
- fact (first): âA quarter of the top thousand websites restrict AI crawlers, decreasing to one-tenth across the broader top millionâ
- fact (second): â60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sitesâ
- what blocking costs: Strategic Response of News Publishers to Generative AI, Zhao, Berman, arXiv 2025, preprint
- did: compared publishers before and after they blocked, against ones that had not yet blocked; traffic from SimilarWeb, Semrush, and the Comscore panel
- fact: âAbout 75% of the top publishers blocked LLM crawlers at different times starting mid-2023â
- fact: âa 7% post-blocking decline in weekly visits measured by SimilarWeb or Semrush within the 6 weeks after blockingâ
- limit: traffic is a vendor estimate; publishers chose when to block, so timing may track something else
- proposals for a better control, none measured in the wild
- A Survey of Web Content Control for Generative AI, Dinzinger, HeĂ, Granitzer, arXiv 2024, preprint: site owners âare overwhelmed by the multitude of recent ad hoc standardsâ
- ai.txt DSL, Li et al., arXiv 2025, preprint: adds âelement-level regulationsâ and ânatural language instructionsâ
- terms.txt, Chowdhury, arXiv 2026, preprint: robots.txt âcannot express identity, purpose, terms, or priceâ; prototype âadds 0.20 to 0.65 ms per requestâ
- The Liabilities of Robots.txt, Chang, He, Computer Law and Security Review (accepted, per arXiv), legal analysis: robots.txt âcan give rise to a unilateral contractâ
- can a site tell an agent visit from a human visit?
- FP-Agent, Wang, Shafiq, Vekaria, arXiv 2026, preprint
- did: âseven AI browsing agents and human usersâ did 3 tasks on an instrumented test site
- fact: âbrowser fingerprints provide limited discriminative powerâ; âdifferences in typing, scrolling, and mouse behavior separate AI browsing agents from humans and one anotherâ
- fact: âFP-Agent detects all seven AI browsing agents, whereas Cloudflare detects only oneâ
- On the Internet, Nobody Knows Youâre an LLM Bot, Fayolle, Bouhenniche, PĂ©lissier, Laperdrix, Maurice, Rudametkin, arXiv 2026, preprint
- did: 6 agents visited test sites protected by robots.txt, CAPTCHAs, proof of work, and Cloudflareâs free tools
- fact: âsome Web Agents were able to bypass all evaluated anti-bot mechanismsâ; âstealth and anti-detection mechanisms often increase detectability rather than decrease itâ
- Whose Agent Are You?, Kang, Jeong, Sheffey, Datta, Houmansadr, arXiv 2026, preprint
- did: fingerprinted âAutoGen, Browser Use, Claude, Gemini, Operator, and Skyvernâ at the TLS, HTTP, and behavior layers; â97% accuracyâ
- fact (limit they state): âClaude and Gemini agents are indistinguishable because they present almost identical TLS and HTTP fingerprintsâ
- Broken Gates, Ousat, Turkmen, Rampersaud, Bailey, Kharraz, arXiv 2026, preprint
- fact: two agents with near identical behavior got different outcomes, âisolating execution-environment authenticity, rather than agent behavior, as the determining factorâ
- Developer Experience with AI Coding Agents: HTTP Behavioral Signatures in Documentation Portals, Borysenko, arXiv 2026, preprint
- did: logged requests from 9 coding agents and 6 assistants to one documentation endpoint
- fact: âAI agent access compresses multi-page navigation into a single or two requestsâ, so âsession depth, time-on-page, click path, and bounce rateâ stop meaning much
- limit: one endpoint, single author; closest thing to a demand-side study that I found
- inference from these five
- all train and test on agents the authors launched themselves
- none reports how many agent visits a real site gets, or how well the classifier works on traffic the authors did not create
- what an agent could be shown: A Whole New World: Creating a Parallel-Poisoned Web Only AI-Agents Can See, Zychlinski, arXiv 2025, preprint
- claim: a site can spot an agent and âdynamically serve a different, âcloakedâ version of its contentâ
- limit: an attack design; no count of sites that do it
- agents as attackers: LLM Agent Honeypot, Reworr, Volkov, arXiv 2024, preprint
- fact: an SSH honeypot with prompt-injection bait logged â8,130,731 hacking attempts and 8 potential AI agentsâ
- inference: bait that only a language model follows is a cheap way to label agent visits; nobody has used it on websites at scale
- traffic volume
- Logrip, Hoetzlein, arXiv 2025, preprint: on one small site â80 to 95 percent of traffic originates from AI crawlersâ; a single case
- LLMs as the Next Challenging Internet Traffic Source, Koneva et al., arXiv 2025, preprint: âThe average size of each prompt query and response is 7,593 bytesâ; about chat traffic, not web visits
- how do websites change for agents?
- Build the web for agents, not agents for the web, LĂč, Kamath, Mosbach, Reddy, arXiv 2025, position paper, preprint
- claim: sites should offer âan interface specifically designed for agentsâ; gives âsix guiding principlesâ
- VOIX, Schultze, Kietzmann, Schönfeld, Stock-Homburg, arXiv 2025, preprint
- did: new HTML tags that declare tools and state; âa three-day hackathon study with 16 developersâ
- webMCP, Perera, arXiv 2025, preprint
- claim: page metadata for agents âreduces processing requirements by 67.6% while maintaining 97.9% task success rates compared to 98.8%â
- limit: the authorâs own test scenarios
- big-picture papers, no measurements
- Agentic Web, Yang et al., arXiv 2025, preprint: frames it as âintelligence, interaction, and economicsâ
- Towards an Agent-First Web, Bandara et al., arXiv 2026, preprint: âagents acting for humans should inherit equivalent access rightsâ
- Beyond Message Passing, Yuan et al., arXiv 2026, preprint: compares â18 representative protocolsâ; they give âlimited protocol-level mechanisms for clarification, context alignment, and verificationâ
- agent payments, measured on the blockchain
- How Agentic Is Agentic Commerce?, Ling, Zhou, Wu, Wang, arXiv 2026, preprint
- fact: âOver a 280-day window Base carries 136,708,672 settlements worth $44,121,383.81â
- fact: â21.20% are fictitious and 63.78% internal settlement within a linked clusterâ
- fact: what is provably independent starts at âthe $187,861.35 that demonstrably reaches a nameable serviceâ
- claim: âSettlement count measures manufacturability, not adoptionâ
- Can Trustless Agents Be Trusted?, Xiong, Li, Wei, Wang, Knottenbelt, Wang, arXiv 2026, preprint
- fact: only â3%, 4%, and 15% across Ethereum, BSC, and Baseâ of registered agents have a valid file âwith at least one live service endpointâ
- fact: â73.5%, 59.2%, and 90.6%â of reviewers âexhibit coordinated Sybil behaviorâ
- The Web4 Agent Economy, Jin, Wu, Chen, Bao, Yang, Chen, arXiv 2026, preprint
- claim: agents run âa highly active machine-to-machine payment economy, processing millions of daily transactionsâ
- inference: Ling et al. looked at who pays whom and reached the opposite reading of the same kind of data; I trust the payment graph more than the raw count
- How Agentic Is Agentic Commerce?, Ling, Zhou, Wu, Wang, arXiv 2026, preprint
- inference: for llms.txt, MCP on websites, agent cards, WebMCP, and Web Bot Auth, I found design papers and one file census, but no study of use
- how do people really use agents?
- The Adoption and Usage of AI Agents: Early Evidence from Perplexity, Yang, Yonack, Zyskowski, Yarats, Ho, Ma, arXiv 2025, preprint by the vendor
- did: classified âhundreds of millions of anonymized user interactionsâ with the Comet browser agent
- fact: âProductivity & Workflow and Learning & Research, account for 57% of all agentic queriesâ; âCourses and Shopping for Goods, make up 22%â
- fact: âPersonal use constitutes 55% of queries, while professional and educational contexts comprise 30% and 16%â
- limit: one product, early adopters, data the public cannot check; nothing on which sites the agent visits or how often it fails
- How People Use ChatGPT, Chatterji, Cunningham, Deming, Hitzig, Ong, Shan, Wadman, NBER working paper 2025, not peer reviewed, vendor data
- fact: adopted âby around 10% of the worldâs adult populationâ by July 2025
- fact: non-work messages âhave grown from 53% to more than 70% of all usageâ
- fact: ââPractical Guidance,â âSeeking Information,â and âWritingâ ⊠collectively account for nearly 80% of all conversationsâ
- Measuring AI agent autonomy in practice, McCain et al., Anthropic research post, Feb 2026, not peer reviewed; read through a page summary tool
- fact: âSoftware engineering accounted for nearly 50% of tool calls on our public APIâ
- fact: âThe 99.9th percentile turn duration nearly doubled, from under 25 minutes to over 45 minutesâ
- fact: âonly 0.8% of actions appear to be irreversibleâ
- limit they state: âWe can only analyze individual tool calls in isolation, rather than full agent sessionsâ
- Agentic Much? Adoption of Coding Agents on GitHub, Robbes, Matricon, Degueule, Hora, Zacchiroli, arXiv 2026, preprint
- did: looked for agent traces such as co-authored commits in â128,018 projectsâ
- fact: âan estimated adoption rate of 22.20%â28.66%â
- inference: the only usage study here that outsiders can redo, because the traces are public
- How are AI agents used? Evidence from 177,000 MCP tools, Stein, arXiv 2026, preprint
- fact: âSoftware development accounts for 67% of all agent tools, and 90% of MCP server downloadsâ
- fact: âthe share of âactionâ tools rose from 27% to 65% of total usageâ
- limit they state: they âdo not observe whether downloaded tools are actually calledâ
- side effects on the public web
- Are Large Language Models a Threat to Digital Public Goods?, del Rio-Chanona, Laurentsyeva, Wachs, arXiv 2023; search results say it later appeared in PNAS Nexus, which I did not check
- fact: âA difference-in-differences model estimates a 16% decrease in weekly posts on Stack Overflowâ
- Big Help or Big Brother?, Vekaria et al., USENIX Security 2025, peer reviewed; read through a page summary tool
- fact: assistant browser extensions share âfull webpage content, including the HTML DOM and user form inputsâ with their servers
- Are Large Language Models a Threat to Digital Public Goods?, del Rio-Chanona, Laurentsyeva, Wachs, arXiv 2023; search results say it later appeared in PNAS Nexus, which I did not check
- catalogs of what exists: The AI Agent Index, Casper et al., arXiv 2025, preprint; The 2025 AI Agent Index, Staufer et al., FAccT 2026 per its arXiv comment
- fact (2025 index): covers â30 state-of-the-art AI agentsâ; âmost developers share little information about safety, evaluations, and societal impactsâ
- inference: all three large usage studies come from the vendorâs own logs; the web side of an agent session is invisible in them
- what do AI search engines cite?
- How Generative AI Disrupts Search, Grossman, Liu, Chen, Smith, Borcea, Chen, SIGIR 2026, peer reviewed
- did: â11,500 user queriesâ sent to Google Search, AI Overviews, and Gemini 2.5 Flash
- fact: âfor 51.5% of representative, real-user queries, AIOs are generatedâ
- fact: sources differ, â<0.2 average Jaccard similarityâ
- fact: 21 popular publishers and several social and review sites âwere never cited by Geminiâ; âall of these websites block the Google-Extended botâ
- limit they state: âsolely focuses on Googleâ; âdescriptive in nature and limited to a single point in timeâ
- Quantifying Uncertainty in AI Visibility, Sielinski, arXiv 2026, preprint
- did: repeated the same queries on Perplexity, SearchGPT, and Gemini, daily for 9 days and every 10 minutes
- fact: âmany apparent differences between domains fall within the noise floor of the measurement processâ
- inference: any citation study that runs each query once is suspect; this is the cheapest lesson in the whole topic
- From Citation Selection to Citation Absorption, Zhang, He, Yao, arXiv 2026, preprint
- did: â602 controlled prompts across ChatGPT, Google AI Overview/Gemini, and Perplexity; 21,143 valid search-layer citationsâ
- fact: âPerplexity and Google cite more sources on average, while ChatGPT cites fewer sources but shows substantially higher average citation influenceâ
- Navigating the Shift, Chen, Wang, Chen, Koudas, EDBT/ICDT 2026 Workshops
- fact: AI answers and web search âdiverge significantly in their consulted source domains, the typology of these domains (e.g., earned media vs. owned, social), query intent, and the freshnessâ
- When Content is Goliath and Algorithm is David, Ma, Qin, Xu, Tan, arXiv 2025, preprint
- fact: Googleâs AI search prefers âcontent characterized by significantly higher predictability for underlying LLMsâ
- inference: pages rewritten by a language model may get cited more; this links to the sibling study of AI-written pages
- inference: nobody has joined the three data sets of site policy, AI citation, and traffic for the same sites over time
- what is in the agent tool stores?
- MCP servers listed in markets and code hosts
- A Measurement Study of Model Context Protocol Ecosystem, Guo, Hao, Zhang, Xu, Lv, Chen, Cheng, arXiv 2025, preprint
- fact: â17,630 raw entries, of which 8,401 valid projects (8,060 servers and 341 clients)â; âmore than half of listed projects are invalid or low-valueâ
- MCP at First Glance, Hasan, Li, Fallahzadeh, Rajbahadur, Adams, Hassan, arXiv 2025, preprint
- fact: of â1,899 open-source MCP serversâ, â7.2% of servers contain general vulnerabilities, and 5.5% exhibit MCP-specific tool poisoningâ
- Rethinking MCP Security, Chen et al., arXiv 2026, preprint; builds on the MCPZoo dataset, Wu et al., arXiv 2025
- fact: â64,611 unique MCP servers ⊠more than 37,288 supporting dynamic analysisâ
- fact: scanners âreport that 96.89% of servers are riskyâ; âless than 50% of sampled alerts are true positivesâ
- MCPEvol-Bench, Liu et al., arXiv 2026, preprint
- fact: when server tools change, âGPT-5.4 and Claude-Sonnet-4-6 exhibit performance declines of 13.7% and 14.4%â
- A Measurement Study of Model Context Protocol Ecosystem, Guo, Hao, Zhang, Xu, Lv, Chen, Cheng, arXiv 2025, preprint
- MCP servers reachable on the internet
- A First Measurement Study on Authentication Security in Real-World Remote MCP Servers, Zhou, Zhang, Zhang, Zhang, Zhang, Yang, arXiv 2026, preprint
- did: found candidates with FOFA and Shodan, then confirmed each with an MCP handshake
- fact: â7,973 live remote MCP servers, finding that 40.55% expose tools without authenticationâ
- fact: of â119 testable real-world OAuth-enabled MCP servers ⊠each server exhibits at least one flawâ; â9 CVE IDsâ
- Exposed by Design, Padilla, arXiv 2026, preprint
- fact: âwe confirm 640 production MCP servers and dynamically audit 414, uncovering 68 reportable vulnerabilitiesâ
- fact: â41.6% of confirmed servers disappear within three days between consecutive measurement runsâ
- inference: the two counts differ by 12 times (7,973 against 640) because âliveâ and âproductionâ are defined differently; nobody has reconciled them
- A First Measurement Study on Authentication Security in Real-World Remote MCP Servers, Zhou, Zhang, Zhang, Zhang, Zhang, Yang, arXiv 2026, preprint
- skills
- Agent Skills in the Wild, Liu, Wang, Feng, Zhang, Xu, Deng, Li, Zhang, arXiv 2026, preprint
- fact: â26.1% of skills contain at least one vulnerabilityâ; â5.2% of skills exhibit high-severity patterns strongly suggesting malicious intentâ
- Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem, Holzbauer, Schmidt, Gegenhuber, Schrittwieser, Ullrich, AgentSkills 2026 workshop, peer reviewed
- fact: â238,180 unique skillsâ; scanner reports âclassify up to 46.8% of skills as maliciousâ; âonly 0.52% remain suspicious after repository-aware analysisâ
- fact: ârepository hijacking risks for seven abandoned repositories referenced by skill indexes, affecting 121 skillsâ
- Exploring the Emerging Threats of the Agent Skill Ecosystem, Beurer-Kellner et al., arXiv 2026, technical report
- fact: â3,984 AI agent skills ⊠76 confirmed malicious payloadsâ
- inference: the first and second papers disagree by a factor of 10 on how much is malicious; the scanner is the unknown, not the store
- Agent Skills in the Wild, Liu, Wang, Feng, Zhang, Xu, Deng, Li, Zhang, arXiv 2026, preprint
- older stores, same pattern
- A First Look at GPT Apps, Zhang, Zhang, Yuan, Zhang, Xu, Qian, arXiv 2024, preprint: ânearly 90% system prompts can be easily accessedâ; âcreator interest plateaus within three monthsâ
- GPT Store Mining and Analysis, Su, Zhao, Hou, Wang, Wang, arXiv 2024, preprint: topic and popularity census
- An Empirical Study on the Security Vulnerabilities of GPTs, Wu, Wu, Zheng, arXiv 2025, preprint: attack suite for âinformation leakage and tool misuseâ
- inference: every store study measures what is published; only Stein uses download counts, and none sees calls
what is missing
- nobody has tested which agents use the agent-facing doors
- evidence: Hoffmann 2026 counts llms.txt files only; Lopez-Fonseca 2026 tests robots.txt only; Borysenko 2026 covers one documentation endpoint
- evidence: search snippets of industry posts (Ahrefs, Originality.ai; I did not open them) say most llms.txt files get no AI requests; that is a vendor claim with no public method
- 2 targeted searches for a controlled study of llms.txt, MCP, or agent card use returned none
- nobody has measured what real sites serve to agents compared with people
- evidence: Liu 2025 and Steinacker-Olsztyn 2026 test whether a crawler name is blocked, a yes or no answer
- evidence: Zychlinski 2025 designs agent-only content as an attack; a search for a wild measurement found only a vendor demo (SPLX; not opened)
- no independent number for how much traffic is agents
- evidence: the fingerprinting papers use only self-launched agents; Kim 2025 has real logs but only bots that name themselves, at one university in early 2025
- evidence: the shares in search results come from Cloudflare, HUMAN, DataDome, and TollBit, who sell blocking
- no causal link from blocking to citation to traffic
- evidence: Grossman 2026 is âlimited to a single point in timeâ; Zhao 2025 has traffic but no citations; Sielinski 2026 shows single runs are noise
- no study of the official MCP servers that large websites run
- evidence: Zhou 2026 and Padilla 2026 scan IP space and code hosts; neither ties a server to the website it belongs to or compares it with the siteâs robots.txt and terms
- no outside view of agent sessions on the web
- evidence: Yang 2025, Chatterji 2025, and McCain 2026 all use private vendor logs; McCain says tool calls are seen âin isolationâ
- these are âI searched and did not findâ, not proof; IMC 2026 (12 to 16 Oct 2026) may publish some of this, and I could not read its accepted list
research we can do
- who walks through which door: a honey site study of agent-facing interfaces
- question: when a site offers llms.txt, a markdown version of each page, an MCP endpoint, an agent card, WebMCP tools, and an x402 paywall, which real assistants and agents use which, and does it change what they answer or do?
- why open: supply is counted (Hoffmann 2026), use is not; see the first gap
- first experiment
- register about 20 fresh domains with made-up but plausible content, such as a product catalog and docs
- each domain offers a different subset of doors; each door holds a different secret fact, as in Seiden 2026
- ask about 20 assistants, browser agents, and coding agents questions that need the site; log requests and see which secret shows up
- repeat weekly for 2 months to catch product changes
- convincing result: a table of product by door with use rates and confidence intervals, plus whether a door changes answer correctness; âno product reads llms.txt in search mode but 5 coding agents doâ would be a clean finding
- cost: about $300 for domains and hosting, $500 to $1,500 for subscriptions and API calls, 6 to 8 weeks for one person
- closest work that could scoop it: the Wenger group (Seiden, Kim, Liu) already runs such sites; Lopez-Fonsecaâs setup extends easily; Hoffmannâs group may add a use study
- risk: fresh domains may never be fetched by search-backed assistants; give the URL in the prompt as a second condition
- what the web shows an agent: person and agent visits compared at scale
- question: on the top 10K to 100K sites, how often does a visit that declares itself an agent get a block, a challenge, a paywall, different text, or text aimed at the model, compared with a personâs browser from the same network?
- why open: prior crawls record block or no block for crawler names; content differences and agent-only instructions are unmeasured; see the second gap
- first experiment
- 4 visitors per site: plain Chrome, Chrome with an agent user-agent (ChatGPT-User, Claude-User, Perplexity-User), headless automation, and one real agent product on a 1K subsample
- visit twice as each, so normal page churn is not counted as a difference
- compare status, main text, links, prices, and hidden text; flag text addressed to a model
- convincing result: rates by site category and CDN with a churn baseline, hand-checked samples, and a set of real agent-only pages
- cost: one crawl machine, 2 to 4 weeks for the top 10K; real agent runs cost about $0.05 to $0.50 per page, so keep that sample small
- closest work that could scoop it: Gundelach 2026 (block rates for automation), Liu 2025 (crawler blocking), the UC Davis group behind FP-Agent
- risk: a spoofed user-agent from the wrong IP range may be treated as fake; report that as its own finding, and it argues for Web Bot Auth
- does blocking AI crawlers cost citations and visits?
- question: when a site starts or stops blocking an AI crawler, do its citations in AI search change, and how fast?
- why open: Grossman 2026 found the link at one moment; Zhao 2025 found a traffic drop without looking at citations
- first experiment
- take monthly robots.txt history from HTTP Archive or Common Crawl for about 5K publisher and reference sites
- each week send a fixed set of 2K queries to 3 or 4 AI search products, 5 runs each, following Sielinski 2026
- compare citation share around each robots.txt change with sites that did not change
- convincing result: a drop or rise that starts after the change, is absent before it, and holds across products
- cost: $1K to $3K per month in API and search result fees, and at least 4 months to see enough changes
- closest work that could scoop it: Grossmanâs group has the query set; Zhao and Berman have the traffic panel
- risk: few sites change policy in a short window; a long window costs money
smaller ideas
- the official agent door of big sites: find first-party remote MCP servers of the top 10K sites; compare their tools with the siteâs robots.txt, terms, and login rules; rescan weekly for churn (Padilla saw 41.6% vanish in 3 days)
- agent traffic share without a vendor: get web logs from a university or an open-source project; label agents by fingerprints from the 4 lab papers plus model-only bait links; report the share that names itself against the share that hides
- reconcile the scanners: run the skill and MCP scanners on one shared sample with hand labels; the 46.8% against 0.52% gap says the tools are the problem
ChatGPTâs opinion
- pending; the coordinator said ChatGPT cannot be used until the human signs in, so I did not run it
what I searched
- when: 7 Oct 2026; web search worked this time
- about 30 search queries, on: AI crawler measurement and robots.txt; llms.txt adoption; MCP ecosystem and remote MCP scans; skill and GPT stores; AI search citations; AI Overviews and traffic; agent traffic and fingerprinting; cloaking for agents; Web Bot Auth; x402 and agent payments; A2A agent cards; agentic web surveys; agent usage field studies; Stack Overflow and Wikipedia effects; agentic browser privacy; residential proxies; IMC 2026 papers
- sources: arXiv abstract pages and full-text HTML, NBER, USENIX, Anthropic, CiteDrive for one ACM abstract
- opened 59 papers and reports: 12 with full text skimmed, the rest abstract only
- found but not opened, so not cited as evidence
- Agarwal and Sen, randomized browser-extension study of AI Overviews and clicks (SSRN returned 403); search snippets say clicks fell about 40%
- Longpre et al., Consent in Crisis; it is in the sibling crawling file
- industry reports from Cloudflare Radar, HUMAN, TollBit, Ahrefs, Originality.ai
- BADPASS (bots using residential proxies), NANDA index, IETF Web Bot Auth and AIPREF drafts
- not covered
- the IMC 2026, CCS 2026, and NDSS 2026 accepted lists; check them before starting idea 1 or 2
- ChatGPT plugin era papers before 2024
- legal cases and licensing deals
- A2A agent cards and Web Bot Auth: I found no measurement paper, only specifications
Last edited: