how researchers crawl the web for measurement studies (authored by agents unless marked đ§)
- written 2026-10-06; quotes are verbatim from the paperâs abstract or body unless I say otherwise
- âfactâ = stated by the source; âinferenceâ / âI thinkâ = mine
- search was limited: web search and most paper APIs were rate limited this session, so discovery leaned on arXiv metadata (DataCite), Crossref, Springer pages, and PDFs I could fetch directly
- consequence: 2023 to 2026 papers that are only in ACM DL or IEEE Xplore may be missing
the picture in plain words
- the question behind every crawl study: âwhat does a normal visitor get when they open this site?â, asked for thousands to millions of sites
- the method is always the same three choices
- which sites and pages: almost always a popularity list (Alexa until 2022, now Tranco or CrUX), and almost always only the home page
- which client: a real browser driven by a script (OpenWPM = Firefox + Selenium, or Chrome + Puppeteer / Playwright), or a plain HTTP fetcher (Heritrix, Common Crawlâs CCBot) when scale matters more than JavaScript
- from where and in what state: usually one cloud or university address, a fresh profile, no login, no clicks, no consent given
- the uncomfortable finding of the last eight years: each of those choices changes the answer, often by tens of percent
- the list: Alexa changed half its entries daily; lists overstate what the wider web does
- the page: internal pages differ from home pages
- the client: sites detect automation and serve something else; headless differs from headful
- the place: cloud addresses see less third-party content than home addresses; EU addresses see consent banners
- the state: before consent, before login, and without interaction you see a thinner site
- chance: two identical crawlers running side by side disagree
- the 2024 to 2026 twist: AI scrapers made site owners block bots much harder, and research crawlers are caught in the same net
- my take: the field has many âX mattersâ papers for tracker and cookie counts, and very few that (a) say how big the error is on your own metric, (b) cover page content and JavaScript behavior rather than privacy metrics, or (c) track the problem over time
crawler tools
- Web Crawling, Olston, Najork, Foundations and Trends in IR, 2010 (survey)
- fact: splits the classic literature into âBuilding an efficient, robust and scalable crawlerâ, âSelecting a traversal order of the web graphâ, âScheduling revisitation of previously crawled contentâ, âAvoiding problematic and undesirable contentâ, and âCrawling so-called âdeep webâ content, which must be accessed via HTML forms rather than hyperlinksâ
- inference: this is the search engine view (collect as many good pages as possible); measurement crawling asks a different question (see a fixed sample the way a user would), so its problems (bot detection, state, repeatability) barely appear here
- IRLbot: Scaling to 6 Billion Pages and Beyond, Lee, Leonard, Wang, Loguinov, WWW 2008
- fact: âIRLbot running on a single server successfully crawled 6.3 billion valid HTML pages (7.6 billion connection requests) and sustained an average download rate of 319 mb/s (1,789 pages/s)â
- fact: the hard parts were âverifying URL uniqueness, BFS crawl order, and fixed per-host ratelimitingâ, plus âhighly-branching spam, legitimate multi-million-page blog sites, and infinite loops created by server-side scriptsâ
- open: no JavaScript; that throughput is impossible with a browser
- Heritrix, Internet Archive (README)
- fact: âthe Internet Archiveâs open-source, extensible, web-scale, archival-quality web crawler projectâ; âdesigned to respect the
robots.txtexclusion directivesâ - no browser; the humanâs note on Danish / Luxembourgish / Finnish news archiving already covers how libraries pair it with Browsertrix
- fact: âthe Internet Archiveâs open-source, extensible, web-scale, archival-quality web crawler projectâ; âdesigned to respect the
- Browsertrix Crawler, Webrecorder (README)
- fact: âa standalone browser-based high-fidelity crawling systemâ; âuses Puppeteer to control one or more Brave Browser browser windows in parallel. Data is captured through the Chrome Devtools Protocol (CDP)â
- Online Tracking: A 1-million-site Measurement and Analysis, Englehardt, Narayanan, CCS 2016 (OpenWPM)
- fact: âuses an automated version of a full-fledged consumer browser. It supports parallelism for speed and scale, automatic recovery from failures of the underlying browser, and comprehensive browser instrumentationâ
- fact from the OpenWPM README: âbuilt on top of Firefox, with automation provided by Seleniumâ
- open: Firefox is a small share of real users, and the instrumentation is itself detectable (Krumnow below)
- Tracker Radar Collector, DuckDuckGo (README)
- fact: âModular, multithreaded, puppeteer-based crawler used to generate third party request data for the Tracker Radarâ
- used as a research crawler by Leaky Forms (below)
- VisibleV8: In-browser Monitoring of JavaScript in the Wild, Jueckstock, Kapravelos, IMC 2019
- fact: âa dynamic analysis framework hosted inside V8 ⊠that logs native function or property accesses during any JS execution. At less than 600 lines (only 67 of which modify V8âs existing behavior)â
- fact: found â46 JavaScript namespace artifacts used by JS code in the wild to detect automated browsing platformsâ and â29% of the Alexa top 50k sites load content which actively probes these artifactsâ
- relevant to the JSphere line: logging inside the engine cannot be seen by page scripts, unlike injected wrappers
- Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web, Ahmad, Dar, Zaffar, Vallina-Rodriguez, Nithyanand, WWW 2020
- fact (abstract, via Semantic Scholar): âcrawlers generally find themselves trading off between computational overhead, developer effort, data accuracy, and completeness. Therefore, the choice of crawler has a critical impact on the data generated and knowledge inferred from itâ
- fact: they âconduct a survey of all research published since 2015 in the premier security and Internet measurement venues to identify and verify the repeatability of crawling methodologiesâ
- I could not fetch the PDF, so I have no numbers; my memory (unverified) is that plain HTTP tools miss a large share of what browsers load
- SoK: State of the Krawlers, Stafeev, Pellegrino, USENIX Security 2024
- what: read 12 years of papers, rebuilt the crawling algorithms in one framework (Arachnarium), compared coverage
- fact: â7,840 papers, identifying 403 conducting a measurementâ
- fact: âOf the 403 surveyed papers, only 32.3% of them (i.e., 130 papers) navigate websites whereas the remaining 273 papers (i.e., 67.7%) visit a single page onlyâ
- fact: â35.4% of the papers navigating websites do not specify (i) the navigation strategy âŠ, (ii) the page similarity âŠ, or (iii) bothâ
- fact: ârandomized algorithms, in particular randomized BFS, offer better performance than standard algorithms, providing a general increase of average coverage, from +8% in LoCs up to +21% for JavaScript source codeâ
- open: coverage means code and link coverage for security scanning; nothing on whether the pages reached are the pages users visit
- Web Execution Bundles: Reproducible, Accurate, and Archivable Web Measurements, Hantke, Snyder, Haddadi, Stock, USENIX Security 2025 (WebREC)
- what: a browser-level recorder and a â.webâ archive format that keeps what happened in the page (which script did what), not just the HTTP traffic
- fact: âmost measurement studies use custom tools and varied archival formats, each of unknown correctness and significant limitationsâ
- fact: â70% of papers discussed in a 2024 web crawling SoK paper could be conducted using WebREC as is, and a larger number (48%) could be leveraged against .web archives without requiring any new crawlingâ
- the âlarger number (48%)â wording is the paperâs own; I read it as 48% of papers need no new crawl at all
- fact: âOpenWPM and similar systems fundamentally cannot address the unpreventable errors caused by the limitations of the underlying capabilities OpenWPM relies onâ
- open: nobody runs a shared, regularly refreshed .web crawl yet, as far as I found
- Sprinter: Speeding Up High-Fidelity Crawling of the Modern Web, Goel et al., NSDI 2024 (I am fairly sure of venue; the copy in the collection has no venue line)
- fact: âto discover and fetch all page resources dependent on JavaScript and modern web APIs, crawlers today have to employ compute-intensive web browsersâ
- fact: âcrawls a small, carefully chosen, subset of pages on each site using a browser, and then efficiently identifies and exploits opportunities to reuse the browserâs computations on other pagesâ; âcrawl a corpus of 50,000 pages 5x faster than browser-based crawling, while still closely matching a browser in the set of resources fetchedâ
- inference: the reuse works because pages of one site share scripts, which is the same structure a âweb atomâ would exploit for change detection
- crawlers built for collecting training text rather than measuring (short, for contrast)
- Craw4LLM, Yu, Liu, Xiong, arXiv 2025: âWith just 21% URLs crawled, LLMs pretrained on Craw4LLM data reach the same downstream performances of previous crawlsâ
- Neural Prioritisation for Web Crawling, Pezzuti, MacAvaney, Tonellotto, ICTIR 2025: âneural crawling policies significantly improve harvest rate, maxNDCG, and search effectiveness during the early stages of crawlingâ
- Efficient Crawling for Scalable Web Data Acquisition, Gauquier, Manolescu, Senellart, EDBT 2026: a bandit learns âwhich hyperlinks lead to pages that link to many targets, based on the paths leading to the links in their enclosing webpagesâ
- inference: all three pick pages by expected value, which makes the sample biased on purpose; the opposite of what a measurement needs
which sites to visit: top lists
- A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists, Scheitle et al., IMC 2018
- fact: âtop lists generally overestimate results compared to the general population by a significant margin, often even an order of magnitudeâ
- fact: âsome top lists have surprising change characteristics, causing high day-to-day fluctuation and leading to result instabilityâ
- fact: â59 studies exclusively use Alexa as a source for domain namesâ
- Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation, Le Pochat et al., NDSS 2019
- fact: âhalf of the Alexa list changes every day and the Umbrella list only has 49% real sitesâ
- fact: âthe ranks of domains in each of the lists are easily altered, in the case of Alexa through as little as a single HTTP requestâ
- fact: â133 top-tier studies over the past four years based their experiments and conclusions on the data from these rankingsâ
- what Tranco is: an average of several lists over 30 days, with a permanent ID per list so a study can cite the exact list
- follow-up: Evaluating the Long-term Effects of Parameters on the Characteristics of the Tranco Top Sites Ranking, CSET 2019: âWe compute how well Tranco captures websites that are responsive, regularly visited and benignâ
- open: Alexa was retired in 2022, so Trancoâs inputs changed; I did not find a paper that re-evaluates todayâs Tranco
- Clustering and the Weekend Effect, Rweyemamu et al., PAM 2019
- fact: âthe weekend effect in Alexa and Umbrella causes these rankings to change their geographical diversity between the workweek and the weekendâ
- fact: âup to 91% of ranked domains appear in alphabetically sorted clusters containing up to 87k domains of presumably equivalent popularityâ
- plain meaning: deep in the list the order is alphabetical, so âtop 500Kâ versus âtop 600Kâ is a cut by first letter
- Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists, Ruth, Kumar, Wang, Valenta, Durumeric, IMC 2022
- what: compared lists against what Cloudflareâs servers see
- fact: âmost lists capture web popularity poorly, with the exception of the Chrome User Experience Report (CrUX) dataset, which is the most accurate top list compared to Cloudflare across all metricsâ
- fact: âresearchers typically use lists as an unordered set of websitesâ
- open: Cloudflare âauthoritatively serves traffic for only about a quarter of top sitesâ, so the reference itself is a biased sample; CrUX gives only rank buckets (top 1K, 10K, 100K, 1M) and only Chrome users who opted in
- A World Wide View of Browsing the World Wide Web, Ruth et al., IMC 2022
- fact: âsix sites account for 25% of page loads on both desktop and mobile, and one site garners 17% of all desktop page loads globallyâ
- fact: âThe top million sites capture over 95% of all page loads and time spent online, but they do so extremely unequallyâ
- fact: âof sites appearing in the top 1K for at least one country, over half do not rank in the top 10K for any other countryâ
- fact: âstudies calculated directly from a simple set of the top million sites place disproportionate focus on the long tail of the webâ
- inference: âx% of the top 1M sites do Yâ and âx% of page loads meet Yâ are different quantities, and most papers report the first while implying the second
- Crawling to the Top: An Empirical Evaluation of Top List Use, Xie, Li, PAM 2024
- what: put test domains into top lists and watched who came (abstract from the Springer page; no numbers there)
- fact: âevaluate how domain traffic changes once placed in top lists, the characteristics of those visiting the domain, and the behavioral patterns of these visitorsâ
- inference: being on a list attracts scanners, so list membership changes the thing being measured
- You Get What You Sample: Evaluating Sampling Strategies for Web Security Measurements, Zhang, Yang, Pellegrino, arXiv 2026
- what: compared âtop Nâ with random and stratified samples on 500k Tranco and 24.8M Common Crawl hosts
- fact: of 107 measurement papers (2020 to 2024), datasets were âTranco (61) and Alexa (41) ⊠followed by CrUX (9), Common Crawl (7), Cisco Umbrella (3), and SecRank (2)â
- fact: âTop N does not reflect the overall distribution of the web. Instead, probability-based strategies yield stable, unbiased estimates for prevalenceâ
- fact: ârelative to the broader Common Crawl baseline, random sampling from Tranco may underestimate the prevalence of the measured security issues by approximately 30%, even with an appropriately large sample sizeâ
- open: sampling of sites only, for security header style metrics; does not touch which pages inside a site
- LLM-Assisted Web Measurements, Bozzolan, Calzavara, Cazzaro, arXiv 2025
- fact: âexisting top lists of popular websites are unlabeled and lack semantic information about the nature of the included websitesâ; they âpropose a practical two-step methodology for scalable targeted web measurements starting from the Tranco listâ using LLMs to label sites
- open: uses the LLM to choose sites, not to drive the crawl
which pages to visit: landing versus internal
- On Landing and Internal Web Pages, Aqeel, Chandrasekaran, Feldmann, Maggs, IMC 2020 (Hispar; already in the humanâs notes)
- fact: âthe insights and claims of nearly two-thirds of the relevant studies would need to be revised for them to apply to internal pagesâ
- fact: Hispar âuses search engine results for discovering internal pagesâ; âFor each web site, we used the Google Search Engine API for the term âsite:.â We fix the userâs location for the search queries to the United States, and limit the search results to pages in the English languageâ
- open: ârepresentative internal pageâ is defined by what a US English Google search returns; no error bars on per-site numbers
- Beyond the Front Page: Measuring Third Party Dynamics in the Field, Urban, Degeling, Holz, Pohlmann, WWW 2020
- fact: landing-page-only measurement âis only able to measure a lower bound as subsites show a significant increase of privacy-invasive techniquesâ
- fact: â50 % of the branches in the third party trees change between repeated visitsâ
- Reproducibility and Replicability of Web Measurement Studies, Demir et al., WWW 2022 (also below)
- fact: âOur total website corpus consists of 10k distinct sites and we found 182,586 subpages on those sitesâ
- Exploring the Cookieverse, Rasaii, Singh, Gosain, Gasser, PAM 2023 also âcompare landing to inner pagesâ (details under consent)
- crawling deeper than links: state and interaction
- Crawling JavaScript-based web applications through dynamic analysis of user interface state changes, Mesbah, van Deursen, Lenselink, TWEB 2012 (Crawljax)
- fact: Ajax techniques âshatter the concept of webpages with unique URLs, on which traditional Web crawlers are basedâ; the crawler âfires events on those candidate elements, and incrementally infers a state machineâ
- Black Widow: Blackbox Data-driven Web Scanning, Eriksson, Pellegrino, Sabelfeld, S&P 2021
- fact: âcode coverage improvements ranging from 63% to 280% compared to other crawlers across all applicationsâ
- YuraScanner: Leveraging LLMs for Task-driven Web App Scanning, Stafeev et al., NDSS 2025
- fact: scanners âstruggle with discovering deeper states in modern web applications due to their limited understanding of workflowsâ; an LLM agent âidentified 12 unique zero-day XSS vulnerabilities, compared to three by Black Widowâ
- open: âevaluated ⊠on 20 diverse web applicationsâ that the authors host; not the live web
- SpiderSapien: Client-Centric Web Crawler and Security Scanner, Olsson, Eriksson, Doupé, Sabelfeld, arXiv 2026
- fact: âincreased average code coverage across applications by at least 46% over any other scannerâ; uses âan LLM to solve formsâ
- Dark Patterns at Scale, Mathur et al., CSCW 2019
- fact: âAnalyzing ~53K product pages from ~11K shopping websites, we discover 1,818 dark pattern instancesâ
- why it is here: an example of a task-specific crawler that has to find product pages and walk a checkout flow
- inference: every one of these judges a crawler by code coverage or bugs found; none asks how the extra pages change a prevalence number
- Crawling JavaScript-based web applications through dynamic analysis of user interface state changes, Mesbah, van Deursen, Lenselink, TWEB 2012 (Crawljax)
how results depend on the setup
the browser and the automation framework
- Towards Realistic and Reproducible Web Crawl Measurements, Jueckstock et al., WWW 2021
- what: ânaive crawling tool defaults vs. careful attempts to match ârealâ users across the Tranco top 25k web domainsâ
- fact: âbrowser configuration alone causing shifts in 19% of known ad and tracking domains encountered and altering the loading frequency of up to 10% of distinct JavaScript code units executedâ
- fact: ânetwork vantage point having similar, though less dramatic, effects on the same web metricsâ
- abstract via Semantic Scholar; I could not fetch the PDF
- Reproducibility and Replicability of Web Measurement Studies, Demir et al., WWW 2022
- what: checklist from 117 papers, then â4.5 million pages with 24 different measurement setupsâ
- fact: âthe identified trackers on pages can vary by 25% based on the used browser configurationâ
- fact, headless versus headful similarity of tracker sets per page: âperfect similarity for 35% of the pages and no similarity for 34%â
- fact on reporting: âapproximately 17.1% of the papers, where the pipeline is described, fail to offer details on the crawling technologyâ; âonly 24% of the analyzed papers make their results openly availableâ
- OmniCrawl, Cassel et al., PETS 2022
- what: â42 different non-emulated browsers simultaneouslyâ, including real phones
- fact: âcommon methodological choices made by web measurement studies, such as the use of emulated mobile browsers and Selenium, can lead to website behavior that deviates from what actual users experienceâ
- fact: âtracking observed on Selenium-driven desktop browsers differs from tracking on non-Selenium-driven browsersâ
- fact: âthe third-party advertising and tracking ecosystem of mobile browsers is more similar to that of desktop browsers than previous findings suggestedâ
- contrast: A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments, Yang, Yue, PETS 2020, with emulated mobile on â23,310 websitesâ: âmobile web tracking has its unique characteristics especially due to mobile-specific trackersâ
- inference: the two papers disagree on how different mobile is, and OmniCrawl blames emulation
- The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing, Zeber et al., WWW 2020
- what: compared OpenWPM crawls with âan opt-in sample of over 50,000 users of the Firefox Web browserâ
- fact: âWe quantify baseline variation of simultaneous crawls, then isolate the effects of time, cloud IP address vs. residential, and operating systemâ
- fact: âcrawlers tend to experience higher rates of third-party activity than human browser users on loading pages from the same domainsâ
- abstract via Semantic Scholar; PDF not fetched
- Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions, Annamalai, Bilogrevic, De Cristofaro, WWW 2025
- what: â30 participants over 10 weeks, capturing telemetry data from real browsing sessions across 3,000 top-ranked websitesâ
- fact: âautomated crawls miss almost half (45%) of the fingerprinting websites encountered by real users. This discrepancy mainly stems from the crawlersâ inability to access authentication-protected pages, circumvent bot detection, and trigger fingerprinting scripts activated by specific user interactionsâ
- inference: Zeber says crawlers see more third parties, this says crawlers see fewer fingerprinters; both can hold because one compares the same pages and the other compares the pages each actually reaches
- Beyond time delays: How web scraping distorts measures of online news consumption, Ulloa et al., arXiv 2024
- what: compared page content saved in the userâs browser at visit time with content scraped later from the same URLs
- fact: âThe ex-situ collection environment is the primary source of the discrepancies (~33.8%), while the time delays in the scraping process play a smaller role (adding ~6.5 percentage points in 90 days)â
- fact: âat least 33.8% of the contents of participantsâ online news exposure cannot be determined using static web scraping approachesâ
- fact: they name âpages that require user interaction (e.g., paywalls) as the most critical source of content distortionâ
- why it matters here: the only setup paper I found whose metric is page text, not trackers; from communication science, news pages only
- Breaking Bad: Quantifying the Addiction of Web Elements to JavaScript, Fouquet, Laperdrix, Rouvoy, ACM TOIT 2023
- fact: on â6,384 pages, including landing and internal web pagesâ, â43% of web pages are not strictly dependent on JavaScript and that more than 67% of pages are likely to be usable as long as the visitor only requires the content from the main section of the pageâ
- inference: roughly a third of pages lose main content without JavaScript, which is a first-order estimate of what a no-JavaScript corpus such as Common Crawl misses
bot detection
- Cloak of Visibility: Detecting When Machines Browse A Different Web, Invernizzi et al., S&P 2016
- fact: studied âten prominent cloaking services marketed within the undergroundâ, including âIP blacklists that contain over 50 million addresses tied to the top five search engines and tens of anti-virus and security crawlersâ
- classic evidence that malicious sites show crawlers a different page
- Fingerprint Surface-Based Detection of Web Bot Detectors, Jonker, Krumnow, Vlot, ESORICS 2019
- fact: âIn a scan of the Alexa Top 1 Million, we find that 12.8% of websites show indications of web bot detectionâ
- FP-Crawlers, Vastel, Rudametkin, Rouvoy, Blanc, MADWeb 2020
- fact: âWe crawled the Alexa top 10K and identified 291 websites that block crawlers. We show that fingerprinting is used by 93 (31.96%) of themâ
- fact: fingerprinting âcan be bypassed with little effort by an adversary with knowledge on the fingerprints collectedâ
- Web Runner 2049: Evaluating Third-Party Anti-bot Services, Amin Azad, Starov, Laperdrix, Nikiforakis, DIMVA 2020
- fact: âmore than 75% of protected websites in our dataset, successfully defend against attacks by basic bots built with Python scripts or PhantomJSâ; yet âby using less popular browsers in terms of automation (e.g., Safari on Mac and Chrome on Android) attackers can successfully bypass the protection of up to 82% of protected websitesâ
- Good Bot, Bad Bot: Characterizing Automated Browsing Activity, Li, Amin Azad, Rahmati, Nikiforakis, S&P 2021
- the server side view: âmore than 86.2% of bots are claiming to be Mozilla Firefox and Google Chrome, yet are built on simple HTTP libraries and command-line toolsâ
- How gullible are web measurement tools? A case study analysing and strengthening OpenWPMâs reliability, Krumnow, Jonker, Karsch, CoNEXT 2022 (arXiv title: âAnalysing and strengthening OpenWPMâs reliabilityâ)
- fact: âOpenWPM is easily detectableâ; on 100,000 sites âit is commonly detected (âŒ14% of front pages)â
- fact: âWe find several new ways in which a malicious website can attack OpenWPMâs data recordingâ
- they ship a stealth extension; open: nobody tracks whether it still works
- HLISA: towards a more reliable measurement tool, GoĂen, Jonker, Karsch, Krumnow, Roefs, IMC 2021
- fact: âthree methods have been used to detect web bots: browser fingerprint, order of site traversal, and aspects of page interactionâ; HLISA is âan API that simulates interaction like humansâ
- Detecting Bot Detection: Prevalence, Techniques, and Implications for Web Measurement Research, Gundelach, MĂŒhlhauser, Herrmann, arXiv 2026
- what: â10,000 websites across four browser configurations (40K page visits in total)â plus a survey of 81 crawl papers from 2020 to 2025
- fact: â83% of papers omit any discussion of bot detection blockingâ
- fact: âChromium headless encounters a 15% soft block rate compared to 7% for other configurationsâ
- fact: â82% of blocks are attributable to bot detection âŠ, predominantly by providers with integrated bot detection such as Cloudflare (37% block rate) and Akamai (26%)â
- fact: â75% of Chromium-headless-only blocks are caused by header-level signals aloneâ
- fact: âbot detection creates systematic, provider-correlated sample loss that the web measurement community neither measures nor reports. The downstream effect on specific measurement outcomes remains future workâ
- the closest paper to research idea 1 below; one snapshot, one vantage point as far as the abstract says
- the bot side is changing because of LLM agents
- Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents, Ousat et al., arXiv 2026: âchallenge-based defenses are broadly ineffective against commercial solversâ; for score-based defenses the deciding factor is âexecution-environment authenticity, rather than agent behaviorâ
- On the Internet, Nobody Knows Youâre an LLM Bot, Fayolle et al., arXiv 2026: âsome Web Agents were able to bypass all evaluated anti-bot mechanismsâ; âstealth and anti-detection mechanisms often increase detectability rather than decrease itâ
- FP-Agent: Fingerprinting AI Browsing Agents, Wang, Shafiq, Vekaria, arXiv 2026: âFP-Agent detects all seven AI browsing agents, whereas Cloudflare detects only oneâ
- inference: the stealth-plugin advice from 2020 to 2022 may now make a research crawler easier to spot; nobody has tested this on research crawlers specifically
vantage point
- The Blind Men and the Internet: Multi-Vantage Point Web Measurements, Jueckstock et al., arXiv 2019
- what: âsynchronized crawls on the Alexa top 5K domains from four distinct network VPs: research university, cloud datacenter, residential network, and Tor gateway proxyâ
- fact: âsome third-party content consistently avoided crawls from our cloud VPâ; âthe added visibility provided by residential VPs over university VPs is marginal compared to the infrastructure complexity and network fragility they introduceâ
- Do You See What I See? Differential Treatment of Anonymous Users, Khattak et al., NDSS 2016
- fact: âThe second-class treatment of anonymous users ranges from outright rejection to limiting their access to a subset of the serviceâs functionality or imposing hurdles such as CAPTCHA-solvingâ
- A Bestiary of Blocking: The Motivations and Modes behind Website Unavailability, Tschantz et al., FOCI 2018
- fact: âthree forms of server-side blocking: blocking visitors from the EU to avoid GDPR compliance, blocking based upon the visitorâs country, and blocking due to security concernsâ
- Measuring Cookies and Web Privacy in a Post-GDPR World, Dabrowski et al., PAM 2019
- fact: âcollected cookies from the Alexa Top 100,000 websites and compared their cookie behavior from different vantage pointsâ; they âdiscuss challenges caused by these new cookie setting policies for Internet measurement studiesâ
- The Impact of User Location on Cookie Notices, van Eijk, Asghari, Winter, Narayanan, ConPro 2019
- fact: âthe websiteâs Top Level Domain explains a substantial portion of the variance in cookie notice metrics, but the userâs vantage point does notâ; exception: â.com domains from inside versus outside of the EUâ
- Not All Roads Lead to Rome: How VPN Selection Alters What We Measure and Infer about Web Infrastructure, Singh, Ricci, Gamero-Garrido, arXiv 2026
- fact: âthe same country measured through different VPN providers yields materially different conclusions about where endpoints sit, who hosts them, and which physical replicas serve themâ
- fact: âcommercial VPN providers operate their own in-country DNS infrastructure, often intercepting queries regardless of client configurationâ
- inference: âwe crawled from country X through a VPNâ is not one setup but one per provider
- the humanâs BGP background fits here: this is the same âwhich vantage points see the same thingâ question as BGP atoms, asked of web servers and CDNs
consent banners
- We Value Your Privacy ⊠Now Take Some Cookies, Degeling et al., NDSS 2019
- fact: â62.1 % of websites in Europe now display cookie consent notices, 16 % more than in January 2018â
- The Internet with Privacy Policies: Measuring The Web Upon Consent, Jha, Trevisan, Vassio, Mellia, TWEB 2022 (Priv-Accept; in the humanâs notes)
- fact: âall measurements performed not dealing with the Privacy Banners offer a very biased and partial view of the Web. After accepting the privacy policies, we observe an increase of up to 70 trackers, which in turn slows down the webpage load time by a factor of 2x-3xâ
- method fact: keyword matching on the accept button; âThe top-98 keywords cover 95% of the websitesâ
- Exploring the Cookieverse, Rasaii, Singh, Gosain, Gasser, PAM 2023 (BannerClick)
- fact: tool can âdetect, accept, and reject cookie banners with an accuracy of 99%, 97%, and 87%, respectivelyâ
- fact: âbanners to be 56% more prevalent when visiting websites from within the EU regionâ; âwebsites send, on average, 5.5Ă more third-party cookies after clicking âacceptââ
- Thou Shalt Not Reject: Analyzing Accept-Or-Pay Cookie Banners on the Web, Rasaii, Gosain, Gasser, IMC 2023
- fact: âcookiewalls on 0.6% of all queried 45k websitesâ; âfor Germany we see cookiewalls on 8.5% of top 1k websitesâ
- follow-up: To Be or Not to Be (in the EU), Stenwreth, TĂ€ng, Morel, ICISSP 2025: âthe presence of a cookie paywall was most affected by the geographic location of the userâ; a âdouble paywallâ on âapproximately 11% of the studied websitesâ
- Intractable Cookie Crumbs, Rasaii et al., PETS 2025
- what: âstateful crawls on over 20k domains âŠ, strategically accepting banners in the first half of domains and measuring intractable cookies in the second halfâ
- fact: âaround 50% of websites send at least one intractable cookieâ
- why it matters for method: what a crawler did on earlier sites changes what later sites do, so stateless per-site crawls and stateful crawls measure different things
- A Large-Scale Study of Cookie Banner Interaction Tools and their Impact on Usersâ Privacy, Demir, Urban, Pohlmann, Wressnegger, PETS 2024
- fact: âstatistically significant differences in which cookies are set, how many of them are set, and which types are setâeven for extensions that aim to implement the same cookie choiceâ
- inference: âwe accepted the bannerâ is not one setup either; the tool used to click matters
- A Cross-Country Analysis of GDPR Cookie Banners and Flexible Methods for Scraping Them, Nouwens et al., CHI 2025
- fact: âtop 10,000 websites across 31 countriesâ; â67% of websites use consent interfaces, but only 15% are minimally compliantâ; âthree organisations hold 37% of the marketâ
- Leaky Forms, Senol, Acar, Humbert, Zuiderveen Borgesius, USENIX Security 2022
- a good template for reporting setup: âtwo vantage points (EU/US), two browser configurations (desktop/mobile), and three consent modesâ
- fact: emails âexfiltrated ⊠before form submission and without giving consent on 1,844 websites in the EU crawl and 2,950 websites in the US crawlâ
- the crawler âfinds and fills email and password fieldsâ, built on Tracker Radar Collector (per the dataset record)
logins and paywalls
- The Prevalence of Single Sign-On on the Web, Ardi, Calder, IMC 2023 (in the humanâs notes)
- fact: â58% of the top 10K websites with logins are accessible with popular 3rd-party SSO providersâ
- SSO-Monitor, Westers et al., S&P 2024
- fact: âautomatically identified 1,632 websites with 3,020 Apple, Facebook, or Google logins within the Tranco 10kâ; âcan automatically login to each SSO websiteâ
- Shepherd: a generic approach to automating website login, Jonker, Karsch, Krumnow, Sleegers, MADWeb 2020
- fact: âable to automatically log in on 7,113 sitesâ, using âlegitimately crowd-sourced credentialsâ
- The Cookie Hunter, Drakonakis, Ioannidis, Polakis, CCS 2020
- fact: automates âthe challenging process of account creationâ and can âfully audit 25K domainsâ
- To Auth or Not To Auth?, Rautenstrauch et al., S&P 2024
- fact: âthe unauthenticated web could provide a significantly skewed picture of security depending on the type of research questionâ; âthe authenticated state has a larger observable attack surface and more vulnerabilitiesâ
- scale fact: started from â4,485 unique sitesâ, âautomatically extracted more than 900 sites where we detected a login and a registration formâ, ended with âaround 200 sites where our automated login continuously succeededâ
- inference: 200 of 4,485 is the honest yield of login automation in 2023, and the surviving sites are not a random subset
- Keeping out the Masses: Understanding the Popularity and Implications of Internet Paywalls, Papadopoulos et al., WWW 2020
- fact: âpaywall use has increased, and at an increasing rate (2Ă more paywalls every 6 months)â; âpaywalls are in general trivial to circumventâ
- both numbers are from 2019; I found no newer paywall prevalence crawl
can web measurements be repeated and compared
- On the Similarity of Web Measurements Under Different Experimental Setups, Demir et al., IMC 2023
- what: âvisiting 1.7M webpages with five different measurement setupsâ, two of which are identical and run in parallel
- fact: âeven identical setups operating in parallel and visiting the same pages can yield significantly different resultsâ
- fact: âwhen comparing two different profiles, 48% of the underlying data variesâ
- fact: âroughly 60% of the nodesâ children show high similarity, the remaining share shows a substantial fluctuationâ
- fact: they drop âroughly 34% of the pagesâ during vetting, and found â14.6 pages per siteâ on average
- open: the unit is the request tree; it does not say which published conclusions would flip
- Zeber 2020, Jueckstock 2021, Demir 2022 above each quantify one or two factors; Urban 2020 gives the 50% branch churn number
- You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements, Hantke et al., CCS 2023
- idea: crawl the Internet Archiveâs copy instead of the live site, so anyone can rerun the study
- fact: âthe IA is the only archive that is powerful enough ⊠it could provide fresh data for around 2,700 domains in both snapshots (roughly 55%); as for the other archives, the best result was 341 domains (7%)â
- fact: âThe IA aggregates information from multiple sources, hence its point of view of the Web might not reflect any actual vantage pointâ
- validated only for âsecurity headers and JavaScript inclusionsâ
- Internet Jones and the Raiders of the Lost Trackers, Lerner et al., USENIX Security 2016
- fact: the Wayback Machineâs âview of past third-party requests, which we find is imperfect â we evaluate its limitations and unearth lessons and strategies for overcoming themâ
- Jawa: Web Archival in the Era of JavaScript, Goel et al., OSDI 2022
- fact: key observations are âthe forms of non-determinism which impair the execution of JavaScript on archived pagesâ and âthe ways in which JavaScriptâs execution fundamentally differs between live web pages and their archived copiesâ
- fact: âreduces overall storage needs by 41%â
- Mahimahi: Accurate Record-and-Replay for HTTP, Netravali et al., USENIX ATC 2015
- fact: âemulating multiple servers is a key factor in accurately measuring Web page load timesâ
- the performance communityâs answer to repeatability: record once, replay locally
- WebREC (above) is the 2025 version of the same answer for security and privacy work
- Improved methodology for longitudinal Web analytics using Common Crawl, Thompson, WebSci 2024
- fact: âwe have identified the least and most representative segments for a number of recent archivesâ; a segment can stand in for a whole archive
- belongs mostly to the Common Crawl file; listed here because it is a repeatability trick
rules of the road: robots.txt, ethics, and the AI crawler backlash
- Where Are the Red Lines? Towards Ethical Server-Side Scans in Security and Privacy Research, Hantke et al., S&P 2024
- fact: âa slight majority (57%) of operators having a positive stance towards such academic researchâ; proposes âa preregistration processâ
- Web Scraping for Research: Legal, Ethical, Institutional, and Scientific Considerations, Brown et al., arXiv 2024
- fact: âplatforms have greatly restricted access to data through official channels. As a result, researchers will likely engage in more web scrapingâ
- Demir 2022: âmore than half (64.1%) of the analyzed papers omit an ethical discussionâ
- Consent in Crisis: The Rapid Decline of the AI Data Commons, Longpre et al., NeurIPS 2024 (venue from memory)
- fact: âin a single year (2023-2024) there has been a rapid crescendo of data restrictions from web sources, rendering ~5%+ of all tokens in C4, or 28%+ of the most actively maintained, critical sources in C4, fully restricted from useâ
- fact: âThe foreclosure of much of the open web will impact not only commercial AI, but also non-commercial AI and academic researchâ
- Somesite I Used To Crawl, Liu et al., arXiv 2024 (I believe IMC 2025)
- fact: Cloudflare-style blocking has ârelatively limited deployment todayâ but offers âstronger protections against AI crawlersâ
- fact: âif a site implemented active blocking on automated requests (like those of the CC crawler), then Common Crawl may record a 403 Forbidden HTTP status code for those sitesâ
- Scrapers selectively respect robots.txt directives, Kim et al., arXiv 2025 (I believe IMC 2025)
- fact: âbots are less likely to comply with stricter robots.txt directives, and ⊠certain categories of bots, including AI search crawlers, rarely check robots.txt at allâ
- Web Crawler Restrictions, AI Training Datasets & Political Biases, Bouchaud, Ramaciotti, arXiv 2025
- fact: âA quarter of the top thousand websites restrict AI crawlers, decreasing to one-tenth across the broader top millionâ
- fact: â9.5% disallow CCBot, the crawler used by CommonCrawlâ
- fact: â34.2% of news outlets disallow OpenAIâs GPTBot, rising to 55% for outlets with high factual reportingâ; âheterogeneous blocking patterns may skew training datasets toward low-quality or polarized contentâ
- Is Misinformation More Open? A Study of robots.txt Gatekeeping on the Web, Steinacker-Olsztyn, Gosain, Dao, arXiv 2025
- fact: â60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sitesâ; âAI-blocking by reputable sites rising from 23% in September 2023 to nearly 60% by May 2025â
- Do Generative AI Assistants Respect robots.txt?, Lopez-Fonseca et al., arXiv 2026
- fact: some assistants âaccessed restricted resources without requesting robots.txt or used generic user-agents that complicated attributionâ
- inference from the last four: who blocks crawlers correlates with content quality, so any corpus built by a polite, named crawler (Common Crawl included) is tilting toward lower-quality sites over time
- this bears directly on DeGenTWeb-style prevalence numbers; see idea 2
known pitfalls and unsolved problems
- the sample is not the web
- top lists overstate (Scheitle: âoften even an order of magnitudeâ); Tranco random samples understate relative to Common Crawl hosts by about 30% on security metrics (Zhang 2026)
- site counts versus page-load weighting give different stories (Ruth 2022)
- most popular sites are country-specific (Ruth 2022), yet most studies use one global list
- one page per site
- 67.7% of 403 crawl papers visit a single page (Stafeev 2024); internal pages differ (Aqeel 2020, Urban 2020)
- no agreed way to pick internal pages; Hispar depends on a commercial search API
- missing data is not random
- blocked sites cluster by CDN (Gundelach 2026); failed logins cluster by site type (Rautenstrauch 2024: 200 of 4,485 survive); 34% of pages dropped in vetting (Demir 2023)
- 83% of papers do not discuss blocking at all (Gundelach 2026)
- the observer changes the observation
- instrumentation is detectable (VisibleV8: 29% of top 50k probe automation artifacts; Krumnow: OpenWPM detected on about 14%)
- list membership attracts bot traffic (Xie 2024)
- earlier visits change later ones in stateful crawls (Rasaii 2025)
- noise floor
- identical parallel setups disagree (Demir 2023); half of third-party branches change between visits (Urban 2020)
- few papers repeat a crawl or report intervals
- setup is under-described
- 35.4% of navigating papers omit the navigation strategy or page similarity rule (Stafeev 2024); 17.1% omit the crawl technology (Demir 2022)
- âfrom a VPN in country Xâ hides the provider (Singh 2026); âaccepted cookiesâ hides the tool (Demir 2024)
- unsolved, as far as I found
- what blocking, consent state, and login state do to page text and to JavaScript API usage, as opposed to tracker counts
- how detection and blocking of research crawlers changed since the 2023 AI scraper wave; all pre-2023 prevalence numbers (12.8%, 14%, 29%) predate it
- a correction method: given a crawlâs known blind spots, how to adjust the estimate rather than only list âlimitationsâ
- whether archive-based or record-replay crawls agree with live crawls for anything beyond headers and script inclusion
research ideas
- ordering is by how much I would bet on them; confidence is about whether the gap is real, given the limited search noted at the top
- what does crawler blocking do to published numbers, and how fast is it getting worse
- question: when a standard research crawler is blocked or served a degraded page, how far do the usual metrics (third parties, cookies, security headers, JS API usage, page text) move, and how did block rates change from 2022 to 2026
- why not answered: Gundelach 2026 measures block rates once and states âThe downstream effect on specific measurement outcomes remains future workâ; Krumnow 2022, Jonker 2019, Vastel 2020 are pre-AI-scraper snapshots; Liu 2024 and Bouchaud 2025 track AI-crawler rules, not research browsers
- what we would build: paired visits to the same pages with a âknown goodâ client (real headful Chrome on a residential line, human-like interaction) and the common research setups (OpenWPM, Playwright headless and headful, cloud IP); label block / challenge / silent degradation; then recompute two or three well-known results with and without the blocked sites, and with reweighting by CDN
- trend: reuse HTTP Archive and Common Crawl response codes and challenge-page signatures per site over years as a free longitudinal signal (Liu 2024 notes Common Crawl records 403 for blocked sites)
- data and tools: Tranco and CrUX; OpenWPM; Playwright; VisibleV8 to see probing; Zenodo has unreviewed 2026 datasets titled âAnti-Bot Adoption Index â WAF / anti-bot vendor scan of the Tranco top 1Mâ that could seed vendor labels
- main risk: the âknown goodâ client is itself not ground truth; silent degradation is hard to label at scale; residential access raises ethics questions
- confidence the gap is real: medium to high for the downstream-effect part, medium for the trend part (someone may have an ACM-only paper I could not see)
- does a polite no-JavaScript corpus see the same text as a browser, and does the difference bias content studies
- question: for the same URLs, how much of the main text is missing or different in Common Crawl compared with a rendered, consent-accepted browser visit, and is the missing part skewed toward particular kinds of sites
- why not answered: setup-comparison papers measure trackers and requests (Jueckstock 2021, Demir 2022, Demir 2023); Ulloa 2024 measures text but only for news pages in a user panel (âat least 33.8% of the contents ⊠cannot be determined using static web scrapingâ); Fouquet 2023 measures breakage without JavaScript on 6,384 pages, not text and not Common Crawl; Bouchaud 2025 and Steinacker-Olsztyn 2025 show that reputable sites block named crawlers more, but only at the robots.txt level
- why the human should care: DeGenTWeb estimates AI-generated site prevalence from Common Crawl; if high-quality sites opt out of CCBot faster than content farms, the estimate drifts upward for reasons unrelated to AI
- what we would measure: sample Common Crawl URLs stratified by site category; refetch with a browser within days; compare extracted main text (same extractor on both); record robots.txt and 403 status for CCBot versus browser; rerun the LLM-text detector on both versions; report a bias factor per site category
- data and tools: Common Crawl index and WARC, Playwright or Browsertrix, BannerClick or Priv-Accept for consent, trafilatura-style extraction
- main risk: overlap with the Common Crawl workerâs file; time gap between the two fetches confounds (Ulloa puts that at about 6.5 points over 90 days, so keep the gap short)
- confidence the gap is real: medium; the selection-bias angle (who blocks CCBot) I am fairly confident nobody has tied to AI-content prevalence
- how many internal pages per site are enough, and can page templates (web atoms) pick them
- question: for a per-site metric, what is the error from visiting 1, 5, or 20 pages, and does sampling one page per template beat random links or search results
- why not answered: Aqeel 2020 and Urban 2020 show internal pages differ but give no estimator; Stafeev 2024 compares navigation algorithms by code and link coverage; Zhang 2026 is the first rigorous sampling paper but samples sites, not pages within a site; Sprinter 2024 exploits per-site script reuse for speed, not for sampling
- what we would build: for a few hundred sites, crawl deeply (thousands of pages each) once as reference; cluster pages by DOM skeleton and shared resources; simulate sampling strategies and plot error versus pages visited for several metrics; test whether clusters that share structure also change together, which would connect to the web atoms idea in the sibling file
- data and tools: Common Crawl gives many URLs per host for free as the candidate pool; Hispar lists as a baseline; Arachnarium from the SoK
- main risk: the deep reference crawl is itself blocked or rate limited on the sites that matter; âtemplateâ may be too site-specific to generalize
- confidence the gap is real: medium; I searched arXiv titles and abstracts for page-level sampling and found only Zhang 2026
- agent-driven crawling for measurement: what a consent-clicking, logging-in, interacting crawler reveals, and at what cost
- question: how much more of a site (text, JS API calls, third parties) appears when an LLM browser agent accepts consent, signs in via SSO, and uses the page, compared with load-and-wait; and are agent crawls repeatable enough to publish numbers from
- why not answered: YuraScanner 2025 and SpiderSapien 2026 use LLMs for security scanning on self-hosted apps; Bozzolan 2025 uses LLMs to label sites; Rautenstrauch 2024 is semi-manual and ends at 200 sites; Ardi 2023 shows 58% of login sites accept SSO but does not crawl behind it; Annamalai 2025 shows real users hit 45% more fingerprinting sites, which is the size of the prize
- what we would build: an agent harness on Playwright with engine-level logging (VisibleV8 or WebREC) so the observation does not depend on the agent; run load-only, scripted (BannerClick + SSO-Monitor), and agent modes on the same sites; repeat each three times to measure run-to-run spread
- fit with JSphere: API usage after interaction is the natural next question
- main risk: agents are now a detection target (FP-Agent 2026: classifier âdetects all seven AI browsing agentsâ; Fayolle 2026), so the agent may be blocked more than a plain crawler; account creation and terms of service; cost per site
- confidence the gap is real: medium to high today, but this is an obvious next step for the CISPA and Caâ Foscari groups, so the window is short
- a variance budget for crawl metrics beyond trackers
- question: for a given metric, how much of the spread comes from time, vantage point, browser, consent state, and plain chance, and how many repeats make a difference between two crawls believable
- why not answered: Zeber 2020 isolates âtime, cloud IP address vs. residential, and operating systemâ and Demir 2023 compares five profiles, both on tracking and request-tree metrics; neither gives a recipe (repeats needed, minimum detectable difference) and neither covers page text or JS API sets
- what we would do: a full factorial crawl on a few thousand pages with replication; fit a variance-components model per metric; publish the table and a calculator
- main risk: reviewers may see it as incremental over Demir 2023; needs a sharp demonstration, such as a published longitudinal trend that falls inside the noise
- confidence the gap is real: low to medium for tracker metrics, medium for content and API metrics
- a shared calibration panel across studies (weakest; listed so the human can discard it knowingly)
- question: if every crawl study also crawled the same few hundred pages, could results from different papers be put on one scale
- nearest work: Trancoâs list IDs fix the site sample; WebRECâs .web archives fix the recording; nothing fixes the comparison across setups
- main risk: adoption problem rather than research problem; may only work as part of idea 5
- confidence: low
opinions from ChatGPT
- I sent the six ideas above to ChatGPT (Extra High) for a novelty check; no answer had arrived when I wrote this file, so nothing here reflects it
gaps in this review
- not fetched as PDF, abstract only: Ahmad 2020, Zeber 2020, Jueckstock 2021, Xie 2024, HLISA 2021, Shepherd 2020, Cookie Hunter 2020, and the Springer papers
- not covered: McDonald et al. â403 Forbidden: A Global View of CDN Geoblockingâ (IMC 2018) and Vallina et al. on domain classification services (IMC 2020), because I could not reach a copy to quote
- thin: HTTP Archive and CrUX as crawl infrastructure (I found no methods paper to quote), mobile app webviews, residential proxy ethics, and the classic 1998 to 2005 crawl-ordering papers beyond what the Olston and Najork survey covers
- venues marked âI believeâ or âfrom memoryâ need a check before citing
Last edited: