how to fight content theft
(authored by agents unless marked đ§)
short version
- words
- content theft: someone takes what you made and uses it without asking, crediting, or paying
- old kind: a person or site copies your article, photo, video, app, or code and shows it as theirs
- new kind: an AI company downloads your work to train a model, or to answer questions with it so nobody visits you
- crawler: a program that downloads web pages in bulk
- robots.txt: a text file on a site that tells crawlers what they may fetch; obeying it is voluntary
- opt-out: any signal that says âdo not use my work for AIâ
- what the literature already shows
- telling crawlers to stay out works only on those who choose to obey; AI search tools obey least
- changing your images so models cannot learn from them (Glaze, Nightshade) is beaten by cheap tricks
- nobody can yet prove from the outside that a real commercial model trained on an ordinary web page
- US courts so far call training fair use; the big payout was for pirated copies, not for training
- AI answers send far fewer visitors back than search did, and blocking the AI crawler can cost you visibility
- for the old kind of theft, finding copies is easy; takedown notices are easy to abuse and dates are easy to fake
- what nobody has measured, as far as I found
- whether opted-out content still reaches AI answers by a side road, e.g. through a search engineâs index
- whether opting out keeps you out of the next model, tested with planted marker text
- whether the newer signals and traps (content signals, RSL, AI Labyrinth, Anubis) change crawler behavior
- how many cloaked images are really on the web
- how often a copy outranks its original in search
- my three strongest ideas, details under âresearch ideasâ
- 1: side-road audit: plant pages that say no to AI but yes to search, then see which AI products still know their secrets
- 2: marker experiment: plant marker text under different opt-out settings and test each new model for it
- 3: thief above author: measure how often a copied or AI-rewritten article outranks its original in search
- the rankings are my opinion; no prototype or pilot exists, and no second agent reviewed them because the Claude usage limit hit
scope and evidence
- reviewed 7 Oct 2026 UTC
- đ§ the question, from the humanâs research index: âhow to fight content theftâ
- đ§ related line in the humanâs photo authentication note: âopting out of ML trainingâ
- evidence labels used below
- read: I, or a helper agent, read the relevant sections of the full text
- abstract: only the abstract was read; a small model fetched the page, and I trust its quoted wording but did not see the raw page
- fetched: same as abstract, for a web page that is not a paper
- snippet: only a search result summary was seen; check before relying on it
- coverage is incomplete
- the web search tool ran out of its shared budget twice, so newer papers and citation chains are thin
- not searched at all: video and reupload theft on YouTube and TikTok, cloned mobile apps, print-on-demand art theft, paywall bypass services, residential proxy scraping
- novelty of the ideas below is therefore unconfirmed
- ChatGPT consultation was not possible: the tool failed with âaccount_ui_login_requiredâ
- related studies in this tree, not repeated here
- crawling, including robots.txt and the AI crawler backlash
- image watermarking and text watermarking
- C2PA and photo provenance
- search spam, which covers content farms
the fight has five places
- 1: at the door: tell or force crawlers to stay out
- 2: in the content: change it so a copy is useless
- 3: after the fact: prove a model used it, or find the copy and prove you were first
- 4: in court and in law
- 5: in the market: get paid instead
1: at the door
who says no
- Longpre et al., Consent in Crisis, NeurIPS 2024, abstract
- robots.txt and terms of service of 14,000 web domains used in AI training sets, over time
- âin a single year (2023-2024) there has been a rapid crescendo of data restrictions from web sources, rendering ~5%+ of all tokens in C4, or 28%+ of the most actively maintained, critical sources in C4, fully restricted from useâ
- âgeneral inconsistencies between websitesâ expressed intentions in their Terms of Service and their robots.txtâ
- Fletcher, Reuters Institute, How many news websites block AI crawlers?, 2024, fetched
- 15 most-used news sites in each of 10 countries, robots.txt from the Wayback Machine for every day of 2023
- â48% of the most widely used news websites across ten countries were blocking OpenAIâs crawlersâ; 24% blocked Googleâs AI crawler
- from 79% in the US down to 20% in Mexico and Poland
- Steinacker-Olsztyn, Gosain, Dao, Is Misinformation More Open?, WWW 2026, abstract
- â60.0% of reputable sites disallow at least one AI crawler, compared to just 9.1% of misinformation sitesâ
- âAI-blocking by reputable sites rising from 23% in September 2023 to nearly 60% by May 2025â
- my reading: good sources leave the training data faster than bad ones
- individual creators mostly cannot say no
- Liu et al., Somesite I Used To Crawl, IMC 2025, read
- survey of 203 professional artists, 1,182 artist sites, crawler tests on their own sites
- â59% of artists have never heard about the term ârobots.txtââ
- most artists use hosting services that do not let them edit robots.txt
- Squarespace has a one-click switch, yet âonly 49 (17%) of the 293 artists who use Squarespace had enabled this optionâ
- âthere are no existing standard mechanisms for explicitly controlling whether publicly accessible Web content is used in training AI modelsâ
- Liu et al., Somesite I Used To Crawl, IMC 2025, read
who obeys
- Kim et al., Scrapers Selectively Respect robots.txt Directives, IMC 2025, read
- changed robots.txt on their universityâs sites and watched who obeyed
- âapproximately 3.9 million external web requests made to a set of 36 websites from February 12 - March 29, 2025â
- â130 self-declared bots (and many anonymous ones) over 40 daysâ
- âbots are less likely to comply with stricter robots.txt directives, and that certain categories of bots, including AI search crawlers, rarely check robots.txt at allâ
- some never fetched the file: â9/34 for the crawl delay experiment, and 15/47 for both the endpoint access and the disallowall experimentsâ
- limits: âonly considered traffic from 36 websites owned by our institutionâ; could not settle whether a bot lies about its name
- changed robots.txt on their universityâs sites and watched who obeyed
- Liu et al., same paper, read
- âmost large AI companies currently do respect robots.txt. However, a number of AI-powered apps and crawlers do not respect it (including crawlers from ByteDance)â
- Lopez-Fonseca et al., Do Generative AI Assistants Respect robots.txt?, arXiv July 2026, abstract
- ten AI assistants with web search, own test pages with âsecret codes embedded in target pagesâ, 200 trials
- âSome systems followed the expected allowed/disallowed access pattern, whereas others accessed restricted resources without requesting robots.txt or used generic user-agents that complicated attributionâ
- this is the closest prior work to idea 1; it tests direct fetches only
- Cloudflare, Perplexity is using stealth, undeclared crawlers, Aug 2025, fetched
- âWe created multiple brand-new domainsâ, with robots.txt that banned all bots; Perplexity still answered questions about them
- Perplexityâs named crawler made 20â25 million requests a day; the unnamed one, 3â6 million
- OpenAIâs fetcher âfetched the robots file and stopped crawling when it was disallowedâ
- limits: one vendorâs blog; Perplexity disputed it, I did not check how
- JaĆșwiĆska and Chandrasekar, Tow Center, AI Search Has A Citation Problem, Columbia Journalism Review, March 2025, fetched
- âPerplexityâs free version correctly identified all ten excerpts from paywalled articles we shared from National Geographic, even though the publisher has disallowed Perplexityâs crawlersâ
- the route is unknown; a side road is one explanation
- industry numbers, not peer reviewed
- TollBit, first half of 2026, as reported by Relevant Audience, fetched: âroughly 15 percent of identified AI page fetching agents reached URLs that robots.txt disallowedâ on European sites
- the same article says OpenAIâs documentation states that ârobots.txt rules may not applyâ when a ChatGPT user starts the request
- BuzzStream, March 2026, as reported by PPC Land, fetched: top 50 news sites that block AI crawlers, â4 million citations across 3,600 promptsâ
- âRoughly 70% of all ChatGPT citations in the dataset came from sites that block ChatGPTâs retrieval botsâ
- it lists four possible roads and settles none: indexed before the block, Common Crawl archives, bots ignoring the block, and AI products that âpull data directly from search engine results pagesâ
- limit: a marketing companyâs study; being cited is not the same as the page text being read
- TollBit, first half of 2026, as reported by Relevant Audience, fetched: âroughly 15 percent of identified AI page fetching agents reached URLs that robots.txt disallowedâ on European sites
- side roads alleged in court, snippet
- Reddit v. SerpApi and Perplexity, S.D.N.Y.: Reddit says Perplexity got Reddit text out of Google results through a scraping company; motion to dismiss âlargely deniedâ 31 July 2026
what crawlers cost
- Wikimedia Foundation, How crawlers impact the operations of the Wikimedia projects, April 2025, fetched
- âSince January 2024, we have seen the bandwidth used for downloading multimedia content grow by 50%â
- âAt least 65% of this resource-consuming traffic we get for the website is coming from botsâ; âThe overall pageviews from bots are about 35% of the totalâ
- Cloudflare, AI Labyrinth, March 2025, fetched: âAI Crawlers generate more than 50 billion requests to the Cloudflare network every day, or just under 1% of all web requestsâ
stronger doors
- blocking by a reverse proxy, a service that sits in front of the site
- Liu et al., read: of top 10k sites on Cloudflare, âonly 107 (5.7%) sites enable Cloudflareâs Block AI Bots optionâ; âCloudflareâs feature blocks 17 AI user agentsâ
- it cannot âstop AI training for Meta, Google, and Webzioâ: they use one crawler for search and AI, so blocking it also removes you from search
- Cloudflare, Content Independence Day, 1 July 2025, fetched: âchanging the default to block AI crawlers unless they pay creators for their contentâ
- proof of work: make every visitorâs browser solve a small puzzle first
- Anubis, open source, 23.1k GitHub stars, fetched
- âAnubis is a bit of a nuclear response. This will result in your website being blocked from smaller scrapers and may inhibit âgood botsâ like the Internet Archiveâ
- no measurement paper found on whether it stops AI crawlers or what it costs real visitors
- Anubis, open source, 23.1k GitHub stars, fetched
- traps: feed a misbehaving crawler endless made-up pages
- Cloudflare AI Labyrinth, fetched: âAI-generated set of linked pages when we detect inappropriate bot activityâ
- the page gives no numbers on whether it works; open source tarpits Nepenthes and iocaine not searched
- new signals, all still requests that nothing enforces
- Dinzinger, HeĂ, Granitzer, A Survey of Web Content Control for Generative AI, arXiv 2024, abstract: site owners âare overwhelmed by the multitude of recent ad hoc standardsâ
- Cloudflare content signals, fetched: lines in robots.txt such as
Content-Signal: search=yes, ai-train=no; Cloudflare says over 3.8 million domains use its managed robots.txt - IETF AIPREF: a draft standard vocabulary for such preferences; version 07, Aug 2026, not final
- RSL, Really Simple Licensing 1.0, Dec 2025, fetched: a site states license and price terms in robots.txt; backed by Cloudflare, Akamai, Fastly, Reddit and others; the page names no AI company that honors it
- Web Bot Auth, an IETF working group: bots sign their requests so a site knows who is really asking
- Liu et al., read: the
noaipage tag is nearly unused; among the top 10k sites, âonly 17 sites having noaiâ - đ§ from the humanâs C2PA paper notes, on Keller and Warso 2023: âC2PA has entry to opt out of ML trainingâ
- EU rule that gives opt-outs teeth, fetched as a summary only, check the wording
- General-Purpose AI Code of Practice, copyright chapter, measure 1.3: signersâ crawlers must follow robots.txt as in RFC 9309 and other widely adopted opt-out signals
- measure 1.2: no getting around paywalls; leave out known piracy sites
- the code is voluntary; I found no audit of whether signers keep it
2: in the content
image cloaks
- Shan et al., Glaze, USENIX Security 2023, abstract
- adds a small change to an artwork so that a model trained on it copies the style badly
- âeven at low perturbation levels (p=0.05), Glaze is highly successful at disrupting mimicry under normal conditions (>92%) and against adaptive countermeasures (>85%)â
- Shan et al., Nightshade, IEEE S&P 2024, abstract
- images that look normal but teach the model the wrong thing about a word
- âcan corrupt an Stable Diffusion SDXL prompt in <100 poison samplesâ
- âa last defense for content creators against web scrapers that ignore opt-out/do-not-crawl directivesâ
- limit: tested on models the authors trained; I found no public sign that a commercial model was hurt
- how many people use them, snippet: MIT Technology Review, Sept 2024, âGlaze has been downloaded nearly 3.5 million times (and Nightshade over 700,000)â
- downloads are not users, and users are not cloaked images on the web
- same idea for web text: Liu et al., ExpShield, NDSS 2026, abstract
- invisible changes to a pageâs text so a model trained on it memorizes less
- âthe Membership Inference Attack (MIA) AUC drops from 0.95 to 0.55 under the defenseâ
- tested on models the authors trained; I did not find an attack paper on it yet
- same idea elsewhere, all snippet: faces (PhotoGuard 2023, Anti-DreamBooth ICCV 2023, MetaCloak CVPR 2024), music (HarmonyCloak, IEEE S&P 2025), code (CoProtector, WWW 2022)
the cloaks get broken
- Hönig, Rando, Carlini, TramÚr, Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI, 2024 (I believe ICLR 2025), abstract
- âlow-effort and âoff-the-shelfâ techniques, such as image upscaling, are sufficient to create robust mimicry methods that significantly degrade existing protectionsâ
- âthey only provide a false sense of securityâ
- Foerster et al., LightShed, USENIX Security 2025, abstract
- learns the cloak pattern from public cloaked examples, detects it, removes it
- âa TPR of 99.98% and TNR of 100% on detecting NightShadeâ: it catches nearly every poisoned image and flags no clean one
- useful for us: a detector for cloaked images exists
- Cao et al., IMPRESS, NeurIPS 2023, abstract: cleans an image by making it agree with its own rebuilt version
- Radiya-Dixit et al., Data Poisoning Wonât Save You From Facial Recognition, ICLR 2022, abstract
- the root problem: you cloak a picture once, then every later model gets a try
- the Glaze teamâs view, Glaze FAQ, fetched
- asked whether Glaze has been broken: âNo, it has notâ
- âAt this time, we do not believe Glaze provides consistent protection against img2img attacks, including style transfer and inpaintingâ
- my take: cloaking is a losing race for the artist, so I would not build a new cloak; measuring cloaks in the wild is still open
3: after the fact
can you prove a model trained on your work
- guessing from the modelâs confidence does not work
- membership inference: ask whether the model is oddly sure about your text
- Duan et al., Do Membership Inference Attacks Work on Large Language Models?, COLM 2024, abstract
- âMIAs barely outperform random guessing for most settings across varying LLM sizes and domainsâ
- earlier wins came from test sets where trained-on and not-trained-on texts differed in date
- Zhang, Das, Kamath, TramĂšr, Membership Inference Attacks Cannot Prove that a Model Was Trained On Your Data, SaTML 2025, abstract
- âfundamentally unsoundâ: you cannot measure how often the test cries wolf, because you cannot retrain the model without your data
- what is sound: âdata extraction attacks and membership inference on special canary dataâ
- canary: made-up marker text that exists nowhere else
- Maini et al., LLM Dataset Inference, NeurIPS 2024, abstract: tests a whole collection instead of one text; works on open models with known training data, at weak confidence, âp-values < 0.1â
- planting markers before you publish does work, in the lab
- Wei, Wang, Jia, Proving membership in LLM pretraining data via data watermarks, Findings of ACL 2024, abstract
- insert random strings or look-alike Unicode letters; because you picked them at random, the false alarm rate is known
- works âprovided that the rightholder contributed multiple training documents and watermarked them before public releaseâ
- on a real large model: âwe can robustly detect hashes from BLOOM-176Bâs training data, as long as they occurred at least 90 timesâ
- Meeus et al., Copyright Traps for Large Language Models, ICML 2024, abstract
- âeven medium-length trap sentences repeated a significant number of times (100) are not detectable using existing methods. However, we show that longer sequences repeated a large number of times can be reliably detected (AUC=0.75)â
- tested on a 1.3B model they trained themselves
- Sander et al., Watermarking Makes Language Models Radioactive, NeurIPS 2024, abstract: training on watermarked model output leaves a trace, found âeven when as little as 5% of training text is watermarkedâ
- for images: Bouaziz, Usunier, El-Mhamdi, Data Taggants, ICLR 2025, abstract: slightly altered images make a trained model answer secret key images in a known way, giving âstatistical certificates with black-box access onlyâ
- Cui, Wei, Swayamdipta, Jia, Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge, Findings of ACL 2025, abstract
- the marker is a made-up fact about a made-up thing, written as normal prose, so data cleaning does not throw it out
- âour data watermarks can be evaluated even under API-only access via question answeringâ: you just ask the model about the made-up thing
- Weinberg, SIGIL, arXiv 2026, abstract: five kinds of marker; results come from a simulator, not from trained models
- also seen by title only: SPECTRA (arXiv 2512.17075), markers for retrieval systems (arXiv 2502.10673)
- the gap: I found no test of an ownerâs planted marker against a commercial model
- Wei, Wang, Jia, Proving membership in LLM pretraining data via data watermarks, Findings of ACL 2024, abstract
- pulling your text back out works for famous works
- Nasr et al., Scalable Extraction of Training Data from (Production) Language Models, arXiv 2023, abstract: an attack makes ChatGPT âemit training data at a rate 150x higher than when behaving properlyâ
- Ahmed, Cooper, Koyejo, Liang, Extracting books from production language models, arXiv Jan 2026, abstract
- âit was unnecessary to jailbreak Gemini 2.5 Pro and Grok 3 to extract text (e.g, nv-recall of 76.8% and 70.3%, respectively, for Harry Potter and the Sorcererâs Stone)â
- nv-recall: the share of the book that came back nearly word for word
- GPT-4.1 âeventually refuses to continue (e.g., nv-recall=4.0%)â
- Chen et al., CopyBench, arXiv 2024, abstract: bigger models copy more, âliteral copying rates increasing from 0.2% to 10.5%â from Llama3-8B to 70B
- Carlini et al., Extracting Training Data from Diffusion Models, 2023, abstract: âwe extract over a thousand training examples from state-of-the-art modelsâ
- Xu et al., LiCoEval, ICSE 2025, abstract: 14 code models produce âa non-negligible proportion (0.88% to 2.01%) of code strikingly similar to existing open-source implementationsâ, mostly without the right license
- limit: famous books appear thousands of times in training data; an ordinary web page appears once
do AI answers credit you
- Liu, Zhang, Liang, Evaluating Verifiability in Generative Search Engines, Findings of EMNLP 2023, abstract: âa mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentenceâ
- Tow Center, March 2025, fetched: eight AI search tools, 1,600 queries asking where a news quote came from
- âprovided incorrect answers to more than 60 percent of queriesâ
- âMore than half of responses from Gemini and Grok 3 cited fabricated or broken URLsâ
finding copies made by people and sites
- matching text is old and cheap
- Broder, On the resemblance and containment of documents, 1997, read by the earlier Codex agent
- compares small random samples of word runs instead of whole pages; tells âroughly the sameâ from âroughly containedâ
- Schleimer, Wilkerson, Aiken, Winnowing, SIGMOD 2003, read by the earlier Codex agent
- catches any shared run longer than a set length, âincluding small partial copiesâ
- breaks under translation or heavy rewriting, which is what an AI rewrite does
- a match is not a verdict; Moss: âthe scores are certainly not a proof of plagiarismâ
- Broder, On the resemblance and containment of documents, 1997, read by the earlier Codex agent
- matching images is easy to dodge
- Jain, Cretu, de Montjoye, Adversarial Detection Avoidance Attacks, USENIX Security 2022, abstract
- perceptual hash: a short fingerprint that stays the same when a picture changes a little
- âmore than 99.9% of images successfully attacked while preserving the content of the imageâ
- Jain, Cretu, de Montjoye, Adversarial Detection Avoidance Attacks, USENIX Security 2022, abstract
- AI rewrites at scale
- NewsGuard, AI Tracking Center, fetched, updated 23 June 2026: âidentified 3,749 AI Content Farm news and information websitesâ
- Puccetti et al., AI âNewsâ Content Farms Are Easy to Make and Hard to Detect, ACL 2024, abstract: âthere are currently no practical methods for detecting synthetic news-like texts âin the wildââ
- search engines say they punish copies; nobody outside checks
- Google, spam policies, fetched: âWhen we receive a significant volume of valid copyright removal requests involving a given site, we are able to use that to demote other content from the site in our resultsâ
- Google, How Google Fights Piracy, 2018, read: âdemoted sites lost an average of 89% of their traffic from Google Searchâ
- I found no independent count of how often a copy outranks its original
takedown
- how it works in the US: you send a notice, the host or search engine removes the link, the other side may object
- notices are often wrong or abused
- Google 2018, read: âContent owners have notified us about 882 Million URLs in 2017 alone. Google removed more than 95% of these webpages, meaning we pushed back on around 54 million removal requests that were incomplete, mistaken, or abusiveâ
- Urban, Karaganis, Schofield, Notice and Takedown in Everyday Practice, 2016, snippet; the paper could not be opened
- about 30% of sampled notices to Google web search raised questions about validity; 70% in a Google image search sample
- Harvard Law Today on Eugene Volokhâs work with the Lumen notice archive, Shedding light on fraudulent takedown notices, fetched
- âclose to 200 out of 700 court orders submitted to Google were ⊠âeither obviously forged or fraudulent or at least highly suspicious cases,â including at least 80 outright forgeriesâ
- back-dating trick, snippet, from a Lumen blog post I could not open: copy a real article, give the copy an earlier date, then report the real one as the copy; about 34,000 such notices from June 2019 to January 2022
- YouTube Content ID, Google 2018, read: âFewer than 1% of Content ID claims are disputed, and of that number, over 60% resolve in favor of the uploaderâ
- takedown removes a link, the copy lives on
- U.S. Copyright Office, Section 512 Report, 2020, read by the earlier Codex agent: pp 54â55 separate removal from stopping the next upload
- EU: platforms must log every removal in a public database
- Kaushal et al., Automated Transparency, FAccT 2024, abstract: â131mâ records from November 2023; âcompliance remains problematicâ; not split out for copyright
- proving who was first
- OpenTimestamps, fetched: âA timestamp proves that some data existed prior to some point in timeâ
- it helps only if you stamped at publish time, and it shows the file existed, not who wrote it
- I found no study of timestamps used in takedown disputes
piracy
- Rafique et al., Itâs Free for a Reason, NDSS 2016, read: free sports streaming sites
- âmore than 850,000 visits by identifying 5,685 free live streaming domainsâ
- âon average, 50% of the time, a click on an overlay ad leads the user to a malware-hosting webpageâ
- âmore than 30% have been reported at least once by copyright ownersâ
- Himmelstein et al., Sci-Hub provides access to nearly all scholarly literature, eLife 2018, read: âSci-Hubâs database contains 68.9% of the 81.6 million scholarly articles registered with Crossrefâ
- Ahn et al., Watch Out Your TV Box, USENIX Security 2025, abstract: took apart a pirate TV box and its peer-to-peer network; â131,175 unique users and 78 serversâ in two months
- Roudot and Sabt, Narrowbeer, USENIX Security 2025, abstract: breaks Widevine, the lock on most streaming video, so that licenses never expire
- piracy now feeds AI: see Bartz v. Anthropic below
4: in court and in law
- US: training looks legal so far; how you got the copies matters
- Bartz v. Anthropic, N.D. Cal.
- June 2025 ruling, snippet: training on books is fair use, âtransformativeâspectacularly soâ; keeping a library of pirated books is not
- settlement, Authors Guild, fetched: âthe landmark $1.5 billion class action settlementâ, final approval âOn July 20, 2026â
- about $3,000 per book, âfour times the statutory minimum for ordinary infringementâ
- covers only âAnthropicâs past acquisition and copying of their worksâthe âinputsâ sideâthrough August 25, 2025â; claims about what the model outputs stay open
- Kadrey v. Meta, June 2025, snippet: fair use for training; the claim about sharing pirated books by torrent continues
- Thomson Reuters v. ROSS, 3rd Circuit, 29 Sept 2026, Ballard Spahr summary, fetched: not fair use, but for a search tool that competed directly; âROSSâs AI platform cannot generate original expressionâ
- New York Times v. OpenAI and Disney v. Midjourney: pending, snippet
- U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training, pre-publication May 2025, read
- p 106: âthe Office recommends allowing the licensing market to continue to develop without government intervention. If market failures are shown as to specific types of works in specific contexts, targeted intervention such as ECL should be consideredâ
- ECL, extended collective licensing: one body licenses a whole class of works for everyone, including owners who never signed up
- Bartz v. Anthropic, N.D. Cal.
- Europe: courts disagree on whether a model contains a copy, snippet
- GEMA v. OpenAI, Munich, Nov 2025: song lyrics a model can recite count as a copy inside the model
- Getty v. Stability AI, UK High Court, Nov 2025: the model is not an âinfringing copyâ
- EU law lets owners reserve their rights in machine-readable form, which is why robots.txt and its successors matter legally there
- my take: courts reward evidence of how content was obtained and of word-for-word output; both are things a measurement person can produce
5: in the market
- deals exist for the big, snippet, press reports: News CorpâOpenAI âmore than $250 millionâ over five years; RedditâGoogle about $60 million a year
- small sites get tools that are not open to them yet
- Cloudflare pay per crawl, fetched: the site answers a crawler with â402 Payment Requiredâ and a price; âPay per crawl is currently in closed betaâ
- no public numbers on money paid
- visitors are falling
- Pew Research Center, Google users are less likely to click on links when an AI summary appears, July 2025, fetched
- 900 US adults, 68,879 Google searches in March 2025
- âUsers who encountered an AI summary clicked on a traditional search result link in 8% of all visits. Those who did not encounter an AI summary clicked on a search result nearly twice as often (15% of visits)â
- a link inside the summary got a click âin just 1% of all visitsâ
- Khosravi and Yoganarasimhan, Impact of AI Search Summaries on Website Traffic, arXiv 2026, abstract
- compares English Wikipedia articles with the same articles in German and French while Googleâs AI summaries reached countries at different times
- âdefault AIO availability reduced English search traffic by 5.45% and 4.82%, respectivelyâ
- del Rio-Chanona, Laurentsyeva, Wachs, Large language models reduce public knowledge sharing on online Q&A platforms, PNAS Nexus 2024, fetched: âWithin 6 months of ChatGPTâs release, activity on Stack Overflow decreased by 25% relative to its Russian and Chinese counterpartsâ
- Cloudflare, crawl-to-refer ratio, July 2025, fetched: pages crawled per visitor sent back; 70,900 to 1 for Anthropic in one June 2025 week
- Cloudflareâs own caveat: visits from apps carry no referrer, so the numbers âmay overstate the respective ratios, but it is unclear by how muchâ
- Pew Research Center, Google users are less likely to click on links when an AI summary appears, July 2025, fetched
- saying no has a price
- Grossman et al., How Generative AI Disrupts Search, SIGIR 2026, abstract: 11,500 queries
- âfor 51.5% of representative, real-user queries, AIOs are generated, and are displayed above the organic search resultsâ
- âwebsites that block Googleâs AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the contentâ
- it is a correlation; sites that block may differ in other ways
- Zhao and Berman, The Impact of LLMs on Online News Consumption and Production, arXiv Dec 2025, snippet; the page fetch failed, so check the wording
- compares large news publishers that blocked AI bots in robots.txt with ones that did not, before and after
- blocking reduced âtotal website traffic by 23% and real consumer traffic by 14% compared to not blockingâ
- about 30 large newspaper sites; small sites not covered
- Grossman et al., How Generative AI Disrupts Search, SIGIR 2026, abstract: 11,500 queries
research ideas
what I would not do
- a new cloak or poison: attackers get the last move
- a new membership inference test: the best paper on it says the approach cannot give proof
- law or pricing: no systems contribution
idea 1: side-road audit of AI opt-outs
- question: when a site says no to AI but yes to search, which AI products still know what is on it, and by which road
- why open
- Lopez-Fonseca 2026 and Cloudflare 2025 test only whether the AI product fetches the page itself
- BuzzStream 2026 sees blocked news sites cited anyway and cannot tell which road the content took
- Redditâs lawsuit and the Tow Center paywall finding point at side roads: search engine indexes, scraping companies, shared caches
- no audit found of content signals, RSL, or the EU codeâs measure 1.3 promise
- build
- a few hundred fresh domains, each page holding a unique secret code, as Lopez-Fonseca did
- vary one setting per group: robots.txt for AI bots, content signals, RSL terms, a login wall, a proxy block
- let search engines index all of them
- ask each AI assistant and AI search product about each page every month
- record all site requests and acquired markers
- distinguish observed direct requests, possible indirect access, and unresolved paths
- an absent recognized crawler cannot exclude another client identity or intermediary
- first step: 20 domains, 5 products, one month
- possible output: product-specific observations under each signal and known access condition
- marker acquisition alone does not establish ignored terms or a particular collection path
- risks
- fresh domains may never get indexed or asked about; seed them with links
- products change weekly, so results are a snapshot; turn that into a monthly tracker
- cheap add-on: put a tarpit and Anubis on some groups and report who falls in and what it costs real browsers; nobody has published that
idea 2: marker experiment, does opting out keep you out of the next model
- question: if I plant marker text today, does it show up in models released next year, and does an opt-out prevent that
- why open
- Zhang 2025: planted markers are one of the few sound proofs
- the cited Wei 2024 and Meeus 2024 experiments used controlled models or BLOOM
- this reading did not establish whether commercial opt-out marker comparisons already exist
- perform a fresh closest-work check before claiming a gap
- build
- reuse idea 1âs domains; each group gets its own random marker strings, repeated on many pages
- Wei 2024 needed about 90 copies in a 176B-parameter model, so plan for hundreds of copies per group
- use Cui 2025âs made-up facts as markers: they survive data cleaning and can be tested by plain questions
- at each new model release, ask about each groupâs made-up facts and about control facts never published; controls give the false alarm rate
- disable exposed web-search features and record the settings
- hidden retrieval, caches, and copied pages can still explain marker acquisition
- claim training membership only when alternative access paths are independently excluded
- possible contribution: measured opt-out differences under known collection and model-access conditions
- commercial marker experiments require a fresh closest-work check
- risks
- a year or more of waiting
- a few hundred small new sites may be too little text for any model to remember
- a miss proves nothing: the marker may just be too weak
- I would start it early because it only costs waiting
idea 3: the price of saying no
- question: what does a site lose in search and AI visibility when it blocks AI crawlers
- why open: Grossman 2026 shows a correlation for Google AI summaries only; owners decide today without numbers
- build
- find sites whose robots.txt started or stopped blocking AI crawlers, from Common Crawl and Wayback Machine history
- compare their visibility before and after against similar sites that did not change: rank lists such as Tranco, how often AI summaries and assistants cite them for a fixed query set
- on own test sites from idea 1, flip the setting on purpose and watch
- why it sells: every publisher asks this question
- risks: public traffic data is coarse; sites that change robots.txt often change other things at once
other ideas, weaker or less checked
- count cloaked images in the wild
- run a LightShed-style detector over images from artist sites and public image crawls; join with the siteâs robots.txt
- answers âdo artists who cloak also opt outâ and âhow much poison is out thereâ; no prior count found
- needs GPU time and LightShedâs code or a rebuild
- how often does the copy outrank the original
- take new articles from feeds, find copies with Broder-style matching plus a meaning-based match for AI rewrites, check search ranks over weeks
- overlaps with the search spam study
- proof of first publication
- stamp pages at publish time, then replay known back-dating cases from the Lumen notice archive to see how many it would have settled
- needs researcher access to Lumen
- carried over from the earlier Codex agent
- follow copies after takedown: with 20 cooperating creators, check weekly for three months whether removed copies return under new addresses
- a copy-evidence tool that shows matched passages and dates, and counts false accusations against quotes and licensed reprints
- check Cloudflareâs crawl-to-refer numbers against sites that log both crawls and real visits, to learn how much app traffic the referrer misses
8 October marker-test constraint
- a returned marker establishes that the product acquired it through some path
- turning off visible search does not prove that hidden retrieval, cached answers, or copied pages are absent
- training membership needs controlled training data or independently established access restrictions
- compare publication with blocking from the start against blocking after indexing
- retain server logs and query histories, but missing fetches do not rule out indirect access
- keep never-published markers and record unintended copies
- a missing marker remains weak negative evidence
- closest controlled-training work already appears above
- Meeus et al., Copyright Traps, ICML 2024
- Wei et al., data watermarks, Findings of ACL 2024
- reproduce those controlled conditions before claiming training detection from a live commercial assistant
Last edited: