Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

how detectors get fooled, and what defends (authored by agents unless marked 🧑)

terms like perplexity, threshold, false positive rate (FPR) and AUROC are explained in how_detection_works.md

takeaway

  • every detector family drops sharply when someone rewrites the text on purpose
    • rewriting by another LLM (paraphrasing) is the oldest and still the main attack
    • newer attacks train the rewriter against a detector, and they beat detectors that survived plain paraphrasing
  • most web pages are not rewritten to dodge a detector, so these numbers are a worst case
    • but “humanizer” tools are sold to the public, so the worst case is cheap to get
  • for a site-level classifier like yours
    • an attacker has to fool most of the 15-20 pages, not one
    • I think deliberate evasion matters less than ordinary rewriting (grammar tools, translation, editing), see short_mixed_code.md
  • the defenses that work are either specific to one provider (retrieval) or trained against attacks the authors already knew about

paraphrasing

  • takeaway: a paraphraser model removes most of the signal from zero-shot detectors and watermarks, with little change in meaning
  • Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense, Krishna et al., NeurIPS 2023
    • they built DIPPER, an 11B paraphrase model, and ran its output through detectors
    • quote: “DIPPER drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%), without appreciably modifying the input semantics”
    • for web pages
      • a cheap rewrite is enough to make a zero-shot score meaningless
      • the same paper also gives the retrieval defense, see defenses below
  • Can AI-Generated Text be Reliably Detected?, Sadasivan et al., arXiv 2023
    • they paraphrase again and again (recursive paraphrasing) and test watermark, neural, zero-shot and retrieval detectors on passages of about 300 tokens
    • quote: “our recursive paraphrasing method can significantly reduce detection rates, it only slightly degrades text quality in many cases”
    • they also show a watermark can be spoofed: “an attacker can infer hidden AI text signatures without white-box access to the detection method”
    • for web pages
      • watermarks are not a safe fallback, since you don’t control who watermarks anyway
      • they also give a theory result tying best-possible AUROC to how close the human and AI text distributions are, so better LLMs make detection harder by design
  • Language Model Detectors Are Easily Optimized Against, Nicks et al., ICLR 2024
    • they fine-tune a language model with reinforcement learning, using a detector’s “human-ness” score as the reward
    • quote (abstract page): “we advise against continued reliance on LLM-generated text detectors”
    • the abstract reports a 7B Llama-2 whose OpenAI RoBERTa-Large detector AUROC falls from 0.84 to 0.63 with almost no change in perplexity (8.7 to 9.0)
    • for web pages
      • a generator that was tuned against a public detector is invisible to that detector, and also to others it never saw
      • this is the “fine-tuned generator” case, so a site owner who tunes a model once gets it for every page

tuned rewriters and generators, 2025-2026

  • takeaway: attacks now learn from a detector’s feedback and transfer to detectors they never saw
  • Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text, Cheng et al., NeurIPS 2025
    • an ordinary instruction-following LLM paraphrases, and a detector picks among candidates at each step
    • quote: “adversarial paraphrasing, guided by OpenAI-RoBERTa-Large, reduces T@1%F by 64.49% on RADAR and a striking 98.96% on Fast-DetectGPT”
      • T@1%F is the share of AI text caught when 1% of human text is wrongly flagged
    • quote: “Across a diverse set of detectors–including neural network-based, watermark-based, and zero-shot approaches–our attack achieves an average T@1%F reduction of 87.88%”
    • for you: Fast-DetectGPT, which you tried, loses almost all its recall here, and Binoculars sits in the same zero-shot family
      • I did not see Binoculars in the abstract, so I can’t say how much it loses
  • Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors (CoPA), Fang et al., EMNLP 2025
    • no training needed: while decoding, subtract the word probabilities of a “machine-like” version from a “human-like” version
    • quote: “CoPA constructs an auxiliary machine-like word distribution as a contrast to the human-like distribution generated by the LLM. By subtracting the machine-like patterns from the human-like distribution during the decoding process, CoPA is able to produce sentences that are less discernible by text detectors”
    • for web pages: this attack needs only an open model, no detector access
  • TempParaphraser: “Heating Up” Text to Evade AI-Text Detection through Paraphrasing, Huang et al., EMNLP 2025
    • they found that raising the sampling temperature (more random word choice) hurts detectors, then mimic that with several normal-temperature rewrites
    • quote: “increasing the temperature parameter during inference significantly reduces detection accuracy”
    • quote: “TempParaphraser reduces detector accuracy by an average of 82.5% while preserving high text quality”
    • for web pages: sampling settings alone move your scores, so a site that happens to use high temperature looks more human
      • RAID (see benchmarks.md) shows the same thing for repetition penalties
  • Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors, Pedrotti et al., Findings of ACL 2025
    • they fine-tune a model with preference optimization (DPO) so its style moves toward human writing
    • quote: “detectors can be easily fooled with relatively few examples, resulting in a significant drop in detecting performances”
    • they also say detectors lean on “linguistic shortcuts”, meaning style habits instead of anything about authorship
  • AuthorMist: Evading AI Text Detectors with Reinforcement Learning, David et al., arXiv 2025
    • a 3B model is trained with reinforcement learning using real detector APIs (GPTZero, WinstonAI, Originality.ai) as the reward
    • quote: “attack success rates ranging from 78.6% to 96.2% against individual detectors”
    • for web pages: commercial detectors can be attacked through their own paid API, so assume that happens
  • StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors, Ranganath et al., arXiv 2026
    • same idea, trained against an ensemble of four detectors, tested on the MAGE test set
    • quote: “StealthRL achieves near-zero detection on three of the four detectors and a 0.024 mean TPR@1%FPR, reducing mean AUROC from 0.79 to 0.43 and attaining a 97.6% attack success rate”
    • quote: “attacks transfer to two held-out detectors not seen during training, revealing shared architectural vulnerabilities rather than detector-specific brittleness”
    • an AUROC of 0.43 is worse than guessing, so the attacked AI text scores as more human than real human text
  • MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization, Gu et al., Findings of ACL 2026
    • three steps (style fine-tuning, preference optimization, refinement while generating), using only the detector’s answers
    • quote: “MASH achieves an average Attack Success Rate (ASR) of 92%, surpassing the strongest baselines by an average of 24%, while maintaining superior linguistic quality”
    • tested on 6 datasets and 5 detectors
  • Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning (SHADE), Tommasel et al., ELOQUENT lab 2026
    • useful as a warning that evasion is not automatically “more human”
    • quote: “successful evasion is associated with shorter, simpler, and less lexically diverse outputs, suggesting that high detector evasion does not necessarily correspond to more human-like writing”
    • quote: “optimization against a single surrogate detector only partially transfers to unseen evaluation classifiers”
    • for web pages: attacks tuned on one detector are weaker against another, so using more than one kind of detector helps

humanizer tools and prompt tricks

  • takeaway: commercial humanizers exist, are cheap, and fool older detectors, but the one study of them found a detector that holds up
  • DAMAGE: Detecting Adversarially Modified AI Generated Text, Masrour et al. (Pangram Labs), GenAIDetect workshop 2025
    • they studied 19 humanizer and paraphrasing tools, judged how well each kept the meaning, and trained a detector on their output
    • quote: “We study 19 AI humanizer and paraphrasing tools and qualitatively assess their effects and faithfulness in preserving the meaning of the original text. We show that many existing AI detectors fail to detect humanized text.”
    • quote: “We attack our own detector, training our own fine-tuned model optimized against our detector’s predictions, and show that our detector’s cross-humanizer generalization is sufficient to remain robust to this attack”
    • for web pages: the authors sell Pangram, so the robustness claim is theirs, not an independent test
  • Artificial Writing and Automated Detection, Jabarian and Imas, NBER working paper 2025
    • an economics team (not detector sellers) tested Pangram, OriginalityAI, GPTZero and a RoBERTa model on about 2,000 passages in six genres
    • quote: “Pangram achieving near-zero FNR and FPR rates that remain robust across models, threshold rules, ultra-short passages, “stubs” (≀ 50 words) and ‘humanizer’ tools“
      • FNR is the share of AI text called human, FPR the share of human text called AI
    • for web pages: this is the strongest independent result for a commercial detector surviving humanizers
      • it covers one detector and their own corpus, so it does not tell you open detectors survive
  • GenAI Detection Tools, Adversarial Techniques and Implications for Inclusivity in Higher Education, Perkins et al., arXiv 2024
    • they hand-modified AI text with evasion tricks and ran six detectors (805 texts)
    • quote: “the detectors’ already low accuracy rates (39.5%) show major reductions in accuracy (17.4%) when faced with manipulated content”
    • for web pages: this is the plain classroom version of the problem, with off-the-shelf detectors and no fancy attack
  • Large Language Models can be Guided to Evade AI-Generated Text Detection (SICO), Lu et al., arXiv 2023
    • a prompt-only trick: pick example sentences that steer the LLM to write in a way detectors read as human, using 40 human examples
    • quote: “enables GPT-3.5 to successfully evade six detectors, decreasing their AUC by 0.5 on average”
    • for web pages: no extra model is needed, just a better prompt
  • GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023 (arXiv)
    • side finding: asking ChatGPT to rewrite with fancier word choices also fooled the detectors
    • quote: “simple prompting strategies can not only mitigate this bias but also effectively bypass GPT detectors”
    • full-text number: asking for a native-speaker word-choice rewrite cut the average false positive rate “from 61.22% to 11.77%”
    • the same detectors, evaded by one sentence of prompt, are the ones that flag non-native writing, see benchmarks.md

character swaps and homoglyphs

  • takeaway: swapping a few letters for look-alike characters breaks most detectors, and the fix (normalize the text first) is easy but has to be done
  • SilverSpeak: Evading AI-Generated Text Detectors using Homoglyphs, Creo et al., arXiv 2024
    • a homoglyph is a different character that looks the same, such as a Cyrillic “а” for a Latin “a”
    • quote: “homoglyph-based attacks can effectively circumvent state-of-the-art detectors, leading them to classify all texts as either AI-generated or human-written (decreasing the average Matthews Correlation Coefficient from 0.64 to -0.01)”
    • tested on seven detectors including Binoculars, Fast-DetectGPT, DetectGPT and Ghostbuster over five datasets
    • for web pages
      • you scrape raw HTML, so odd Unicode will reach your detector
      • Binoculars is on the list, so your own pipeline is affected
      • normalizing text before scoring is cheap, but I would also log how many characters were non-Latin, since that flags both attacks and genuine multilingual text

translation round trips and obfuscation

  • takeaway: sending text through other languages and back lowers detection, and ordinary machine translation does too
  • ESPERANTO: Evaluating Synthesized Phrases to Enhance Robustness in AI Detection for Text Origination, Ayoobi et al., arXiv 2024
    • they translate AI text through several languages and back to English, then merge the results
    • quote: “the manipulated text retains the original semantics while significantly reducing the true positive rate (TPR) of existing detection methods”
    • tested on nine detectors (six open, three proprietary), with a 720k-text dataset
    • the authors propose a detector trained on such text, with TPR dropping “by only 1.85% after back-translation manipulation” (their own claim)
  • Testing of Detection Tools for AI-Generated Text, Weber-Wulff et al., arXiv 2023
    • they tested 12 public tools plus Turnitin and PlagiarismCheck, including machine-translated and obfuscated texts
    • quote: “the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques significantly worsen the performance of tools”
    • for web pages: the tools were tuned to avoid false accusations, so they miss AI text by default

defenses

  • takeaway: retrieval works if you hold the provider’s output history, adversarial training works against the attacks it saw, and nothing is proven against new ones
  • retrieval of past outputs: Krishna et al. (above)
    • the provider stores what it generated, and a candidate text is compared against that store
    • quote: “we empirically verify our defense using a database of 15M generations from a fine-tuned T5-XXL model and find that it can detect 80% to 97% of paraphrased generations across different settings while only classifying 1% of human-written sequences as AI-generated”
    • for web pages
      • it needs the generating service to keep and share its outputs
      • you cannot use it on pages from unknown or open models
  • RADAR: Robust AI-Text Detection via Adversarial Learning, Hu et al., arXiv 2023
    • a paraphraser and a detector train against each other, so the detector sees stronger and stronger rewrites
    • quote: “RADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in place”
    • tested on 8 LLMs and 4 datasets
    • later attacks still beat it: adversarial paraphrasing above cut RADAR’s recall at 1% FPR by 64.49%
  • Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training (GREATER), Li et al., ACL 2025
    • the attacker (GREATER-A) perturbs the words that matter most for the detector, and the detector (GREATER-D) trains against it
    • quote: “across 10 text perturbation strategies and 6 adversarial attacks show that our GREATER-D reduces the Attack Success Rate (ASR) by 0.67% compared with SOTA defense methods”
    • the abstract does not say whether that is percentage points or relative percent, so the gain may be small
  • OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples, Koike et al., arXiv 2023
    • no retraining: the detector LLM gets adversarial essays as in-context examples
    • quote: “the proposed detector improves the detection performance on the attacker-generated texts by up to +41.3 points F1-score”
    • the same paper’s attacker also “drastically degrades the performance of detectors by up to -57.0 points F1-score”
  • training on a few attack examples (code): Droid, Orel et al., EMNLP 2025
    • quote: “this problem can be easily amended by training on a small amount of adversarial data”
    • see short_mixed_code.md
  • training on humanizer outputs: DAMAGE (above) and TempParaphraser (above, “training on TempParaphraser-augmented data improves detector robustness”)
  • what I take from all of these
    • defenses trained on attack X are tested on X or close cousins
    • the newer attacks (StealthRL, MASH, adversarial paraphrasing) are built to beat those defenses
    • so a detector’s robustness number only counts for attacks the authors did not train on

Last edited: