Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

how detectors are tested, and what goes wrong (authored by agents unless marked 🧑)

terms like threshold, false positive rate (FPR), AUROC and perplexity are explained in how_detection_works.md

takeaway

  • a benchmark score tells you how a detector does on that benchmark’s text, and little about your web pages
    • older benchmarks use old generators, one domain, clean text and no attacks, and detectors score near 100% on them
    • newer ones add many generators, domains, languages, decoding settings and attacks, and the same detectors fall
  • three checks to run before trusting any number
    • what is the FPR, and at what threshold: average accuracy hides false accusations
    • was the threshold picked on the same data it is scored on
    • does the test text look like the text you will see (new models, short pages, edited text, non-native writers)
  • for your project: pick the threshold on human pages from the sites you care about, not on a benchmark

the benchmarks, oldest first

  • takeaway: each one fixed a gap in the one before, and each reported that detectors break when the setting shifts
  • TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation, Uchendu et al., arXiv 2021
    • 200K human and machine texts from 19 generators (GPT-1 to 3, GROVER, CTRL, XLNet and others), with a leaderboard
    • quote: “FAIR_wmt20 and GPT-3 are the current winners, among all language models tested, in generating the most human-like indistinguishable texts with the lowest F1 score by five state-of-the-art TT detection models”
    • meaning: all the generators are from 2019-2021, so it says nothing about today’s models
  • MAGE: Machine-generated Text Detection in the Wild, Li et al., arXiv 2023
    • texts from many human sources and many LLMs, with test sets for unseen domains and unseen models
    • quote: “Empirical results show challenges in distinguishing machine-generated texts from human-authored ones across various scenarios, especially out-of-distribution”
    • quote: “the top-performing detector can identify 86.54% out-of-domain texts generated by a new LLM”
    • meaning: the best case on a new model is about 86%, and the abstract gives no FPR
  • M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection, Wang et al., arXiv 2023 (EACL 2024)
    • quote: “it is challenging for detectors to generalize well on instances from unseen domains or LLMs. In such cases, detectors tend to misclassify machine-generated text as human-written”
    • meaning: a detector shifted to new data tends to miss AI text, so the errors go in the direction you least want if you are counting AI pages
  • M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection, Wang et al., arXiv 2024
    • adds model attribution, mixed human+machine text with a boundary word, and a human-performance test
    • quote: “obtaining good performance in MGT detection usually requires an access to the training data from the same domain and generators”
  • MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark, Macko et al., arXiv 2023
    • 74,081 texts in 11 languages from 8 multilingual LLMs
    • quote: “available benchmarks which lack authentic texts in languages other than English and predominantly cover older generators”
    • meaning: if your crawl includes non-English sites, check that the detector’s training data does too
  • SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection, Wang et al., SemEval 2024
    • a shared task with three parts: human vs machine, which model wrote it, and where a text switches from human to machine
    • quote: “Subtask C aims to identify the changing point within a text, at which the authorship transitions from human to machine”
    • quote: “For all subtasks, the best systems used LLMs”
    • meaning: the winners were fine-tuned LLM classifiers, which do well in-domain and often poorly out of it (see the M4 result above)
  • RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Dugan et al., ACL 2024
    • 6 million texts, 11 models, 8 domains, 11 attacks, 4 decoding settings, with a public leaderboard
    • quote: “Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets”
    • quote: “current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models”
    • full-text figure caption: “few detectors can operate at FPR<1%”, and “Binoculars works significantly better than other detectors at low FPR”
    • meaning: this is the best match for your setup, because Binoculars is the detector that held up at low FPR
      • it still loses ground under attacks and under repetition penalties
  • DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios, Wu et al., arXiv 2024
    • human text from domains where misuse is likely, text from four popular LLMs, and added human-like noise such as word swaps and spelling mistakes
    • quote: “even state-of-the-art (SOTA) detection techniques still underperformed in this task”
    • full-text findings: “shorter training data is beneficial for building robust detectors, while longer test data improves detector performance”
    • meaning: short pages are harder, which matters for pages with little text
  • GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge, Dugan et al., GenAIDetect 2025
    • RAID used as a shared task where all domains and models are seen in training
    • quote: “multiple participants were able to obtain accuracies of over 99% on machine-generated text from RAID while maintaining a 5% False Positive Rate”
    • meaning: a 5% FPR is far too high for accusations, and training on the same domains is easy to pass, so this shows the ceiling and not real-world use
  • EvoBench: Towards Real-world LLM-Generated Text Detection Benchmarking for Evolving Large Language Models, Yu et al., Findings of ACL 2025
    • 7 model families and 29 versions (updates over time, fine-tuned, pruned), 14 detectors
    • quote: “Relying on existing static benchmarks could create a misleading sense of security, overestimating the real-world effectiveness of detection methods”
    • quote: “they all struggle to maintain generalization when confronted with evolving LLMs”
    • meaning: a detector tuned today ages as models change, so record which model versions you tested and re-check
  • DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection, Wu et al., ACL 2026
    • 8 languages, 6 domains, 4 commercial LLMs, plus polishing, expanding and condensing and attack variants
    • quote: “include typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patterns”
    • meaning: this is the closest benchmark to “AI-assisted but not fully AI” web text, in many languages
  • Droid, AICD Bench and other code benchmarks are in short_mixed_code.md
  • MultiSocial, DetectAIRev, the peer-review benchmark and other short-text benchmarks are in short_mixed_code.md

what goes wrong in evaluation

  • takeaway: most of the headline numbers come from tests that are easier than real use, and the errors fall on specific groups of people
  • average accuracy hides false positives
    • Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection, Ayoobi et al., arXiv 2025
      • quote: “Current approaches predominantly report conventional metrics like AUROC, overlooking that even modest false positive rates constitute a critical impediment to practical deployment of detection systems”
      • quote: “real-world deployment necessitates predetermined threshold configuration, making detector stability 
 a critical factor”
        • the “
” replaces the paper’s own gloss in brackets
      • their SHIELD benchmark scores reliability and stability together and includes a “humanification” step with a hardness setting
      • meaning: AUROC does not tell you where a fixed threshold lands on a new domain, and you have to pick one
    • the RAID text and Jabarian and Imas (see attacks_and_paraphrase.md) both frame results at fixed low FPR, which is the right habit
      • zero false positives in 300 human texts only shows the FPR is below about 1% (rule of three: 3 divided by 300), so you need thousands of human texts to claim 0.1%
  • bias against non-native English writers
    • GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023 (arXiv)
      • they ran seven GPT detectors on 91 TOEFL essays (written by non-native speakers) and 88 US 8th-grade essays
      • quote: “these detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified”
      • full-text numbers: “average false positive rate: 61.22%” for TOEFL essays, and 97.80% were flagged by at least one detector
      • quote: “GPT detectors may unintentionally penalize writers with constrained linguistic expressions”
      • meaning: low perplexity means “AI-like” to these detectors, and plain simple writing has low perplexity
        • a site written by non-native writers could be wrongly flagged as AI
        • these were 2023 detectors, so I would test Binoculars on non-native text before assuming the problem went away
  • old models in the test set, new models on the web
    • Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing, Saha et al., arXiv 2025 (Findings of ACL 2025)
      • quote: “detectors frequently flag even minimally polished text as AI-generated, struggle to differentiate between degrees of AI involvement, and exhibit biases against older and smaller models”
      • meaning: the detectors treat polish by an older or smaller model as more AI-like than polish by a newer model, so a test of old models flatters them (details in short_mixed_code.md)
    • TuringBench, MULTITuDE and others above say the same: “predominantly cover older generators”
  • test data that differs from the wild, or leaks
    • Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy, Pudasaini et al., arXiv 2026
      • they trained detectors on two benchmark corpora (PAN CLEF 2025, COLING 2025) and then tested across domains and generators
      • quote: “classifiers that excel in-domain degrade significantly under distribution shift”
      • quote: “detectors often rely on dataset-specific stylistic cues rather than stable signals of machine authorship”
      • meaning: a detector may be learning “how this dataset’s AI text was made” such as formatting, not “AI”
    • Benchmarking of LLM Detection: Comparing Two Competing Approaches, Pröhl et al., arXiv 2024
      • quote: “the construction and independence of the evaluation dataset is often not comprehensible. As a result, discrepancies in the performance evaluation of LLM detectors are often visible due to the different benchmarking datasets”
      • meaning: two papers can report different scores for the same detector because their test sets differ
    • I did not find a paper that proves a specific detector’s test set was in its training data, so I leave “leaks” as a risk rather than a result
      • the closest items are Pudasaini (dataset-specific cues) and EvoBench (new model versions)
  • commercial detectors tested by outsiders
    • Artificial Writing and Automated Detection, Jabarian and Imas, NBER 2025
      • quote: “Pangram is the only tool to satisfy a strict cap (FPR ≀ 0.005) without sacrificing accuracy”
      • meaning: other commercial tools failed that cap on their data, which tells you how much threshold choice matters
    • Testing of Detection Tools for AI-Generated Text, Weber-Wulff et al., arXiv 2023
      • quote: “the available detection tools are neither accurate nor reliable”
      • meaning: that was ChatGPT-era (2023) text and tools, so it is old evidence
    • Pangram 4 Technical Report, Glickenhaus et al. (Pangram Labs), arXiv 2026
      • quote: “We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%”
      • meaning: this is the vendor’s own number on its own test, so treat it as a claim and not an independent test

Last edited: