how detectors are tested, and what goes wrong (authored by agents unless marked đ§)
terms like threshold, false positive rate (FPR), AUROC and perplexity are explained in how_detection_works.md
takeaway
- a benchmark score tells you how a detector does on that benchmarkâs text, and little about your web pages
- older benchmarks use old generators, one domain, clean text and no attacks, and detectors score near 100% on them
- newer ones add many generators, domains, languages, decoding settings and attacks, and the same detectors fall
- three checks to run before trusting any number
- what is the FPR, and at what threshold: average accuracy hides false accusations
- was the threshold picked on the same data it is scored on
- does the test text look like the text you will see (new models, short pages, edited text, non-native writers)
- for your project: pick the threshold on human pages from the sites you care about, not on a benchmark
the benchmarks, oldest first
- takeaway: each one fixed a gap in the one before, and each reported that detectors break when the setting shifts
- TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation, Uchendu et al., arXiv 2021
- 200K human and machine texts from 19 generators (GPT-1 to 3, GROVER, CTRL, XLNet and others), with a leaderboard
- quote: âFAIR_wmt20 and GPT-3 are the current winners, among all language models tested, in generating the most human-like indistinguishable texts with the lowest F1 score by five state-of-the-art TT detection modelsâ
- meaning: all the generators are from 2019-2021, so it says nothing about todayâs models
- MAGE: Machine-generated Text Detection in the Wild, Li et al., arXiv 2023
- texts from many human sources and many LLMs, with test sets for unseen domains and unseen models
- quote: âEmpirical results show challenges in distinguishing machine-generated texts from human-authored ones across various scenarios, especially out-of-distributionâ
- quote: âthe top-performing detector can identify 86.54% out-of-domain texts generated by a new LLMâ
- meaning: the best case on a new model is about 86%, and the abstract gives no FPR
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection, Wang et al., arXiv 2023 (EACL 2024)
- quote: âit is challenging for detectors to generalize well on instances from unseen domains or LLMs. In such cases, detectors tend to misclassify machine-generated text as human-writtenâ
- meaning: a detector shifted to new data tends to miss AI text, so the errors go in the direction you least want if you are counting AI pages
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection, Wang et al., arXiv 2024
- adds model attribution, mixed human+machine text with a boundary word, and a human-performance test
- quote: âobtaining good performance in MGT detection usually requires an access to the training data from the same domain and generatorsâ
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark, Macko et al., arXiv 2023
- 74,081 texts in 11 languages from 8 multilingual LLMs
- quote: âavailable benchmarks which lack authentic texts in languages other than English and predominantly cover older generatorsâ
- meaning: if your crawl includes non-English sites, check that the detectorâs training data does too
- SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection, Wang et al., SemEval 2024
- a shared task with three parts: human vs machine, which model wrote it, and where a text switches from human to machine
- quote: âSubtask C aims to identify the changing point within a text, at which the authorship transitions from human to machineâ
- quote: âFor all subtasks, the best systems used LLMsâ
- meaning: the winners were fine-tuned LLM classifiers, which do well in-domain and often poorly out of it (see the M4 result above)
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Dugan et al., ACL 2024
- 6 million texts, 11 models, 8 domains, 11 attacks, 4 decoding settings, with a public leaderboard
- quote: âMany commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasetsâ
- quote: âcurrent detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative modelsâ
- full-text figure caption: âfew detectors can operate at FPR<1%â, and âBinoculars works significantly better than other detectors at low FPRâ
- meaning: this is the best match for your setup, because Binoculars is the detector that held up at low FPR
- it still loses ground under attacks and under repetition penalties
- DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios, Wu et al., arXiv 2024
- human text from domains where misuse is likely, text from four popular LLMs, and added human-like noise such as word swaps and spelling mistakes
- quote: âeven state-of-the-art (SOTA) detection techniques still underperformed in this taskâ
- full-text findings: âshorter training data is beneficial for building robust detectors, while longer test data improves detector performanceâ
- meaning: short pages are harder, which matters for pages with little text
- GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge, Dugan et al., GenAIDetect 2025
- RAID used as a shared task where all domains and models are seen in training
- quote: âmultiple participants were able to obtain accuracies of over 99% on machine-generated text from RAID while maintaining a 5% False Positive Rateâ
- meaning: a 5% FPR is far too high for accusations, and training on the same domains is easy to pass, so this shows the ceiling and not real-world use
- EvoBench: Towards Real-world LLM-Generated Text Detection Benchmarking for Evolving Large Language Models, Yu et al., Findings of ACL 2025
- 7 model families and 29 versions (updates over time, fine-tuned, pruned), 14 detectors
- quote: âRelying on existing static benchmarks could create a misleading sense of security, overestimating the real-world effectiveness of detection methodsâ
- quote: âthey all struggle to maintain generalization when confronted with evolving LLMsâ
- meaning: a detector tuned today ages as models change, so record which model versions you tested and re-check
- DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection, Wu et al., ACL 2026
- 8 languages, 6 domains, 4 commercial LLMs, plus polishing, expanding and condensing and attack variants
- quote: âinclude typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patternsâ
- meaning: this is the closest benchmark to âAI-assisted but not fully AIâ web text, in many languages
- Droid, AICD Bench and other code benchmarks are in short_mixed_code.md
- MultiSocial, DetectAIRev, the peer-review benchmark and other short-text benchmarks are in short_mixed_code.md
what goes wrong in evaluation
- takeaway: most of the headline numbers come from tests that are easier than real use, and the errors fall on specific groups of people
- average accuracy hides false positives
- Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection, Ayoobi et al., arXiv 2025
- quote: âCurrent approaches predominantly report conventional metrics like AUROC, overlooking that even modest false positive rates constitute a critical impediment to practical deployment of detection systemsâ
- quote: âreal-world deployment necessitates predetermined threshold configuration, making detector stability ⊠a critical factorâ
- the ââŠâ replaces the paperâs own gloss in brackets
- their SHIELD benchmark scores reliability and stability together and includes a âhumanificationâ step with a hardness setting
- meaning: AUROC does not tell you where a fixed threshold lands on a new domain, and you have to pick one
- the RAID text and Jabarian and Imas (see attacks_and_paraphrase.md) both frame results at fixed low FPR, which is the right habit
- zero false positives in 300 human texts only shows the FPR is below about 1% (rule of three: 3 divided by 300), so you need thousands of human texts to claim 0.1%
- Beyond Easy Wins: A Text Hardness-Aware Benchmark for LLM-generated Text Detection, Ayoobi et al., arXiv 2025
- bias against non-native English writers
- GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023 (arXiv)
- they ran seven GPT detectors on 91 TOEFL essays (written by non-native speakers) and 88 US 8th-grade essays
- quote: âthese detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identifiedâ
- full-text numbers: âaverage false positive rate: 61.22%â for TOEFL essays, and 97.80% were flagged by at least one detector
- quote: âGPT detectors may unintentionally penalize writers with constrained linguistic expressionsâ
- meaning: low perplexity means âAI-likeâ to these detectors, and plain simple writing has low perplexity
- a site written by non-native writers could be wrongly flagged as AI
- these were 2023 detectors, so I would test Binoculars on non-native text before assuming the problem went away
- GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023 (arXiv)
- old models in the test set, new models on the web
- Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing, Saha et al., arXiv 2025 (Findings of ACL 2025)
- quote: âdetectors frequently flag even minimally polished text as AI-generated, struggle to differentiate between degrees of AI involvement, and exhibit biases against older and smaller modelsâ
- meaning: the detectors treat polish by an older or smaller model as more AI-like than polish by a newer model, so a test of old models flatters them (details in short_mixed_code.md)
- TuringBench, MULTITuDE and others above say the same: âpredominantly cover older generatorsâ
- Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing, Saha et al., arXiv 2025 (Findings of ACL 2025)
- test data that differs from the wild, or leaks
- Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy, Pudasaini et al., arXiv 2026
- they trained detectors on two benchmark corpora (PAN CLEF 2025, COLING 2025) and then tested across domains and generators
- quote: âclassifiers that excel in-domain degrade significantly under distribution shiftâ
- quote: âdetectors often rely on dataset-specific stylistic cues rather than stable signals of machine authorshipâ
- meaning: a detector may be learning âhow this datasetâs AI text was madeâ such as formatting, not âAIâ
- Benchmarking of LLM Detection: Comparing Two Competing Approaches, Pröhl et al., arXiv 2024
- quote: âthe construction and independence of the evaluation dataset is often not comprehensible. As a result, discrepancies in the performance evaluation of LLM detectors are often visible due to the different benchmarking datasetsâ
- meaning: two papers can report different scores for the same detector because their test sets differ
- I did not find a paper that proves a specific detectorâs test set was in its training data, so I leave âleaksâ as a risk rather than a result
- the closest items are Pudasaini (dataset-specific cues) and EvoBench (new model versions)
- Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy, Pudasaini et al., arXiv 2026
- commercial detectors tested by outsiders
- Artificial Writing and Automated Detection, Jabarian and Imas, NBER 2025
- quote: âPangram is the only tool to satisfy a strict cap (FPR †0.005) without sacrificing accuracyâ
- meaning: other commercial tools failed that cap on their data, which tells you how much threshold choice matters
- Testing of Detection Tools for AI-Generated Text, Weber-Wulff et al., arXiv 2023
- quote: âthe available detection tools are neither accurate nor reliableâ
- meaning: that was ChatGPT-era (2023) text and tools, so it is old evidence
- Pangram 4 Technical Report, Glickenhaus et al. (Pangram Labs), arXiv 2026
- quote: âWe achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%â
- meaning: this is the vendorâs own number on its own test, so treat it as a claim and not an independent test
- Artificial Writing and Automated Detection, Jabarian and Imas, NBER 2025
Last edited: