Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

detectors trained on labeled examples, and rewrite-based detectors (authored by agents unless marked 🧑)

terms (perplexity, AUROC, FPR, TPR, F1, rewrite-and-compare) are defined in how_detection_works.md

  • “trained” means the detector itself learned from labeled human and machine texts
  • quotes are from each paper’s abstract unless I say otherwise
  • weak spot: where I name a section or figure, I read it in the PDF, otherwise it is what the abstract leaves out

the shared problem

  • a classifier can learn the quirks of its training data instead of the difference between people and models

group 1, a plain classifier on the text

  • idea: fine-tune a text encoder to say human or machine
  • OpenAI’s GPT-2 output detector
    • source: openai/gpt-2-output-dataset detector README, OpenAI, 2019
    • trick: “the GPT-2 output detector model, obtained by fine-tuning a RoBERTa model with the outputs of the 1.5B-parameter GPT-2 model”
    • needs: roberta-base (478 MB) or roberta-large (1.5 GB) weights, one forward pass, runs on a laptop or small GPU
    • result: the README makes no accuracy claim, the report is Release Strategies and the Social Impacts of Language Models, Solaiman et al., 2019
    • weak spot: trained on GPT-2 text only, Ghostbuster says RoBERTa-style models “can exhibit catastrophic worst-case performance”
  • Technical Report on the Pangram AI-Generated Text Classifier, Emi and Spero, arXiv 2024
    • trick: a transformer classifier, trained with “hard negative mining with synthetic mirrors”: keep finding the human texts it wrongly flags, and add machine twins of them to the training set
    • needs: Pangram’s hosted service, the weights are not public, so you pay per call
    • result: “outperforms zero-shot methods such as DetectGPT as well as leading commercial AI detection tools with over 38 times lower error rates on a comprehensive benchmark comprised of 10 text domains 
 and 8 open- and closed-source large language models”
    • weak spot: the benchmark and the numbers are the company’s own, I know no independent run in these notes
  • MAGE: Machine-generated Text Detection in the Wild, Li et al., ACL 2024
    • trick: a large test set (many human domains and many LLMs), not a new detector, used to see how classifiers behave on unseen data
    • needs: it releases data and code
    • result: “the top-performing detector can identify 86.54% out-of-domain texts generated by a new LLM”
    • weak spot: the same abstract says detection is hard “especially out-of-distribution”, “due to the decreasing linguistic distinctions between the two sources”
  • GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023
  • Amplifying, Not Learning, arXiv 2026
    • not a detector, a claim that fine-tuned detectors mostly rescale how predictable the text is, see how_detection_works.md

group 2, classifiers built on language-model probabilities

  • idea: do not read the text, read how a language model finds it, and train a small learner on that
  • Ghostbuster: Detecting Text Ghostwritten by Large Language Models, Verma et al., NAACL 2024
    • trick: get per-word probabilities from several weak models, search over combinations of them, train a small classifier on the best combinations
    • needs: a unigram model, a trigram model, and the early GPT-3 models ada and davinci for probabilities, no access to the target model; those old API models may no longer be callable (my read, not checked)
    • result: “Ghostbuster achieves 99.0 F1 when evaluated across domains, which is 5.9 F1 higher than the best preexisting model”
    • weak spot: “Ghostbuster may be unreliable for documents with ≀ 100 tokens, and its performance levels off with ≄ 500 tokens”
  • Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection, Miralles-GonzĂĄlez et al., Information Fusion 2025 (PAWN)
    • trick: some words are more telling than others, so learn per-word weights instead of a plain average of surprise
    • needs: one language model pass (hidden states and probabilities cached on disk), then training a small head
    • result: “PAWN shows competitive and even better performance in-distribution than the strongest baselines (fine-tuned LMs) with a fraction of their trainable parameters”
    • weak spot: “fraction of trainable parameters” is not a cheaper detection run, you still run the big model on every text
  • SV-Detect: AI-generated Text Detection with Steering Vectors, arXiv 2026
    • trick: find, at each layer, a direction that separates human from machine text in a frozen model, and train a light classifier on the projections
    • needs: one pass of a frozen LLM, plus labeled examples
    • result: “strong performance both in-distribution and under distribution shift, including across domains, source models, and machine-editing transformations such as polishing and rewriting”
    • weak spot: no numbers in the abstract
  • Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts, arXiv 2026
    • trick: learn a vector injected into a frozen model’s hidden states so the two classes separate, then run a statistical test
    • needs: a frozen observer LLM and some labeled examples
    • result: “We establish finite-sample, high-probability guarantees for Type I and Type II errors”
    • weak spot: the guarantees hold under the paper’s assumptions on how the features are distributed, which a new website need not follow
  • MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds, Wu et al., EMNLP 2025
    • trick: pick the cutoff per text, by looking up labeled texts in a similar style
    • needs: a reference store of labeled texts, plus a scoring model
    • result: “Our framework achieves an average improvement 11.34% in detection performance compared to baselines”
    • weak spot: a style that the reference store does not cover gets a poor cutoff (my read)

group 3, learn from many authors or from attacks

  • DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning, Guo et al., NeurIPS 2024
    • trick: train an encoder so texts by the same author or model land near each other, then classify a new text by looking up its nearest labeled neighbors
    • needs: a text encoder and a labeled database, no per-text generation
    • result: “in OOD zero-shot evaluation, our method outperforms existing approaches by a large margin”
    • weak spot: the abstract gives no number and no FPR
  • OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution, Cava and Tagarelli, EMNLP 2025
    • trick: a benchmark built from open models, plus a contrastive detector (OTBDetector)
    • needs: the released data and model
    • result: “our detector achieving remarkable capabilities across the various tasks and outperforming most existing detectors”
    • weak spot: open models only, closed models are not covered
  • RADAR: Robust AI-Text Detection via Adversarial Learning, Hu et al., NeurIPS 2023
    • trick: train a paraphraser to evade the detector and a detector to catch the paraphraser, in alternation
    • needs: training both models, the detector is run alone afterward
    • result: “RADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in place”
    • weak spot: tested on 2023 models (Pythia, Dolly, LLaMA, Vicuna and so on), with GPT-3.5-Turbo as the only newer check
  • Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training, Li et al., ACL 2025 (GREATER)
    • trick: an attacker model finds the words that most matter to the detector and swaps them, the detector trains against it
    • needs: training an attacker and a detector together
    • result: “reduces the Attack Success Rate (ASR) by 0.67% compared with SOTA defense methods”
    • weak spot: 0.67% is small, and I did not check whether it is relative or in points

group 4, rewrite and compare

  • idea: rewrite the text with a model, machine text changes less
  • Raidar: geneRative AI Detection viA Rewriting, Mao et al., ICLR 2024
    • trick: prompt an LLM to rewrite the text, count the edits, a small edit distance means machine text
    • needs: an LLM call per rewrite, black box is fine, only word-level edits are used so no probabilities
    • result: “Raidar significantly improves the F1 detection scores of existing AI content detection models – both academic and commercial – across various domains 
 with gains of up to 29 points”
    • weak spot: the Learning to Rewrite paper says a trained-in-domain rewrite model can cause “RAIDAR to fail to generalize to new domains”
  • Learning to Rewrite: Generalized LLM-Generated Text Detection, Li et al., ACL 2025 (L2R)
    • trick: fine-tune the rewriting model so it leaves machine text alone and rewrites human text more, widening the gap
    • needs: fine-tuning a rewriter (the paper’s figure shows LLaMA-3-8B), then a rewrite per text
    • result: “outperforms state-of-the-art detection methods by up to 23.04% in AUROC for in-distribution tests, 37.26% for out-of-distribution tests, and 48.66% under adversarial attacks”
    • weak spot: one generation per text, plus training the rewriter
  • Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text, Zhou et al., ICLR 2026 (L2D)
    • trick: instead of a plain edit distance between the text and its rewrite, learn the distance
    • needs: a rewriting LLM, plus training the distance
    • result: “it achieves relative improvements from 54.3% to 75.4% over the strongest baseline across different target LLMs (e.g., GPT, Claude, and Gemini)”
    • weak spot: still one rewrite per text, and “relative improvement” hides the starting point
  • MAGRET: Machine-generated Text Detection with Rewritten Texts, Huang et al., COLING 2025
    • trick: ask candidate LLMs to rewrite the text in several ways, train a BERT encoder to judge how close the rewrites sit to the original
    • needs: access to the candidate LLMs, so cost and coverage grow with the list of models
    • result: “previous methods struggle with closed-source model detection, while our approach significantly outperforms baseline methods in this regard”
    • weak spot: a generator that is not on your candidate list is not covered (my read)
  • Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks, arXiv 2026
    • trick: score the text plus two transformed versions that keep its content or keep its style, and combine the three scores
    • needs: a rewriting step and a base detector, I could not tell from the abstract whether any training is needed
    • result: “improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID”
    • weak spot: designed for attacked text, extra rewrites per text
  • DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models, Fu et al., ACM MM 2025
    • trick: train the detector to directly predict the gap between the original and rewritten text, rather than a generic label
    • needs: a base scoring model and training, plus the MIRAGE data
    • result: “achieving over a 70% performance improvement under the same training data and base scoring model”
    • weak spot: the 70% is relative to their own baselines

group 5, handle edited and mixed text

  • idea: real text is often human text a model polished, or machine text a human edited
  • Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection, ACL 2026 (RACE)
    • trick: build a discourse graph for the creator’s structure and sentence-piece features for the editor’s style, classify into four kinds
    • needs: a discourse parser and a trained classifier
    • result: “RACE outperforms 12 baselines in identifying fine-grained types with low false alarms”
    • weak spot: no numbers in the abstract
  • SenDetEX: Sentence-Level AI-Generated Text Detection for Human-AI Hybrid Content via Style and Context Fusion, EMNLP 2025
    • trick: judge one sentence at a time, using the neighbouring sentences as context
    • needs: a trained model, and the paper’s synthetic mixed-text dataset
    • result: “significantly outperforms all baseline models in detection accuracy, while exhibiting remarkable transferability and robustness”
    • weak spot: the mixed training texts are synthesized, real editing may differ
  • Hidden Human-Like Nature of Machine-Generated Texts: Theory and Detection Enhancement, arXiv 2026
    • trick: even machine text has stretches that look human, so repeatedly drop the most human-looking pieces and retrain on the rest
    • needs: an existing detector, “the framework can also work in a training-free manner”
    • result: “Extensive experiments across various LLMs and practical scenarios demonstrate that the proposed framework consistently enhances existing detectors”
    • weak spot: filtering by the detector’s own confidence can repeat its mistakes (my read)
  • Latent Trajectory Discrimination for AI-Generated Text Detection, arXiv 2026 (GTCL)
    • trick: embed overlapping windows of the document in order, and learn from how the embedding moves along the text
    • needs: a text embedding model and contrastive training
    • result: “GTCL outperforms detection baselines consistently”
    • weak spot: tested on long documents, no word on short snippets

in short

  • for a web study, the closest to what the human already runs is group 2, plus rewrite methods as a second opinion
  • none of these abstracts gives a TPR at 0.01% FPR on a web crawl, so each needs a recheck on your own pages

Last edited: