detectors trained on labeled examples, and rewrite-based detectors (authored by agents unless marked đ§)
terms (perplexity, AUROC, FPR, TPR, F1, rewrite-and-compare) are defined in how_detection_works.md
- âtrainedâ means the detector itself learned from labeled human and machine texts
- quotes are from each paperâs abstract unless I say otherwise
- weak spot: where I name a section or figure, I read it in the PDF, otherwise it is what the abstract leaves out
the shared problem
- a classifier can learn the quirks of its training data instead of the difference between people and models
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts, Tulchinskii et al., NeurIPS 2023, related work section
- supervised detectors âdo not generalize to other text domains, generation models, and even sampling strategiesâ
- so every trained detector below should be read for how it handles new topics, new generators and rewrites
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts, Tulchinskii et al., NeurIPS 2023, related work section
group 1, a plain classifier on the text
- idea: fine-tune a text encoder to say human or machine
- OpenAIâs GPT-2 output detector
- source: openai/gpt-2-output-dataset detector README, OpenAI, 2019
- trick: âthe GPT-2 output detector model, obtained by fine-tuning a RoBERTa model with the outputs of the 1.5B-parameter GPT-2 modelâ
- needs: roberta-base (478 MB) or roberta-large (1.5 GB) weights, one forward pass, runs on a laptop or small GPU
- result: the README makes no accuracy claim, the report is Release Strategies and the Social Impacts of Language Models, Solaiman et al., 2019
- weak spot: trained on GPT-2 text only, Ghostbuster says RoBERTa-style models âcan exhibit catastrophic worst-case performanceâ
- Technical Report on the Pangram AI-Generated Text Classifier, Emi and Spero, arXiv 2024
- trick: a transformer classifier, trained with âhard negative mining with synthetic mirrorsâ: keep finding the human texts it wrongly flags, and add machine twins of them to the training set
- needs: Pangramâs hosted service, the weights are not public, so you pay per call
- result: âoutperforms zero-shot methods such as DetectGPT as well as leading commercial AI detection tools with over 38 times lower error rates on a comprehensive benchmark comprised of 10 text domains ⊠and 8 open- and closed-source large language modelsâ
- weak spot: the benchmark and the numbers are the companyâs own, I know no independent run in these notes
- MAGE: Machine-generated Text Detection in the Wild, Li et al., ACL 2024
- trick: a large test set (many human domains and many LLMs), not a new detector, used to see how classifiers behave on unseen data
- needs: it releases data and code
- result: âthe top-performing detector can identify 86.54% out-of-domain texts generated by a new LLMâ
- weak spot: the same abstract says detection is hard âespecially out-of-distributionâ, âdue to the decreasing linguistic distinctions between the two sourcesâ
- GPT detectors are biased against non-native English writers, Liang et al., Patterns 2023
- not a detector, a test of several, see how_detection_works.md
- Amplifying, Not Learning, arXiv 2026
- not a detector, a claim that fine-tuned detectors mostly rescale how predictable the text is, see how_detection_works.md
group 2, classifiers built on language-model probabilities
- idea: do not read the text, read how a language model finds it, and train a small learner on that
- Ghostbuster: Detecting Text Ghostwritten by Large Language Models, Verma et al., NAACL 2024
- trick: get per-word probabilities from several weak models, search over combinations of them, train a small classifier on the best combinations
- needs: a unigram model, a trigram model, and the early GPT-3 models ada and davinci for probabilities, no access to the target model; those old API models may no longer be callable (my read, not checked)
- result: âGhostbuster achieves 99.0 F1 when evaluated across domains, which is 5.9 F1 higher than the best preexisting modelâ
- weak spot: âGhostbuster may be unreliable for documents with †100 tokens, and its performance levels off with â„ 500 tokensâ
- Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection, Miralles-GonzĂĄlez et al., Information Fusion 2025 (PAWN)
- trick: some words are more telling than others, so learn per-word weights instead of a plain average of surprise
- needs: one language model pass (hidden states and probabilities cached on disk), then training a small head
- result: âPAWN shows competitive and even better performance in-distribution than the strongest baselines (fine-tuned LMs) with a fraction of their trainable parametersâ
- weak spot: âfraction of trainable parametersâ is not a cheaper detection run, you still run the big model on every text
- SV-Detect: AI-generated Text Detection with Steering Vectors, arXiv 2026
- trick: find, at each layer, a direction that separates human from machine text in a frozen model, and train a light classifier on the projections
- needs: one pass of a frozen LLM, plus labeled examples
- result: âstrong performance both in-distribution and under distribution shift, including across domains, source models, and machine-editing transformations such as polishing and rewritingâ
- weak spot: no numbers in the abstract
- Steer-to-Detect: Probing Hidden Representations for Detection of LLM-Generated Texts, arXiv 2026
- trick: learn a vector injected into a frozen modelâs hidden states so the two classes separate, then run a statistical test
- needs: a frozen observer LLM and some labeled examples
- result: âWe establish finite-sample, high-probability guarantees for Type I and Type II errorsâ
- weak spot: the guarantees hold under the paperâs assumptions on how the features are distributed, which a new website need not follow
- MoSEs: Uncertainty-Aware AI-Generated Text Detection via Mixture of Stylistics Experts with Conditional Thresholds, Wu et al., EMNLP 2025
- trick: pick the cutoff per text, by looking up labeled texts in a similar style
- needs: a reference store of labeled texts, plus a scoring model
- result: âOur framework achieves an average improvement 11.34% in detection performance compared to baselinesâ
- weak spot: a style that the reference store does not cover gets a poor cutoff (my read)
group 3, learn from many authors or from attacks
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning, Guo et al., NeurIPS 2024
- trick: train an encoder so texts by the same author or model land near each other, then classify a new text by looking up its nearest labeled neighbors
- needs: a text encoder and a labeled database, no per-text generation
- result: âin OOD zero-shot evaluation, our method outperforms existing approaches by a large marginâ
- weak spot: the abstract gives no number and no FPR
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution, Cava and Tagarelli, EMNLP 2025
- trick: a benchmark built from open models, plus a contrastive detector (OTBDetector)
- needs: the released data and model
- result: âour detector achieving remarkable capabilities across the various tasks and outperforming most existing detectorsâ
- weak spot: open models only, closed models are not covered
- RADAR: Robust AI-Text Detection via Adversarial Learning, Hu et al., NeurIPS 2023
- trick: train a paraphraser to evade the detector and a detector to catch the paraphraser, in alternation
- needs: training both models, the detector is run alone afterward
- result: âRADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in placeâ
- weak spot: tested on 2023 models (Pythia, Dolly, LLaMA, Vicuna and so on), with GPT-3.5-Turbo as the only newer check
- Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial Training, Li et al., ACL 2025 (GREATER)
- trick: an attacker model finds the words that most matter to the detector and swaps them, the detector trains against it
- needs: training an attacker and a detector together
- result: âreduces the Attack Success Rate (ASR) by 0.67% compared with SOTA defense methodsâ
- weak spot: 0.67% is small, and I did not check whether it is relative or in points
group 4, rewrite and compare
- idea: rewrite the text with a model, machine text changes less
- Raidar: geneRative AI Detection viA Rewriting, Mao et al., ICLR 2024
- trick: prompt an LLM to rewrite the text, count the edits, a small edit distance means machine text
- needs: an LLM call per rewrite, black box is fine, only word-level edits are used so no probabilities
- result: âRaidar significantly improves the F1 detection scores of existing AI content detection models â both academic and commercial â across various domains ⊠with gains of up to 29 pointsâ
- weak spot: the Learning to Rewrite paper says a trained-in-domain rewrite model can cause âRAIDAR to fail to generalize to new domainsâ
- Learning to Rewrite: Generalized LLM-Generated Text Detection, Li et al., ACL 2025 (L2R)
- trick: fine-tune the rewriting model so it leaves machine text alone and rewrites human text more, widening the gap
- needs: fine-tuning a rewriter (the paperâs figure shows LLaMA-3-8B), then a rewrite per text
- result: âoutperforms state-of-the-art detection methods by up to 23.04% in AUROC for in-distribution tests, 37.26% for out-of-distribution tests, and 48.66% under adversarial attacksâ
- weak spot: one generation per text, plus training the rewriter
- Learn-to-Distance: Distance Learning for Detecting LLM-Generated Text, Zhou et al., ICLR 2026 (L2D)
- trick: instead of a plain edit distance between the text and its rewrite, learn the distance
- needs: a rewriting LLM, plus training the distance
- result: âit achieves relative improvements from 54.3% to 75.4% over the strongest baseline across different target LLMs (e.g., GPT, Claude, and Gemini)â
- weak spot: still one rewrite per text, and ârelative improvementâ hides the starting point
- MAGRET: Machine-generated Text Detection with Rewritten Texts, Huang et al., COLING 2025
- trick: ask candidate LLMs to rewrite the text in several ways, train a BERT encoder to judge how close the rewrites sit to the original
- needs: access to the candidate LLMs, so cost and coverage grow with the list of models
- result: âprevious methods struggle with closed-source model detection, while our approach significantly outperforms baseline methods in this regardâ
- weak spot: a generator that is not on your candidate list is not covered (my read)
- Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks, arXiv 2026
- trick: score the text plus two transformed versions that keep its content or keep its style, and combine the three scores
- needs: a rewriting step and a base detector, I could not tell from the abstract whether any training is needed
- result: âimproves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAIDâ
- weak spot: designed for attacked text, extra rewrites per text
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models, Fu et al., ACM MM 2025
- trick: train the detector to directly predict the gap between the original and rewritten text, rather than a generic label
- needs: a base scoring model and training, plus the MIRAGE data
- result: âachieving over a 70% performance improvement under the same training data and base scoring modelâ
- weak spot: the 70% is relative to their own baselines
group 5, handle edited and mixed text
- idea: real text is often human text a model polished, or machine text a human edited
- Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection, ACL 2026 (RACE)
- trick: build a discourse graph for the creatorâs structure and sentence-piece features for the editorâs style, classify into four kinds
- needs: a discourse parser and a trained classifier
- result: âRACE outperforms 12 baselines in identifying fine-grained types with low false alarmsâ
- weak spot: no numbers in the abstract
- SenDetEX: Sentence-Level AI-Generated Text Detection for Human-AI Hybrid Content via Style and Context Fusion, EMNLP 2025
- trick: judge one sentence at a time, using the neighbouring sentences as context
- needs: a trained model, and the paperâs synthetic mixed-text dataset
- result: âsignificantly outperforms all baseline models in detection accuracy, while exhibiting remarkable transferability and robustnessâ
- weak spot: the mixed training texts are synthesized, real editing may differ
- Hidden Human-Like Nature of Machine-Generated Texts: Theory and Detection Enhancement, arXiv 2026
- trick: even machine text has stretches that look human, so repeatedly drop the most human-looking pieces and retrain on the rest
- needs: an existing detector, âthe framework can also work in a training-free mannerâ
- result: âExtensive experiments across various LLMs and practical scenarios demonstrate that the proposed framework consistently enhances existing detectorsâ
- weak spot: filtering by the detectorâs own confidence can repeat its mistakes (my read)
- Latent Trajectory Discrimination for AI-Generated Text Detection, arXiv 2026 (GTCL)
- trick: embed overlapping windows of the document in order, and learn from how the embedding moves along the text
- needs: a text embedding model and contrastive training
- result: âGTCL outperforms detection baselines consistentlyâ
- weak spot: tested on long documents, no word on short snippets
in short
- for a web study, the closest to what the human already runs is group 2, plus rewrite methods as a second opinion
- none of these abstracts gives a TPR at 0.01% FPR on a web crawl, so each needs a recheck on your own pages
Last edited: