detectors that need no labeled training data (authored by agents unless marked đ§)
terms (perplexity, rank, AUROC, FPR, TPR) are defined in how_detection_works.md
- âzero-shotâ here means the detector is not trained on labeled human/machine examples
- all of them still need a language model to score the text, and a labeled sample to pick the cutoff
- quotes are from each paperâs abstract unless I say otherwise
- weak spots: where I name a paper section, I read it in the PDF, otherwise the weak spot is what the abstract leaves out
group 1, read the text once and look at how expected it is
- idea: machine text sits where the scoring model puts high probability
- GLTR: Statistical Detection and Visualization of Generated Text, Gehrmann et al., ACL demo 2019
- trick: color every word by its rank under the model, green for top 10, red for rare, a human reads the colors
- needs: one GPT-2 117M pass, the deployed demo also offers BERT
- result: âthe annotation scheme provided by GLTR improves the human detection-rate of fake text from 54% to 72% without any prior trainingâ
- weak spot: built for GPT-2 era text sampled from the head of the distribution, nothing here tests modern chat models
- log-likelihood, rank and log-rank baselines
- trick: average the per-word log probability, or the log of the wordâs rank, and threshold it
- needs: one pass of the scoring model
- these are the baselines inside the DetectGPT and Fast-DetectGPT tables
- on GPT-J outputs the Fast-DetectGPT paper reports LogRank 0.8818 AUROC and Likelihood 0.8480
- weak spot: they fail once the generator is not the scoring model, and flag plain human text
- Training-free LLM-generated Text Detection by Mining Token Probability Sequences, Xu et al., ICLR 2025 (Lastde)
- trick: treat the per-word probabilities as a time series, and measure how they wiggle locally as well as on average
- needs: one scoring model pass, Lastde++ is a faster variant
- result: âour method consistently achieves state-of-the-art performanceâ on six datasets, âgreater robustness against paraphrasing attacksâ
- weak spot: the abstract gives no number at low FPR
- Diversity Boosts AI-Generated Text Detection, Basani and Chen, arXiv 2025 (DivEye)
- trick: human text has more up-and-down in surprise from word to word, so use features of how surprise fluctuates
- needs: one scoring model pass, then a few simple statistics
- result: âoutperforms existing zero-shot detectors by up to 33.2% and achieves competitive performance with fine-tuned baselines across multiple benchmarksâ
- weak spot: âup toâ is the best case, and the abstract does not say what language models it needs
- When AI Settles Down: Late-Stage Stability as a Signature of AI-Generated Text Detection, Sun et al., arXiv 2026
- trick: machine text gets steadier toward the end, so measure variation in the second half only
- needs: one scoring pass, âWithout perturbation sampling or additional model accessâ
- result: âThis divergence peaks in the second half of sequences, where AI-generated text shows 24â32% lower volatilityâ
- weak spot: needs enough text to have a second half, so short pages are out
- AI-Generated Text is Non-Stationary: Detection via Temporal Tomography, arXiv 2025 (TDT)
- trick: turn per-word scores into a signal and take a wavelet transform to see where and at what scale anomalies sit
- needs: scoring pass plus a transform, âonly 13% computational overheadâ
- result: âOn the RAID benchmark, TDT achieves 0.855 AUROC (7.1% improvement over the best baseline)â
- weak spot: AUROC only, nothing about FPR on real pages
- DWT-Fusion: A Signal-Based Framework for Training-Free LLM-Generated Text Detection, arXiv 2026
- trick: same wavelet idea on the log-probability sequence, then vote across several settings
- needs: one proxy model, tested with GPT-Neo-2.7B, GPT-J-6B, Falcon-7B and LLaMA-3-8B
- result: âThe best single wavelet configurations achieve AUROC values of 0.9872, 0.8185, and 0.7138 on HC3, M4, and MAGE, respectivelyâ
- weak spot: that spread is the story, 0.99 on HC3 but 0.71 on MAGE
group 2, compare the text against slightly different versions of itself
- idea: machine text sits at a peak of the modelâs probability, so nearby alternatives are less likely, while human text does not sit at a peak
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature, Mitchell et al., ICML 2023
- trick: rewrite small spans of the text many times, check whether the original is clearly more likely than the rewrites
- needs: the generating modelâs probabilities plus a mask-filling model; the paper uses 100 samples from T5-3B per passage
- result: ânotably improving detection of fake news articles generated by 20B parameter GPT-NeoX from 0.81 AUROC for the strongest zero-shot baseline to 0.95 AUROC for DetectGPTâ
- weak spot: slow, the Fast-DetectGPT paper times the full benchmark at 79,113 seconds, about 22 hours
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text, Su et al., EMNLP Findings 2023 (LRR and NPR)
- trick: LRR divides log-likelihood by log-rank, NPR does the DetectGPT rewrite test on ranks
- needs: LRR one pass, NPR many rewrites
- result: âour proposed methods improve over the state of the art by 3.9 and 1.75 AUROC points absoluteâ
- weak spot: NPR keeps the rewrite cost, which the abstract calls âmore accurate, but slower due to the need for perturbationsâ
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature, Bao et al., ICLR 2024
- trick: skip the rewriting, ask how much more likely the real word is than the words the model itself would have sampled there
- needs: one or two model passes, no generation; in the black-box setting GPT-J samples and GPT-Neo-2.7B scores, timing used a single Tesla A100
- result: ânot only surpasses DetectGPT by a relative around 75% in both the white-box and black-box settings but also accelerates the detection process by a factor of 340â
- the 75% is relative to the AUROC gap left to 1.0, not 75 points
- Table 2: 0.9887 average AUROC with source-model access, 0.9677 black-box
- weak spot: the ChatGPT/GPT-4 AUROC in Table 1 is 0.9338, lower than on open models, and AUROC says nothing about the 0.01% corner
- Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection, Bao et al., ICLR 2025
- trick: closed APIs return only top probabilities, so guess the rest of the distribution, then run Fast-DetectGPT and friends on the guess
- needs: an API that returns top token probabilities for supplied text, for example GPT-3.5
- result: âGlimpse with Fast-DetectGPT and GPT-3.5 achieves an average AUROC of about 0.95 in five latest source modelsâ
- weak spot: depends on an API feature that vendors can remove, and your page text goes to the vendor
- Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy, arXiv 2026 (LAPD)
- trick: chat models are tuned from base models, so the probability ratio between a tuned and a base model carries an imprint of that tuning
- needs: a tuned model and its base model
- result: âLAPD achieves an improvement 45.82% relative to the strongest existing baselinesâ
- weak spot: the guarantee that it beats Fast-DetectGPT is theory under assumptions, the abstract does not report FPR numbers
- Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text Detection, ACL 2026
- trick: some words carry more evidence than others, find them by where two modelsâ internal states disagree, and weight them more
- needs: two models
- result: âit attains a 2.2% relative improvement in average AUROC over the strongest prior baseline on DetectRLâ
- weak spot: a small gain, on one benchmark
- Zero-Shot Detection of LLM-Generated Text using Temperature Sensitivity, Ma et al., ACL 2026 (NTS)
- trick: change the sampling temperature of the scoring model and see how its word probabilities react, machine text reacts more
- needs: one surrogate model, run at several temperatures
- result: âLLM-generated text tends to exhibit higher TS than human-written textâ
- weak spot: I did not open the full paper, so I cannot report its limits
group 3, contrast two models
- idea: a single modelâs surprise depends on how hard the text is, so divide by what another model expects
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text, Hans et al., ICML 2024
- trick: perplexity under one model divided by how surprised that model is by the other modelâs predictions (âcross-perplexityâ), so hard prompts do not fool it
- needs: two models with the same tokenizer, the paper uses Falcon-7B and Falcon-7B-Instruct, two forward passes, a GPU that fits both
- result: âBinoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT dataâ
- weak spot: section 6 says âwe do not consider explicit efforts to bypass detectionâ, it skips source code, and the authors say they lacked GPU memory to test 30B+ models
- MOSAIC: Multiple Observers Spotting AI Content, Dubois et al., arXiv 2024 (revised 2025)
- trick: Binoculars with more than two models, combined by a principled weighting
- needs: several scoring models, so more GPU memory and time
- result: âusing a fixed pair of models can induce brittleness in performanceâ, their ensemble gives ârobust detection performance across multiple domainsâ
- weak spot: the abstract gives no numbers
- Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection, Chen et al., AAAI 2025 (ImBD)
- trick: fine-tune the scoring model briefly on machine text first, then run a Fast-DetectGPT style test, aimed at text a person wrote and a machine polished
- needs: a scoring model and a short tuning run (the abstract says âjust samples and five minutes of SPOâ), so lightly trained, not truly zero-shot
- result: âNotably, our method surpasses the commercially trained GPT-Zero with just samples and five minutes of SPO, demonstrating its efficiency and effectivenessâ
- weak spot: needs some machine-style samples, and results are for revised text, not whole websites
- kNNProxy: Efficient Training-Free Proxy Alignment for Black-Box Zero-Shot LLM-Generated Text Detection, arXiv 2026
- trick: the scoring model rarely matches the unknown generator, so mix its predictions with nearest-neighbor lookups from a small store of text from the target
- needs: the proxy model plus a small datastore built once from target-style text
- result: the abstract only says âExtensive experiments demonstrate strong detection performance of our methodâ
- weak spot: no numbers in the abstract, and you need sample text from the generator you want to catch
group 4, regenerate and compare
- idea: cut the text, let a model continue it, see how much its continuation matches the real continuation
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text, Yang et al., ICLR 2024
- trick: keep the first part, have the model regenerate the rest K times, machine text overlaps its own regenerations in word sequences more than human text
- needs: black-box API calls K times per text (the black-box score is n-gram overlap), or probabilities for the white-box score
- result: âoutperforming OpenAIâs own classifier, which is trained on millions of textâ on âfour English and one German datasetâ
- weak spot: K generations per page cost money, and it targets GPT-family models
- A Training-free Method for LLM Text Attribution, arXiv 2025
- trick: model the text as a random process and run a hypothesis test, designed to hold the false positive rate fixed
- needs: the candidate LLMâs probabilities, or sampling for black-box access
- result: âboth Type I and Type II errors decay exponentially with text lengthâ
- weak spot: it answers âdid this particular model write thisâ, and the abstract admits some model pairs cannot be separated
group 5, look at the geometry of the text
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts, Tulchinskii et al., NeurIPS 2023
- trick: embed the text, estimate how many degrees of freedom its points have, human text is about 1.5 higher
- needs: a text embedding model, no generator access
- result: âthe average intrinsic dimensionality of fluent texts in a natural language is hovering around the value 9 for several alphabet-based languages and around 7 for Chinese, while the average intrinsic dimensionality of AI-generated texts for each language is â 1.5 lowerâ
- weak spot: section 6 lists three limits, the first is âit is stochastic in natureâ, scores of the same generator vary widely and noise adds up
group 6, bring in outside text
- Enhancing LLM Text Detection with Retrieved Contexts and Logits Distribution Consistency, Huang et al., EMNLP 2025 (HALO)
- trick: fetch similar human text and machine-rewritten text, see how the pageâs word probabilities shift under each
- needs: a retrieval corpus and a rewriting model, on top of the scoring model
- result: âachieves state-of-the-art performance in AUROC, both in cross-domain and domain-specific scenariosâ
- weak spot: the retrieval corpus and rewrites are part of the cost and could leak into the test
the ones I would start from
- Binoculars and Fast-DetectGPT remain the best-documented, cheapest ones to run
- anything with âAUROCâ and no low-FPR number needs a recheck on your data before it counts as better
- the 2026 papers above are abstracts I read, not results I reproduced
Last edited: