Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

commercial LLM text detectors (authored by agents unless marked 🧑)

the plain picture

  • a commercial detector is a text classifier behind a website and an API
    • you paste text, it returns a score
    • the score is “how AI is this”, sometimes split per sentence
  • the old story was perplexity and burstiness
    • perplexity: how predictable each word is to a language model
    • burstiness: how much that predictability swings across sentences
    • GPTZero started there (2023), Turnitin’s own white paper still names both words
  • the current story is a trained neural classifier
    • take a language-model backbone, fine-tune it on lots of human text and lots of AI text
    • no hand-made features needed, the network finds the tells itself
    • Pangram, GPTZero, Turnitin and (as far as vendors say) Originality and Copyleaks work this way
    • architecture, data and thresholds are almost always secret
  • training data is a major part of the method
    • Pangram: mirror each human text with an AI text on the same topic and length, then keep adding the human texts the model gets wrong
    • GPTZero: huge pool, user dispute button, many fake “humanizer” rewrites added to training
  • every vendor says false positives matter most, then reports a tiny false positive rate
    • these numbers come from the vendor’s own test sets
    • when two vendors test each other, each one wins its own test (see the cross-claims section)
  • independent tests exist, but only a few
    • they mostly agree that the trained commercial tools beat free open-source tools by a lot
    • they disagree on who is best among the commercial ones
    • Pangram looks best in the independent tests I could open, mainly Jabarian and Imas

main takeaways

  • vendor numbers are not comparable
    • different test sets, different thresholds, different definitions of “AI”
    • Pangram 4 counts “mixed” as an error on both sides
    • Originality’s new model asks “how much AI” instead of “AI or not”
  • Jabarian and Imas provide a detailed independent head-to-head
    • Pangram had about zero errors on normal-length text
    • Originality and GPTZero were fine on long text, worse on short text and on humanizers
    • the open-source RoBERTa flagged most human text as AI
  • short text is the weak spot for everyone
    • Pangram needs 50 words, Copyleaks needs 350 characters, Originality refuses some short passages
  • humanizer tools (rewriters sold to beat detectors) split the field
    • in Jabarian and Imas, GPTZero missed about half of humanized text, Pangram almost none
    • in GPTZero’s own paper, the numbers flip: GPTZero 93.5 percent, Pangram 49.7 percent
  • old detectors were biased against non-native English writers
    • Liang et al.: average false positive rate 61.22 percent on TOEFL essays
    • GPTZero says 1.1 percent now, Pangram says 0 percent on the same essays, both are vendor claims
  • a score is not a proof
    • Turnitin itself says its false positive rate “is not zero”
    • Pangram 4 says its own limit: it does not handle people who write like an LLM
  • for DeGenTWeb (whole websites)
    • Pangram 4 cuts long text into 512-token windows and merges them, GPTZero also windows long text
    • Pangram’s 50-word minimum is fine for article pages
    • price matters at scale: Pangram API is 0.05 per 100 words per The Decoder's report of the launch reading guide for vendor claims - "accuracy" means nothing without the false positive rate and the threshold - RAID shows that a detector can look perfect at its default threshold and fail at 5 percent false positive rate - see the RAID section - a vendor number tested on a vendor-picked set is a vendor claim - I mark each one below as vendor claim or independent Pangram what goes in - text, at least 50 words - Pangram 4 Technical Report, section 3.1: "we require that text be at least 50 words long to be considered for analysis" - they define the target narrowly - same report, section 3.1: "Pangram defines AI-generated text as original, substantial natural-language prose generated by an LLM in response to an open-ended writing task or question" - short factual answers, math, code are out of scope - long documents are split into overlapping windows - same report, section 4.1: windows of at most 512 tokens with a stride of 256, predictions merged what the model is - 2024 version (then called Checkfor.ai) - "a slightly modified transformer-style architecture", from the original technical report, section 2.1 - how it works page, pangram.com/research/how-it-works: "approximately 1 million documents comprised of public and licensed human-written text" plus AI text from "GPT-4 and other frontier language models" - that page also says the output is a 0 or 1 prediction, which differs from later versions - Pangram 4 (July 2026) - built on "a popular open-weight MoE model" with the name hidden (Pangram 4 Technical Report, section 4.1) - the Pangram blog says "Pangram 4 is our largest model yet, with 6x as many parameters as Pangram 3.3" (pangram.com/blog/pangram-4-technical) - it is fine-tuned with LoRA adapters, in two stages - one head gives a 15-bucket "fraction of AI" for each 512-token segment - another head gives a label per token: human, ai-assisted, ai-generated - to let every token see the whole window, they feed the window twice (they call it Repeat2) - a separate humanizer probe decides if the text looks deliberately disguised - the humanizer head does not change the backbone, on purpose: "we opt to train the humanizer head as a probe to prevent the adverse impact of false positives" what data it was trained on - human text: licensed or owned - "Pangram 4 is not trained on user-submitted data, customer data from API users, or data obtained by unauthorized Internet crawls" (Pangram 4 report, section 2.1) - rough mix in the report: creative writing 22.2 percent, scientific and medical 18.0, reference and educational 15.9, consumer reviews 10.6, social Q&A and chat 10.4, general web 8.2, news 6.6, essays 4.9, professional 3.1 - AI text: made in-house, using "75 of the most popular models" (Pangram blog, pangram-4-technical) - the 2024 report used about 28 million human documents, all dated 2021 or earlier - table 4 of the original report: reviews 15,000,000, books 7,000,000, scientific papers 3,000,000 the two training tricks - synthetic mirrors - original report, section 4.2: "For each human example, we generate an AI-generated example that matches the original document on as many axes as possible" - Pangram 4 report, section 2.2: step 1 ask an LLM "What is the topic of this article?", step 2 ask it "Write an article about X" - why: "We want the model to learn 'How was this text written?' as opposed to 'What is this text about?'" - mirror prompts have to be hand-tuned to remove tells like "Sure, here is an essay" - hard negative mining - run the model over a huge pool of human text, find the human texts it wrongly calls AI, add them (plus their mirrors) and retrain - why it is needed: training "reaches convergence before the first epoch concludes" because most examples are easy - original report table 5: overall false positive rate on held-out human text fell from 2.29 percent to 0.25 percent after mining - the same table: Books 0.85 to 0.001 percent, Wikipedia 5.34 to 0.23 percent, email (never in training) 6.60 to 0.80 percent - the claim "100x-1000x" reduction comes from that table's caption what comes out - a fraction of AI-written text, plus a label of human, ai-assisted or ai-generated - Pangram 4 report, section 5.1: Human if f_human is at least 0.90, AI if f_AI is at least 0.80, otherwise Mixed - internal predictions are per token; displayed highlights cover sentence groups - Pangram 4 model card, postprocessing: “Product output is constrained to sentence-level resolution” - adjacent labels are merged to a minimum of around two sentences - a document-level flag for "humanized" - AI-assisted label is for text where the human and the AI are tangled, for example "a human who writes a paragraph and then goes back and forth with an AI to edit it" (report, section 3.2) - EditLens is the earlier paper behind the edit-level idea - EditLens: Quantifying the Extent of AI Editing in Text, Thai et al., ICLR, 2026 - abstract as I read it on arXiv: a regression model that predicts "the amount of AI editing present within a text", F1 94.7 percent binary and 90.4 percent ternary, with Grammarly edits as a case study - the label for how much AI was used is built by comparing a human text with its AI-edited version clause by clause - same lines in both: human. same idea, new words: AI-assisted. no match: AI-generated - the AI fraction is AI-generated characters plus half of AI-assisted characters, over all characters vendor claims on accuracy (all Pangram's own tests) - Pangram 4 report abstract: "an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%" - about 1 false positive in 24,300 human texts, over a million human English texts - false negatives: 1,766 of 519,993 AI texts from 26 models, prompts taken from Chatbot Arena - the strict version: a "Mixed" answer counts as a miss on both sides - the blog on false positives (Bradley Emi, 15 Sep 2026) says "approximately 1 in 10,000" overall - domain rates from 0.0 percent (code documentation, congressional transcripts, movie scripts, political speeches) to 0.23 percent (recipes) - they report three kinds of test set: 10,000 to 10,000,000 examples, about 1,000, and about 10 to 100 - they recommend against screening short bullet lists, math, single sentences and very formulaic text - their own comparison numbers: Turnitin "0.51% on academic writing", GPTZero 2.01 percent measured versus 1 percent claimed, Copyleaks claims 0.2 percent - non-native English - Pangram 4 report, table 9: 1 false positive in 24,586 learner texts, 0 of 89 in the Liang TOEFL set - other languages - reports 14 false positives among 996,273 examples in 104 languages - its multilingual missed-detection aggregate needs clarification - the current model card lists 23,578 misses among 190,149 texts alongside 1.24 percent - my arithmetic gives 12.4 percent from those counts - use the language-specific results until the vendor reconciles this discrepancy - length - their figure 4 shows error rates by length bucket from 50 to 499 words, short is worse - AI-polished text - 11,363 human student texts edited lightly with Grammarly, Apple Intelligence, Gemini in Docs and frontier LLMs - Pangram 4 called 1 of them AI (0.01 percent), Pangram 3.3.2 called 20 (0.18 percent) - how much "AI-assisted" it catches - edited with clear intent to change a lot: 55.01 percent flagged Mixed - WildChat-style edits with AI fraction 0.25 to 0.75: 65.17 percent - so a large share of heavily edited text still passes as Human (41.37 percent in the first set) - humanizers - 97.67 percent flagged AI, 98.83 percent AI or Mixed, across 13 commercial humanizers (report, section 5.9.1) - the humanizer probe alone: accuracy 96.82 percent, false positive rate 1.93 percent - BLADER, a popular GitHub prompt pack for "remove signs of AI-generated writing from text", missed 0.430 percent - red-teaming - they gave two Codex agents (GPT-5.6 Sol) 24 hours on the API - the false-positive hunter found nothing new - the evasion agent found one bypass, dictated surgical pathology notes, and the authors call it out of scope - limits in their own words (report, conclusion) - "Predictions for the same text in different contexts may also be inconsistent" - "Pangram 4 does not account for people who intentionally write like LLMs or who have absorbed elements of LLM style into their writing" - original 2024 report - 1,976 documents, 10 domains, 8 LLMs, compared with GPTZero, Originality and DetectGPT - the arXiv v2 abstract says "over 9 times lower error rates", the current arXiv abstract says "over 38 times lower error rates", so they updated the claim - in v2: Pangram accuracy 99 percent, "commercial competitors which do not even clear 95%" - v2 said GPTZero's false negative rate was 10.02 percent and Originality's false positive rate was 9.24 percent - v2 explained why GPTZero and Originality failed on GPT-4 text: their "perplexity and burstiness features" lose on stronger models humanizer paper - DAMAGE: Detecting Adversarially Modified AI Generated Text, Masrour et al., arXiv, 2025 - abstract: "We study 19 AI humanizer and paraphrasing tools" and "show that many existing AI detectors fail to detect humanized text" - fix: "a data-centric augmentation approach", meaning add humanized AI text to training - they also attacked their own detector with a fine-tuned attacker and say it "remains fairly robust" - I did not read the result tables of this paper independent tests of Pangram - Jabarian and Imas (see section below): near-zero errors - Russell et al.: tied with the best human experts - People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text, Russell et al., ACL, 2025 - table 2: Pangram Humanizers overall TPR 99.3, FPR 2.7, Pangram 98.0 (2.0), GPTZero 85.3 (0.7), Fast-DetectGPT 80.0 (7.2), Binoculars 66.7 (1.3), RADAR 15.3 (2) - the five-expert majority vote also reached 99.3 TPR with 0 FPR - Pangram and GPTZero gave the authors API credits: "We thank Pangram and GPTZero for providing credits to access their APIs" - Russell is also an author of the Pangram 4 report, so Pangram is not far from this study - an outsider trying to fool it - Pangram (AI detection software) can be evaded, Eye You, LessWrong, 30 Mar 2026 - method as I read it: ask GPT-5.4-thinking for text in the style of Socrates dialogues and accept its offers of "more authentic" versions - results: first output 63 percent AI, then 20 percent, then 100 percent human - the author also writes that Pangram "is unreliable on text of fewer than 200 words" - this is one blog post and one text, not a measurement - Pangram's page of third-party studies (vendor-curated list) - pangram.com/blog/third-party-pangram-evals lists the VUB study (June 2026), Jabarian and Imas, a Maryland study, Tom's Guide, ZDNET and one Medium post - I could not open the VUB paper, only Pangram's table of it (see cross-claims) price - individual plan 20 a month, 300,000 words a month
  • professional plan 200 of API credits monthly“
    • both from pangram.com/pricing
  • API: 0.05 per 100 words, representing a significant price increase“)
    • I could not confirm that from Pangram’s own page, only from The Decoder
  • from Jabarian and Imas, older API prices: about 0.0016 per short one
  • no vendor accuracy claim should be read without this note: The Decoder says the numbers “derive from the company’s own benchmarks, not independent testing” (my summary of its caveat, not a quote)

GPTZero what goes in

  • pasted text, docx, pdf or image files, up to 50 files at a time (gptzero.me/technology) what the model is
  • 2023: a stack of parts
    • “Inside GPTZero AI Detection”, June 22, 2023, lists 7 layers: burstiness, GPTZeroX sentence classifier, perplexity, an education module, internet text search, a shield against bypass tools, and deep learning
    • perplexity as they explain it: “After each word in the text, our AI model develops suggestions of what word is coming next. It checks if our suggestions match what is actually there in the text”
    • burstiness: “A human written document will have changes in style and tone throughout the text, whereas AI content remains similar throughout”
  • mid 2023: a deep learning model added
    • “Deep learning model updates” page: updated bi-weekly, AUC “0.85” before and “0.95” after on GPT-4 data, vendor numbers
  • now: one trained network
    • the technology page: “an end-to-end deep learning approach, trained on text datasets from the web, education, and AI-generated from a range of LLMs”
    • GPTZero: Robust Detection of LLM-Generated Texts, Adam et al., arXiv, 2026, section 3.2
    • three classes at the top: Human, AI, Mixed
    • under AI: Pure AI, Polished (human text cleaned by an AI), AI Paraphrased
    • it predicts the document class and each sentence’s AI probability in one pass
    • the benefit of the Mixed class, in the paper’s words: it “resolves the confusing ambiguity of binary classifiers, where a 50% AI-generated score could either mean the detector is 100% confident that 50% of the text is AI-generated, or that the detector is 50% confident that the document is entirely AI-generated”
    • “Architecture and hyperparameters are proprietary”
    • the text is cleaned, windowed if long, and a remapping function adjusts the probabilities to cut false positives
    • Deep Scan: a way to say which sentences moved the score most, built from saliency and occlusion what data it was trained on
  • paper appendix, table 4: academic 1.25M documents, conversation 32K, encyclopedia 2.8M, essay 234K, news 16M, Q&A 65K, reviews 173K, web articles 8M
  • public datasets plus their own generations from OpenAI, Anthropic, Google and open models
  • a dispute button for users, used to find weak spots
  • red-teaming by augmenting with four kinds of attack: paraphrase prompts and translation chains, paraphrase models (Dipper, TempParaphraser), 9 black-box bypasser services, and white-box word-swap attacks what comes out
  • API: HUMAN_ONLY, MIXED, AI_ONLY with probabilities and a confidence category of high, medium or low (gptzero.me/technology)
  • sentence highlights vendor claims
  • technology page and FAQ: “Our false positive rate is under 1%”, 96.5 percent accuracy on mixed documents, TOEFL false positive rate 1.1 percent, “high” confidence has error rate under 1 percent
  • paper (own tests, 1000 human and 1000 AI texts per domain, AI from GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.5, Grok 4 Fast)
    • table 1 and 2 recall: GPTZero 99.2 abstracts, 97.4 creative, 99.7 essays, 99.7 paper reviews, 98.3 product reviews
    • Pangram 3.1 recall: 87.2, 96.0, 99.8, 97.2, 88.8
    • Originality lite-102 recall: 95.1, 92.1, 99.2, 94.9, 82.6
    • their reading: “GPTZero is the only detector which consistently has a sub 1% false positive rate across all domains while achieving a recall >97%”
    • humanizers: “GPTZero has a recall of 93.5%, while Originality and Pangram have a recall of 57.3% and 49.7% respectively”
    • note they tested Pangram 3.1, an older version than the one in Pangram’s 2026 report
    • open-source baselines (HC3, Radar, Fast-DetectGPT, Binoculars) fell near chance on abstracts and reviews
  • what the paper admits: “in-distribution performance metrics are overly optimistic” and the lack of shared test sets “introduces the risk of cherry-picking” independent tests
  • Jabarian and Imas: lowest false negative rate of the three in one summary, but weaker on short text and humanizers
    • Chicago Booth Review summary: GPTZero false negatives between roughly 0 and 2 percent, Pangram 2 to 4 percent, Originality 10 to 40 percent
    • that conflicts with the paper abstract’s “near-zero FNR” for Pangram, so the two should be read side by side (the Booth number is a journalist’s summary)
  • RAID: at default thresholds GPTZero had a false positive rate of 0.03 percent or lower; “unusually robust to adversarial attacks”
    • but accuracy at 5 percent FPR dropped to 9.4 percent on one model/setting and 34.6 or 4.8 on others (table 5), so it is brittle on unusual generators
  • Russell et al.: 85.3 TPR, 0.7 FPR, “GPTZero struggles significantly on O1-PRO with and without humanization”
  • Weber-Wulff (2023, old model): GPTZero had the most false positives among 14 tools, “from 0% (Turnitin) to 50% (GPT Zero)”
    • same study: false negatives from 8 percent (GPTZero) to 100 percent (Content at Scale)
  • researchers who use it as a measuring stick
  • the pricing page I fetched did not show numbers in the text I could extract
  • Jabarian and Imas, older API data: average 0.0040 per short one

Turnitin what goes in

  • student submissions inside the Turnitin product, not a public API
  • I could not open the Turnitin white paper itself (gated) or its FAQ page (HTTP 403)
    • what follows comes from two Turnitin blog posts I opened, the white paper’s landing page, a University of Bristol blog, and one secondary review what the model is
  • the white paper landing page, dated 18 Oct 2024, says it covers “the system’s architecture, testing protocol, and recent enhancements to the core AI detection model, including a new AI paraphrase detection feature”
  • the secondary review (fast.io, not Turnitin) says a paraphrase model called AIR-1 arrived in July 2024, and I take that as unconfirmed by Turnitin text I opened
  • Bristol blog, Sep 2025: Turnitin “expanded its capabilities to detect the output of AI tools that re-arrange AI-generated text in an attempt to bypass AI-detection tools (sometimes called ‘humanisers’)”
    • Bristol says there are no visual changes to the report
    • older submissions must be resubmitted to get it what comes out
  • a percent of the document that looks AI-written, plus highlighted sentences
  • the fast.io review says scores of 1 to 19 percent show an asterisk instead of a number because Turnitin’s tests “found higher false positive rates in this range” (secondary) vendor claims
  • Annie Chechitelli, Chief Product Officer, 16 Mar 2023: “less than 1% false positive rate, to ensure that students are not falsely accused”
    • same post: “given that our false positive rate is not zero, you as the instructor will need to apply your professional judgment”
  • same author, 14 Jun 2023: “Our sentence-level false positive rate is around 4%”
    • and “54% of the time, these sentences are located right next to actual AI writing”
    • meaning a wrongly highlighted sentence is usually near real AI text, not random independent tests
  • Weber-Wulff et al. 2023: Turnitin ranked first
    • Testing of detection tools for AI-generated text, Weber-Wulff et al., International Journal for Educational Integrity, 2023
    • 14 tools, Turnitin top accuracy: 76 percent (strict), 79 percent, 81 percent depending on how partial answers are scored
    • 0 false positives on their human documents, but missed a lot after paraphrase and machine translation
    • the paper’s own words: “the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text”
    • Turnitin approached the researchers and “offered a login”, and the conference they present at is sponsored by Turnitin, which the paper says did not influence them
  • Pangram’s re-run on the public VUB split of fully AI texts (Pangram 4 report, table 20)
    • Turnitin put all 39 of 39 AI documents in the 0 to 20 percent bucket
    • the same table: GPTZero 0 of 39 at 80 to 100 percent, Copyleaks 0 of 39
    • this is Pangram reporting on someone else’s data, so I mark it as a vendor-reported number
  • Pangram’s own test: Turnitin “0.51% on academic writing” false positive rate (Pangram false-positive blog, so Pangram measured a competitor)
  • Liang et al. (below) did not test Turnitin price
  • institutional licence, not stated in anything I opened

Originality.ai what goes in

  • pasted text, doc, pdf files, or a website address (originality.ai/pricing)
  • one credit is 100 words what the model is
  • the model release history on originality.ai/blog/ai-content-detection-accuracy
    • first model Nov 2022 “released before Chat-GPT”
    • 3.0 Turbo Feb 2024, “Trained on the newest LLMs (Grok, Mixtral, GPT-4 Turbo, Gemini, Claude 2)”
    • Lite (July 2024) allows “lightly AI-edited content (like Grammarly’s grammar and spelling suggestions)”
    • Turbo 3.0.2 and Academic 0.0.5 in Sep 2025
    • AI Allowance, July 2026: you choose “0%, 5%, 15%, 25%, or 40%” of AI you are willing to accept, and it measures a spectrum
  • what the network is, how big, which backbone: not stated on any page I opened what data it was trained on
  • not stated, beyond the listed LLM names what comes out
  • a score like “60% Original and 40% AI”, which the site says means “it is 60% confident that the content is original” (originality.ai/ai-content-detector-false-positives)
    • so the percent is confidence, not share of AI text, in the older models
    • AI Allowance is different: it asks how much of the text is AI vendor claims
  • the release history gives: Lite 1.0.2 “99% accuracy” with false positive rate 0.5 percent
    • Turbo 3.0.2 “up to 97% accuracy on the latest AI humanizers”, false positive rate 1.5 percent
    • Academic: false positive “<1%”
    • Multilingual 2.0.0, 30 languages, accuracy 97.8 percent, false positive rate 2.4 percent
    • AI Allowance: “99%+ accurate at 15% settings”, “96.53% overall accuracy at the strict 5% boundary” on an internal benchmark of 167,980 samples
  • the old Turbo 3.0 had a stated false positive rate of 2.8 percent, and the page mixes “accuracy” and false positive numbers in ways that are not comparable across models independent tests
  • Jabarian and Imas: second tier
    • false positive rate low (0.0011 at most thresholds), but false negatives much higher, and rejects some short passages with a “length filter” error
    • false negative rate on StealthGPT text about 0.05 on long text, up to 0.21 on short
  • RAID (Dugan et al.): accuracy at 5 percent FPR was 98.6 or 99.9 on easy settings, but 51.2 on the hardest column of table 5
    • in a later table (an attack breakdown, I did not work out which attack) one Originality cell fell by 75.7 points to 9.3
    • “Originality achieved high precision in some constrained scenarios”
    • false positive rates at default threshold 0.47 percent, 0.25 percent, 0.17 percent, 0.07 percent at four thresholds
  • GPTZero’s own paper: high recall but false positive rate “high enough such that users would disqualify it”
  • Pangram’s early report: false positive 9.24 percent on its 2024 set
  • these four tests disagree a lot, which is the main point price
  • 12.95 a month billed yearly, 2,000 credits a month, i.e. 200,000 words (originality.ai/pricing, as I read the page text)
  • Jabarian and Imas, older API data: average 0.0029 per short one

Copyleaks what goes in

  • text of at least 350 characters, the minimum the product takes (Copyleaks testing methodology page) what the model is
  • not described beyond “V11 model”
  • the testing page says model tested was V11, test date September 10, 2026 what data it was trained on
  • not stated on pages I opened
  • they only say the test data “differed from training data” what comes out
  • not described on the pages I opened vendor claims (own tests)
  • data science team: 500,000 texts, 300,000 human and 200,000 AI, TPR 0.989, TNR 0.999, “Internal extra-hard datasets, including adversarial attacks and special tools”
  • QA team: 326,120 human texts, 923 flagged as AI, accuracy 0.9970
    • web pages 0.9992, student essays 0.9999, scholarly papers 0.9997, but “Articles, news, blogs, social posts” 0.9929
  • Pangram quotes Copyleaks as claiming “0.2% false positive rate”
  • the blog “The AI Content Detector Continues To Be Confirmed As Most Accurate By Third-Party Studies” (July 8, 2025)
    • this is Copyleaks picking studies that rank it first
    • it lists a 2023 arXiv study with 124 student submissions where “GPTZero” scored 54.39 percent on human data and Copyleaks 99.12 percent
    • treat that list as advertising independent tests
  • RAID (2024) did not include Copyleaks as a commercial detector
  • Perkins et al. (via Pangram’s table): baseline AI mean accuracy 73.9 percent, manipulated 58.7 percent for Copyleaks, compared with Pangram 4’s 100 and 94.1
    • this table is Pangram’s, I did not open Perkins et al.
  • VUB table in the Pangram 4 report: Copyleaks gave none of 39 AI documents a score above 40 percent price
  • not found on the pages I could read

Winston AI

  • homepage claim: “99.87% Accurate”, “the only AI detector with a 99,87% accuracy rate” (gowinston.ai)
    • the page does not give a test set, a model description or a false positive rate
  • input: pasted text, .docx, .png, .jpg, with OCR
  • output: 0 to 100 prediction of human or AI, and a sentence “AI Prediction Map”
  • 14 languages listed, including English, French, Spanish, German, Chinese
  • the page does not mention humanizer detection
  • independent: Weber-Wulff et al. 2023 (older version, “Go Winston”)
    • 67 percent strict accuracy, 75 percent in the more forgiving scoring, ranked 4th or 5th of 14
    • RAID: Winston false positive rate 0.55 percent at threshold 0.5, accuracy at 5 percent FPR 97.2 percent on one setting, 68.2 percent on another, 29.5 percent on a third (table 5)
  • price: no figure on the page I opened

independent evaluations, in more detail Jabarian and Imas

  • Artificial Writing and Automated Detection, Jabarian et al., NBER Working Paper 34223 (also Becker Friedman Institute BFI 2025-116), 2025
    • not peer reviewed, per the working paper cover
    • declared no conflicts of interest, funded by Chicago Booth and Google Cloud research program
  • set up
    • 1,992 human passages: 200 Amazon reviews, 200 blogs, 300 news, 1,000 novel excerpts, 100 restaurant reviews, 192 resumes, all from before 2020
    • AI versions from GPT-4.1, Claude Opus 4, Claude Sonnet 4, Gemini 2.0 Flash, matched on content and length
    • detectors: Pangram, Originality.ai, GPTZero, and RoBERTa-base as the free baseline
    • for each, they report AUROC, error rates at fixed thresholds, error rates at the best threshold (Youden), short “stubs” under 50 words, and StealthGPT humanizer text
  • findings
    • abstract: “Commercial detectors outperform open-source, with Pangram achieving near-zero FNR and FPR rates that remain robust across models, threshold rules, ultra-short passages, ‘stubs’ (< 50 words) and ‘humanizer’ tools”
    • false positive rates at fixed thresholds (their table 3): GPTZero 0.0071, Originality 0.0011, Pangram 0.0000 at threshold 0.5 and up, RoBERTa 0.95
    • stubs: Originality gave a “length filter” error for some, so was left out
    • StealthGPT: “GPTZero largely loses its capacity to detect AI-generated text, with FNR scores around 0.50 and above”
    • Pangram on long text detects “nearly 100%” even after StealthGPT
  • policy cap idea
    • choose the false positive rate you can live with (for example 0.5 percent), set each detector’s threshold to hit it, then compare how much AI each one still catches
    • “Pangram is the only tool to satisfy a strict cap (FPR 0.005) without sacrificing accuracy”
  • limits the authors state
    • results “will likely be a continuous technical arms race” so they ask for routine published audits
    • they also say a conservative cap would let light AI help like Grammarly through
  • my caution
    • one fixed corpus, English, pre-2020 human text, passages up to about 1,000 words
    • the Pangram 4 report lists this benchmark as one of its own check marks, so Pangram has since tuned against it Weber-Wulff et al.
  • see the Turnitin section for the setting: 14 tools, Turnitin and PlagiarismCheck commercial
  • the paper’s own words on use: “Reports provided by detection tools for AI-generated text cannot be used as the only basis for reporting students for cheating”
  • noteworthy: “machine translation such as Google translate or DeepL can lead to a higher number of false positives”
  • these are 2023 tools, so none of the numbers describe today’s versions Liang et al. on non-native writers
  • GPT detectors are biased against non-native English writers, Liang et al., Patterns, 2023
  • seven detectors, including Originality.AI, Sapling, Crossplag, GPTZero, ZeroGPT, tested on 91 TOEFL essays and 88 US 8th grade essays
    • “average false positive rate: 61.22%” on TOEFL
    • all seven flagged 18 of 91 TOEFL essays (19.78 percent) as AI
    • 89 of 91 were flagged by at least one detector
    • near perfect on US 8th grade essays
  • why: low perplexity
    • rewriting the TOEFL essays with “Enhance the word choices to sound more like that of a native speaker” cut the average false positive rate from 61.22 percent to 11.77 percent
    • making the US essays sound simpler pushed false positives from 5.19 percent to 56.65 percent
  • also a bypass: a second-round prompt “Elevate the provided text by employing literary language” lowered detection from 100 percent to 13 percent
  • limits they state: small sample, and “most of the detectors assessed in this study utilize GPT-2 as their underlying backbone model”
  • so this tells us about perplexity-style detectors of 2023, not the trained ones now
  • GPTZero (1.1 percent) and Pangram (0 percent) both claim to have fixed this, each on the same 91 essays RAID
  • RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Dugan et al., ACL, 2024
    • over 6 million texts, 11 generators, 8 domains, 11 attacks, 4 decoding strategies
    • 12 detectors: 8 open-source, 4 commercial (GPTZero, Originality, Winston, ZeroGPT)
  • false positive rates at naive thresholds (table 4), thresholds 0.25, 0.5, 0.75, 0.95
    • GPTZero 0.03, 0.00, 0.00, 0.00 percent
    • Originality 0.47, 0.25, 0.17, 0.07 percent
    • Winston 0.75, 0.55, 0.38, 0.21 percent
    • ZeroGPT 1.71, 1.42, 1.21, 0.90 percent
    • open-source RoBERTa and GLTR were often over 10 percent
  • accuracy at 5 percent FPR (table 5) varies hugely by setting
    • Originality: 98.6 and 86.3 on two chat-model columns, 99.9 and 64.1 on another, 89.0 and 51.2 on another
    • GPTZero: 98.8 and 93.7 on the first, 74.7 and 34.6, then 9.4 and 4.8
    • so no single number tells you how a detector behaves
  • authors’ words: “detectors tend not to generalize across different models or generation settings in the same domain”
  • the repetition penalty (a setting that discourages repeated words) drastically cut accuracy for every class of detector
  • good signs they list: “Binoculars 
 performed impressively well across models even at extremely low false positive rates, Originality achieved high precision in some constrained scenarios, and GPTZero was unusually robust to adversarial attacks”
  • I opened the paper, not the live leaderboard, and I do not know the current standings
    • the leaderboard is where detector scores are submitted, and vendors can tune to it Sadasivan et al. stress test
  • Can AI-Generated Text be Reliably Detected?, Sadasivan et al., arXiv, 2023
    • recursive paraphrasing “can significantly reduce detection rates” while “only slightly” lowering text quality
    • attacked watermarks, neural detectors, zero-shot and retrieval-based detectors
    • they give a theory: as AI text gets closer to human text, even the best detector’s AUROC is bounded
    • I read the abstract only, and it predates the current commercial models Russell et al. on people
  • see Pangram section for the detector table
  • 300 non-fiction articles, annotators who “frequently use LLMs for writing tasks”
    • “The majority vote among five such ‘expert’ annotators misclassified only 1 of 300 articles”
    • per-annotator results were uneven (one expert got 59.3 percent TPR overall) Pangram-style measurement studies
  • Pangram 4 report intro claims: “9% of news articles (Russell et al., 2026a)” and “21% of machine-learning conference reviews (Pangram Labs, 2025)” were largely or fully AI generated
    • I did not open the news study or the Pangram review study, so I only report the claim as Pangram’s summary
  • GPTZero used the same way
    • Latona et al. (above), 15.8 percent of ICLR 2024 reviews
  • not Pangram or GPTZero

cross-claims and surprises

  • every vendor wins its own test
    • Pangram’s paper beats GPTZero and Originality; GPTZero’s paper beats Pangram and Originality; Copyleaks and Originality each publish studies where they rank first
  • GPTZero versus Pangram on humanizers is a flip
    • Jabarian and Imas (independent): GPTZero FNR around 0.50 or more, Pangram almost zero
    • GPTZero’s paper: GPTZero 93.5 percent recall, Pangram 49.7 percent
    • the GPTZero paper used Pangram 3.1; Pangram 4 was released after
  • Pangram edits its own claims
    • original report: “9 times lower error rates”, later arXiv version: “38 times”
  • “AI-assisted” detection is still weak
    • Pangram 4 itself flags only about 55 to 65 percent of heavily AI-edited text as Mixed
  • the DeGenTWeb link
    • these detectors score a single text, they are built for 50+ words of prose
    • none of them publishes how a site-level aggregate should be built
    • Pangram 4’s own per-token output (fraction of AI) could replace Binoculars page scores, but needs price check at about $0.05 per 100 words

what I could not open

  • Turnitin white paper text and FAQ page (gated or HTTP 403)
  • the VUB paper itself, Perkins et al., Epoch AI style imitation, the Pangram news and ICLR-review measurement papers
  • Copyleaks pricing and model description, GPTZero pricing numbers
  • web search ran out of quota near the end, so Winston and Copyleaks independent tests are thin

Last edited: