commercial LLM text detectors (authored by agents unless marked đ§)
the plain picture
- a commercial detector is a text classifier behind a website and an API
- you paste text, it returns a score
- the score is âhow AI is thisâ, sometimes split per sentence
- the old story was perplexity and burstiness
- perplexity: how predictable each word is to a language model
- burstiness: how much that predictability swings across sentences
- GPTZero started there (2023), Turnitinâs own white paper still names both words
- the current story is a trained neural classifier
- take a language-model backbone, fine-tune it on lots of human text and lots of AI text
- no hand-made features needed, the network finds the tells itself
- Pangram, GPTZero, Turnitin and (as far as vendors say) Originality and Copyleaks work this way
- architecture, data and thresholds are almost always secret
- training data is a major part of the method
- Pangram: mirror each human text with an AI text on the same topic and length, then keep adding the human texts the model gets wrong
- GPTZero: huge pool, user dispute button, many fake âhumanizerâ rewrites added to training
- every vendor says false positives matter most, then reports a tiny false positive rate
- these numbers come from the vendorâs own test sets
- when two vendors test each other, each one wins its own test (see the cross-claims section)
- independent tests exist, but only a few
- they mostly agree that the trained commercial tools beat free open-source tools by a lot
- they disagree on who is best among the commercial ones
- Pangram looks best in the independent tests I could open, mainly Jabarian and Imas
main takeaways
- vendor numbers are not comparable
- different test sets, different thresholds, different definitions of âAIâ
- Pangram 4 counts âmixedâ as an error on both sides
- Originalityâs new model asks âhow much AIâ instead of âAI or notâ
- Jabarian and Imas provide a detailed independent head-to-head
- Pangram had about zero errors on normal-length text
- Originality and GPTZero were fine on long text, worse on short text and on humanizers
- the open-source RoBERTa flagged most human text as AI
- short text is the weak spot for everyone
- Pangram needs 50 words, Copyleaks needs 350 characters, Originality refuses some short passages
- humanizer tools (rewriters sold to beat detectors) split the field
- in Jabarian and Imas, GPTZero missed about half of humanized text, Pangram almost none
- in GPTZeroâs own paper, the numbers flip: GPTZero 93.5 percent, Pangram 49.7 percent
- old detectors were biased against non-native English writers
- Liang et al.: average false positive rate 61.22 percent on TOEFL essays
- GPTZero says 1.1 percent now, Pangram says 0 percent on the same essays, both are vendor claims
- a score is not a proof
- Turnitin itself says its false positive rate âis not zeroâ
- Pangram 4 says its own limit: it does not handle people who write like an LLM
- for DeGenTWeb (whole websites)
- Pangram 4 cuts long text into 512-token windows and merges them, GPTZero also windows long text
- Pangramâs 50-word minimum is fine for article pages
- price matters at scale: Pangram API is 0.05 per 100 words per The Decoder's report of the launch reading guide for vendor claims - "accuracy" means nothing without the false positive rate and the threshold - RAID shows that a detector can look perfect at its default threshold and fail at 5 percent false positive rate - see the RAID section - a vendor number tested on a vendor-picked set is a vendor claim - I mark each one below as vendor claim or independent Pangram what goes in - text, at least 50 words - Pangram 4 Technical Report, section 3.1: "we require that text be at least 50 words long to be considered for analysis" - they define the target narrowly - same report, section 3.1: "Pangram defines AI-generated text as original, substantial natural-language prose generated by an LLM in response to an open-ended writing task or question" - short factual answers, math, code are out of scope - long documents are split into overlapping windows - same report, section 4.1: windows of at most 512 tokens with a stride of 256, predictions merged what the model is - 2024 version (then called Checkfor.ai) - "a slightly modified transformer-style architecture", from the original technical report, section 2.1 - how it works page, pangram.com/research/how-it-works: "approximately 1 million documents comprised of public and licensed human-written text" plus AI text from "GPT-4 and other frontier language models" - that page also says the output is a 0 or 1 prediction, which differs from later versions - Pangram 4 (July 2026) - built on "a popular open-weight MoE model" with the name hidden (Pangram 4 Technical Report, section 4.1) - the Pangram blog says "Pangram 4 is our largest model yet, with 6x as many parameters as Pangram 3.3" (pangram.com/blog/pangram-4-technical) - it is fine-tuned with LoRA adapters, in two stages - one head gives a 15-bucket "fraction of AI" for each 512-token segment - another head gives a label per token: human, ai-assisted, ai-generated - to let every token see the whole window, they feed the window twice (they call it Repeat2) - a separate humanizer probe decides if the text looks deliberately disguised - the humanizer head does not change the backbone, on purpose: "we opt to train the humanizer head as a probe to prevent the adverse impact of false positives" what data it was trained on - human text: licensed or owned - "Pangram 4 is not trained on user-submitted data, customer data from API users, or data obtained by unauthorized Internet crawls" (Pangram 4 report, section 2.1) - rough mix in the report: creative writing 22.2 percent, scientific and medical 18.0, reference and educational 15.9, consumer reviews 10.6, social Q&A and chat 10.4, general web 8.2, news 6.6, essays 4.9, professional 3.1 - AI text: made in-house, using "75 of the most popular models" (Pangram blog, pangram-4-technical) - the 2024 report used about 28 million human documents, all dated 2021 or earlier - table 4 of the original report: reviews 15,000,000, books 7,000,000, scientific papers 3,000,000 the two training tricks - synthetic mirrors - original report, section 4.2: "For each human example, we generate an AI-generated example that matches the original document on as many axes as possible" - Pangram 4 report, section 2.2: step 1 ask an LLM "What is the topic of this article?", step 2 ask it "Write an article about X" - why: "We want the model to learn 'How was this text written?' as opposed to 'What is this text about?'" - mirror prompts have to be hand-tuned to remove tells like "Sure, here is an essay" - hard negative mining - run the model over a huge pool of human text, find the human texts it wrongly calls AI, add them (plus their mirrors) and retrain - why it is needed: training "reaches convergence before the first epoch concludes" because most examples are easy - original report table 5: overall false positive rate on held-out human text fell from 2.29 percent to 0.25 percent after mining - the same table: Books 0.85 to 0.001 percent, Wikipedia 5.34 to 0.23 percent, email (never in training) 6.60 to 0.80 percent - the claim "100x-1000x" reduction comes from that table's caption what comes out - a fraction of AI-written text, plus a label of human, ai-assisted or ai-generated - Pangram 4 report, section 5.1: Human if f_human is at least 0.90, AI if f_AI is at least 0.80, otherwise Mixed - internal predictions are per token; displayed highlights cover sentence groups - Pangram 4 model card, postprocessing: âProduct output is constrained to sentence-level resolutionâ - adjacent labels are merged to a minimum of around two sentences - a document-level flag for "humanized" - AI-assisted label is for text where the human and the AI are tangled, for example "a human who writes a paragraph and then goes back and forth with an AI to edit it" (report, section 3.2) - EditLens is the earlier paper behind the edit-level idea - EditLens: Quantifying the Extent of AI Editing in Text, Thai et al., ICLR, 2026 - abstract as I read it on arXiv: a regression model that predicts "the amount of AI editing present within a text", F1 94.7 percent binary and 90.4 percent ternary, with Grammarly edits as a case study - the label for how much AI was used is built by comparing a human text with its AI-edited version clause by clause - same lines in both: human. same idea, new words: AI-assisted. no match: AI-generated - the AI fraction is AI-generated characters plus half of AI-assisted characters, over all characters vendor claims on accuracy (all Pangram's own tests) - Pangram 4 report abstract: "an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%" - about 1 false positive in 24,300 human texts, over a million human English texts - false negatives: 1,766 of 519,993 AI texts from 26 models, prompts taken from Chatbot Arena - the strict version: a "Mixed" answer counts as a miss on both sides - the blog on false positives (Bradley Emi, 15 Sep 2026) says "approximately 1 in 10,000" overall - domain rates from 0.0 percent (code documentation, congressional transcripts, movie scripts, political speeches) to 0.23 percent (recipes) - they report three kinds of test set: 10,000 to 10,000,000 examples, about 1,000, and about 10 to 100 - they recommend against screening short bullet lists, math, single sentences and very formulaic text - their own comparison numbers: Turnitin "0.51% on academic writing", GPTZero 2.01 percent measured versus 1 percent claimed, Copyleaks claims 0.2 percent - non-native English - Pangram 4 report, table 9: 1 false positive in 24,586 learner texts, 0 of 89 in the Liang TOEFL set - other languages - reports 14 false positives among 996,273 examples in 104 languages - its multilingual missed-detection aggregate needs clarification - the current model card lists 23,578 misses among 190,149 texts alongside 1.24 percent - my arithmetic gives 12.4 percent from those counts - use the language-specific results until the vendor reconciles this discrepancy - length - their figure 4 shows error rates by length bucket from 50 to 499 words, short is worse - AI-polished text - 11,363 human student texts edited lightly with Grammarly, Apple Intelligence, Gemini in Docs and frontier LLMs - Pangram 4 called 1 of them AI (0.01 percent), Pangram 3.3.2 called 20 (0.18 percent) - how much "AI-assisted" it catches - edited with clear intent to change a lot: 55.01 percent flagged Mixed - WildChat-style edits with AI fraction 0.25 to 0.75: 65.17 percent - so a large share of heavily edited text still passes as Human (41.37 percent in the first set) - humanizers - 97.67 percent flagged AI, 98.83 percent AI or Mixed, across 13 commercial humanizers (report, section 5.9.1) - the humanizer probe alone: accuracy 96.82 percent, false positive rate 1.93 percent - BLADER, a popular GitHub prompt pack for "remove signs of AI-generated writing from text", missed 0.430 percent - red-teaming - they gave two Codex agents (GPT-5.6 Sol) 24 hours on the API - the false-positive hunter found nothing new - the evasion agent found one bypass, dictated surgical pathology notes, and the authors call it out of scope - limits in their own words (report, conclusion) - "Predictions for the same text in different contexts may also be inconsistent" - "Pangram 4 does not account for people who intentionally write like LLMs or who have absorbed elements of LLM style into their writing" - original 2024 report - 1,976 documents, 10 domains, 8 LLMs, compared with GPTZero, Originality and DetectGPT - the arXiv v2 abstract says "over 9 times lower error rates", the current arXiv abstract says "over 38 times lower error rates", so they updated the claim - in v2: Pangram accuracy 99 percent, "commercial competitors which do not even clear 95%" - v2 said GPTZero's false negative rate was 10.02 percent and Originality's false positive rate was 9.24 percent - v2 explained why GPTZero and Originality failed on GPT-4 text: their "perplexity and burstiness features" lose on stronger models humanizer paper - DAMAGE: Detecting Adversarially Modified AI Generated Text, Masrour et al., arXiv, 2025 - abstract: "We study 19 AI humanizer and paraphrasing tools" and "show that many existing AI detectors fail to detect humanized text" - fix: "a data-centric augmentation approach", meaning add humanized AI text to training - they also attacked their own detector with a fine-tuned attacker and say it "remains fairly robust" - I did not read the result tables of this paper independent tests of Pangram - Jabarian and Imas (see section below): near-zero errors - Russell et al.: tied with the best human experts - People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text, Russell et al., ACL, 2025 - table 2: Pangram Humanizers overall TPR 99.3, FPR 2.7, Pangram 98.0 (2.0), GPTZero 85.3 (0.7), Fast-DetectGPT 80.0 (7.2), Binoculars 66.7 (1.3), RADAR 15.3 (2) - the five-expert majority vote also reached 99.3 TPR with 0 FPR - Pangram and GPTZero gave the authors API credits: "We thank Pangram and GPTZero for providing credits to access their APIs" - Russell is also an author of the Pangram 4 report, so Pangram is not far from this study - an outsider trying to fool it - Pangram (AI detection software) can be evaded, Eye You, LessWrong, 30 Mar 2026 - method as I read it: ask GPT-5.4-thinking for text in the style of Socrates dialogues and accept its offers of "more authentic" versions - results: first output 63 percent AI, then 20 percent, then 100 percent human - the author also writes that Pangram "is unreliable on text of fewer than 200 words" - this is one blog post and one text, not a measurement - Pangram's page of third-party studies (vendor-curated list) - pangram.com/blog/third-party-pangram-evals lists the VUB study (June 2026), Jabarian and Imas, a Maryland study, Tom's Guide, ZDNET and one Medium post - I could not open the VUB paper, only Pangram's table of it (see cross-claims) price - individual plan 20 a month, 300,000 words a month
- professional plan 200 of API credits monthlyâ
- both from pangram.com/pricing
- API: 0.05 per 100 words, representing a significant price increaseâ)
- I could not confirm that from Pangramâs own page, only from The Decoder
- from Jabarian and Imas, older API prices: about 0.0016 per short one
- no vendor accuracy claim should be read without this note: The Decoder says the numbers âderive from the companyâs own benchmarks, not independent testingâ (my summary of its caveat, not a quote)
GPTZero what goes in
- pasted text, docx, pdf or image files, up to 50 files at a time (gptzero.me/technology) what the model is
- 2023: a stack of parts
- âInside GPTZero AI Detectionâ, June 22, 2023, lists 7 layers: burstiness, GPTZeroX sentence classifier, perplexity, an education module, internet text search, a shield against bypass tools, and deep learning
- perplexity as they explain it: âAfter each word in the text, our AI model develops suggestions of what word is coming next. It checks if our suggestions match what is actually there in the textâ
- burstiness: âA human written document will have changes in style and tone throughout the text, whereas AI content remains similar throughoutâ
- mid 2023: a deep learning model added
- âDeep learning model updatesâ page: updated bi-weekly, AUC â0.85â before and â0.95â after on GPT-4 data, vendor numbers
- now: one trained network
- the technology page: âan end-to-end deep learning approach, trained on text datasets from the web, education, and AI-generated from a range of LLMsâ
- GPTZero: Robust Detection of LLM-Generated Texts, Adam et al., arXiv, 2026, section 3.2
- three classes at the top: Human, AI, Mixed
- under AI: Pure AI, Polished (human text cleaned by an AI), AI Paraphrased
- it predicts the document class and each sentenceâs AI probability in one pass
- the benefit of the Mixed class, in the paperâs words: it âresolves the confusing ambiguity of binary classifiers, where a 50% AI-generated score could either mean the detector is 100% confident that 50% of the text is AI-generated, or that the detector is 50% confident that the document is entirely AI-generatedâ
- âArchitecture and hyperparameters are proprietaryâ
- the text is cleaned, windowed if long, and a remapping function adjusts the probabilities to cut false positives
- Deep Scan: a way to say which sentences moved the score most, built from saliency and occlusion what data it was trained on
- paper appendix, table 4: academic 1.25M documents, conversation 32K, encyclopedia 2.8M, essay 234K, news 16M, Q&A 65K, reviews 173K, web articles 8M
- public datasets plus their own generations from OpenAI, Anthropic, Google and open models
- a dispute button for users, used to find weak spots
- red-teaming by augmenting with four kinds of attack: paraphrase prompts and translation chains, paraphrase models (Dipper, TempParaphraser), 9 black-box bypasser services, and white-box word-swap attacks what comes out
- API: HUMAN_ONLY, MIXED, AI_ONLY with probabilities and a confidence category of high, medium or low (gptzero.me/technology)
- sentence highlights vendor claims
- technology page and FAQ: âOur false positive rate is under 1%â, 96.5 percent accuracy on mixed documents, TOEFL false positive rate 1.1 percent, âhighâ confidence has error rate under 1 percent
- paper (own tests, 1000 human and 1000 AI texts per domain, AI from GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.5, Grok 4 Fast)
- table 1 and 2 recall: GPTZero 99.2 abstracts, 97.4 creative, 99.7 essays, 99.7 paper reviews, 98.3 product reviews
- Pangram 3.1 recall: 87.2, 96.0, 99.8, 97.2, 88.8
- Originality lite-102 recall: 95.1, 92.1, 99.2, 94.9, 82.6
- their reading: âGPTZero is the only detector which consistently has a sub 1% false positive rate across all domains while achieving a recall >97%â
- humanizers: âGPTZero has a recall of 93.5%, while Originality and Pangram have a recall of 57.3% and 49.7% respectivelyâ
- note they tested Pangram 3.1, an older version than the one in Pangramâs 2026 report
- open-source baselines (HC3, Radar, Fast-DetectGPT, Binoculars) fell near chance on abstracts and reviews
- what the paper admits: âin-distribution performance metrics are overly optimisticâ and the lack of shared test sets âintroduces the risk of cherry-pickingâ independent tests
- Jabarian and Imas: lowest false negative rate of the three in one summary, but weaker on short text and humanizers
- Chicago Booth Review summary: GPTZero false negatives between roughly 0 and 2 percent, Pangram 2 to 4 percent, Originality 10 to 40 percent
- that conflicts with the paper abstractâs ânear-zero FNRâ for Pangram, so the two should be read side by side (the Booth number is a journalistâs summary)
- RAID: at default thresholds GPTZero had a false positive rate of 0.03 percent or lower; âunusually robust to adversarial attacksâ
- but accuracy at 5 percent FPR dropped to 9.4 percent on one model/setting and 34.6 or 4.8 on others (table 5), so it is brittle on unusual generators
- Russell et al.: 85.3 TPR, 0.7 FPR, âGPTZero struggles significantly on O1-PRO with and without humanizationâ
- Weber-Wulff (2023, old model): GPTZero had the most false positives among 14 tools, âfrom 0% (Turnitin) to 50% (GPT Zero)â
- same study: false negatives from 8 percent (GPTZero) to 100 percent (Content at Scale)
- researchers who use it as a measuring stick
- The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates, Latona et al., arXiv, 2024
- abstract: âat least 15.8% of reviews were written with AI assistanceâ using âthe GPTZero LLM detectorâ
- the vendor page quotes this paper, so they use it as advertising too price
- the pricing page I fetched did not show numbers in the text I could extract
- Jabarian and Imas, older API data: average 0.0040 per short one
Turnitin what goes in
- student submissions inside the Turnitin product, not a public API
- I could not open the Turnitin white paper itself (gated) or its FAQ page (HTTP 403)
- what follows comes from two Turnitin blog posts I opened, the white paperâs landing page, a University of Bristol blog, and one secondary review what the model is
- the white paper landing page, dated 18 Oct 2024, says it covers âthe systemâs architecture, testing protocol, and recent enhancements to the core AI detection model, including a new AI paraphrase detection featureâ
- the secondary review (fast.io, not Turnitin) says a paraphrase model called AIR-1 arrived in July 2024, and I take that as unconfirmed by Turnitin text I opened
- Bristol blog, Sep 2025: Turnitin âexpanded its capabilities to detect the output of AI tools that re-arrange AI-generated text in an attempt to bypass AI-detection tools (sometimes called âhumanisersâ)â
- Bristol says there are no visual changes to the report
- older submissions must be resubmitted to get it what comes out
- a percent of the document that looks AI-written, plus highlighted sentences
- the fast.io review says scores of 1 to 19 percent show an asterisk instead of a number because Turnitinâs tests âfound higher false positive rates in this rangeâ (secondary) vendor claims
- Annie Chechitelli, Chief Product Officer, 16 Mar 2023: âless than 1% false positive rate, to ensure that students are not falsely accusedâ
- same post: âgiven that our false positive rate is not zero, you as the instructor will need to apply your professional judgmentâ
- same author, 14 Jun 2023: âOur sentence-level false positive rate is around 4%â
- and â54% of the time, these sentences are located right next to actual AI writingâ
- meaning a wrongly highlighted sentence is usually near real AI text, not random independent tests
- Weber-Wulff et al. 2023: Turnitin ranked first
- Testing of detection tools for AI-generated text, Weber-Wulff et al., International Journal for Educational Integrity, 2023
- 14 tools, Turnitin top accuracy: 76 percent (strict), 79 percent, 81 percent depending on how partial answers are scored
- 0 false positives on their human documents, but missed a lot after paraphrase and machine translation
- the paperâs own words: âthe available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated textâ
- Turnitin approached the researchers and âoffered a loginâ, and the conference they present at is sponsored by Turnitin, which the paper says did not influence them
- Pangramâs re-run on the public VUB split of fully AI texts (Pangram 4 report, table 20)
- Turnitin put all 39 of 39 AI documents in the 0 to 20 percent bucket
- the same table: GPTZero 0 of 39 at 80 to 100 percent, Copyleaks 0 of 39
- this is Pangram reporting on someone elseâs data, so I mark it as a vendor-reported number
- Pangramâs own test: Turnitin â0.51% on academic writingâ false positive rate (Pangram false-positive blog, so Pangram measured a competitor)
- Liang et al. (below) did not test Turnitin price
- institutional licence, not stated in anything I opened
Originality.ai what goes in
- pasted text, doc, pdf files, or a website address (originality.ai/pricing)
- one credit is 100 words what the model is
- the model release history on originality.ai/blog/ai-content-detection-accuracy
- first model Nov 2022 âreleased before Chat-GPTâ
- 3.0 Turbo Feb 2024, âTrained on the newest LLMs (Grok, Mixtral, GPT-4 Turbo, Gemini, Claude 2)â
- Lite (July 2024) allows âlightly AI-edited content (like Grammarlyâs grammar and spelling suggestions)â
- Turbo 3.0.2 and Academic 0.0.5 in Sep 2025
- AI Allowance, July 2026: you choose â0%, 5%, 15%, 25%, or 40%â of AI you are willing to accept, and it measures a spectrum
- what the network is, how big, which backbone: not stated on any page I opened what data it was trained on
- not stated, beyond the listed LLM names what comes out
- a score like â60% Original and 40% AIâ, which the site says means âit is 60% confident that the content is originalâ (originality.ai/ai-content-detector-false-positives)
- so the percent is confidence, not share of AI text, in the older models
- AI Allowance is different: it asks how much of the text is AI vendor claims
- the release history gives: Lite 1.0.2 â99% accuracyâ with false positive rate 0.5 percent
- Turbo 3.0.2 âup to 97% accuracy on the latest AI humanizersâ, false positive rate 1.5 percent
- Academic: false positive â<1%â
- Multilingual 2.0.0, 30 languages, accuracy 97.8 percent, false positive rate 2.4 percent
- AI Allowance: â99%+ accurate at 15% settingsâ, â96.53% overall accuracy at the strict 5% boundaryâ on an internal benchmark of 167,980 samples
- the old Turbo 3.0 had a stated false positive rate of 2.8 percent, and the page mixes âaccuracyâ and false positive numbers in ways that are not comparable across models independent tests
- Jabarian and Imas: second tier
- false positive rate low (0.0011 at most thresholds), but false negatives much higher, and rejects some short passages with a âlength filterâ error
- false negative rate on StealthGPT text about 0.05 on long text, up to 0.21 on short
- RAID (Dugan et al.): accuracy at 5 percent FPR was 98.6 or 99.9 on easy settings, but 51.2 on the hardest column of table 5
- in a later table (an attack breakdown, I did not work out which attack) one Originality cell fell by 75.7 points to 9.3
- âOriginality achieved high precision in some constrained scenariosâ
- false positive rates at default threshold 0.47 percent, 0.25 percent, 0.17 percent, 0.07 percent at four thresholds
- GPTZeroâs own paper: high recall but false positive rate âhigh enough such that users would disqualify itâ
- Pangramâs early report: false positive 9.24 percent on its 2024 set
- these four tests disagree a lot, which is the main point price
- 12.95 a month billed yearly, 2,000 credits a month, i.e. 200,000 words (originality.ai/pricing, as I read the page text)
- Jabarian and Imas, older API data: average 0.0029 per short one
Copyleaks what goes in
- text of at least 350 characters, the minimum the product takes (Copyleaks testing methodology page) what the model is
- not described beyond âV11 modelâ
- the testing page says model tested was V11, test date September 10, 2026 what data it was trained on
- not stated on pages I opened
- they only say the test data âdiffered from training dataâ what comes out
- not described on the pages I opened vendor claims (own tests)
- data science team: 500,000 texts, 300,000 human and 200,000 AI, TPR 0.989, TNR 0.999, âInternal extra-hard datasets, including adversarial attacks and special toolsâ
- QA team: 326,120 human texts, 923 flagged as AI, accuracy 0.9970
- web pages 0.9992, student essays 0.9999, scholarly papers 0.9997, but âArticles, news, blogs, social postsâ 0.9929
- Pangram quotes Copyleaks as claiming â0.2% false positive rateâ
- the blog âThe AI Content Detector Continues To Be Confirmed As Most Accurate By Third-Party Studiesâ (July 8, 2025)
- this is Copyleaks picking studies that rank it first
- it lists a 2023 arXiv study with 124 student submissions where âGPTZeroâ scored 54.39 percent on human data and Copyleaks 99.12 percent
- treat that list as advertising independent tests
- RAID (2024) did not include Copyleaks as a commercial detector
- Perkins et al. (via Pangramâs table): baseline AI mean accuracy 73.9 percent, manipulated 58.7 percent for Copyleaks, compared with Pangram 4âs 100 and 94.1
- this table is Pangramâs, I did not open Perkins et al.
- VUB table in the Pangram 4 report: Copyleaks gave none of 39 AI documents a score above 40 percent price
- not found on the pages I could read
Winston AI
- homepage claim: â99.87% Accurateâ, âthe only AI detector with a 99,87% accuracy rateâ (gowinston.ai)
- the page does not give a test set, a model description or a false positive rate
- input: pasted text, .docx, .png, .jpg, with OCR
- output: 0 to 100 prediction of human or AI, and a sentence âAI Prediction Mapâ
- 14 languages listed, including English, French, Spanish, German, Chinese
- the page does not mention humanizer detection
- independent: Weber-Wulff et al. 2023 (older version, âGo Winstonâ)
- 67 percent strict accuracy, 75 percent in the more forgiving scoring, ranked 4th or 5th of 14
- RAID: Winston false positive rate 0.55 percent at threshold 0.5, accuracy at 5 percent FPR 97.2 percent on one setting, 68.2 percent on another, 29.5 percent on a third (table 5)
- price: no figure on the page I opened
independent evaluations, in more detail Jabarian and Imas
- Artificial Writing and Automated Detection, Jabarian et al., NBER Working Paper 34223 (also Becker Friedman Institute BFI 2025-116), 2025
- not peer reviewed, per the working paper cover
- declared no conflicts of interest, funded by Chicago Booth and Google Cloud research program
- set up
- 1,992 human passages: 200 Amazon reviews, 200 blogs, 300 news, 1,000 novel excerpts, 100 restaurant reviews, 192 resumes, all from before 2020
- AI versions from GPT-4.1, Claude Opus 4, Claude Sonnet 4, Gemini 2.0 Flash, matched on content and length
- detectors: Pangram, Originality.ai, GPTZero, and RoBERTa-base as the free baseline
- for each, they report AUROC, error rates at fixed thresholds, error rates at the best threshold (Youden), short âstubsâ under 50 words, and StealthGPT humanizer text
- findings
- abstract: âCommercial detectors outperform open-source, with Pangram achieving near-zero FNR and FPR rates that remain robust across models, threshold rules, ultra-short passages, âstubsâ (< 50 words) and âhumanizerâ toolsâ
- false positive rates at fixed thresholds (their table 3): GPTZero 0.0071, Originality 0.0011, Pangram 0.0000 at threshold 0.5 and up, RoBERTa 0.95
- stubs: Originality gave a âlength filterâ error for some, so was left out
- StealthGPT: âGPTZero largely loses its capacity to detect AI-generated text, with FNR scores around 0.50 and aboveâ
- Pangram on long text detects ânearly 100%â even after StealthGPT
- policy cap idea
- choose the false positive rate you can live with (for example 0.5 percent), set each detectorâs threshold to hit it, then compare how much AI each one still catches
- âPangram is the only tool to satisfy a strict cap (FPR 0.005) without sacrificing accuracyâ
- limits the authors state
- results âwill likely be a continuous technical arms raceâ so they ask for routine published audits
- they also say a conservative cap would let light AI help like Grammarly through
- my caution
- one fixed corpus, English, pre-2020 human text, passages up to about 1,000 words
- the Pangram 4 report lists this benchmark as one of its own check marks, so Pangram has since tuned against it Weber-Wulff et al.
- see the Turnitin section for the setting: 14 tools, Turnitin and PlagiarismCheck commercial
- the paperâs own words on use: âReports provided by detection tools for AI-generated text cannot be used as the only basis for reporting students for cheatingâ
- noteworthy: âmachine translation such as Google translate or DeepL can lead to a higher number of false positivesâ
- these are 2023 tools, so none of the numbers describe todayâs versions Liang et al. on non-native writers
- GPT detectors are biased against non-native English writers, Liang et al., Patterns, 2023
- seven detectors, including Originality.AI, Sapling, Crossplag, GPTZero, ZeroGPT, tested on 91 TOEFL essays and 88 US 8th grade essays
- âaverage false positive rate: 61.22%â on TOEFL
- all seven flagged 18 of 91 TOEFL essays (19.78 percent) as AI
- 89 of 91 were flagged by at least one detector
- near perfect on US 8th grade essays
- why: low perplexity
- rewriting the TOEFL essays with âEnhance the word choices to sound more like that of a native speakerâ cut the average false positive rate from 61.22 percent to 11.77 percent
- making the US essays sound simpler pushed false positives from 5.19 percent to 56.65 percent
- also a bypass: a second-round prompt âElevate the provided text by employing literary languageâ lowered detection from 100 percent to 13 percent
- limits they state: small sample, and âmost of the detectors assessed in this study utilize GPT-2 as their underlying backbone modelâ
- so this tells us about perplexity-style detectors of 2023, not the trained ones now
- GPTZero (1.1 percent) and Pangram (0 percent) both claim to have fixed this, each on the same 91 essays RAID
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, Dugan et al., ACL, 2024
- over 6 million texts, 11 generators, 8 domains, 11 attacks, 4 decoding strategies
- 12 detectors: 8 open-source, 4 commercial (GPTZero, Originality, Winston, ZeroGPT)
- false positive rates at naive thresholds (table 4), thresholds 0.25, 0.5, 0.75, 0.95
- GPTZero 0.03, 0.00, 0.00, 0.00 percent
- Originality 0.47, 0.25, 0.17, 0.07 percent
- Winston 0.75, 0.55, 0.38, 0.21 percent
- ZeroGPT 1.71, 1.42, 1.21, 0.90 percent
- open-source RoBERTa and GLTR were often over 10 percent
- accuracy at 5 percent FPR (table 5) varies hugely by setting
- Originality: 98.6 and 86.3 on two chat-model columns, 99.9 and 64.1 on another, 89.0 and 51.2 on another
- GPTZero: 98.8 and 93.7 on the first, 74.7 and 34.6, then 9.4 and 4.8
- so no single number tells you how a detector behaves
- authorsâ words: âdetectors tend not to generalize across different models or generation settings in the same domainâ
- the repetition penalty (a setting that discourages repeated words) drastically cut accuracy for every class of detector
- good signs they list: âBinoculars ⊠performed impressively well across models even at extremely low false positive rates, Originality achieved high precision in some constrained scenarios, and GPTZero was unusually robust to adversarial attacksâ
- I opened the paper, not the live leaderboard, and I do not know the current standings
- the leaderboard is where detector scores are submitted, and vendors can tune to it Sadasivan et al. stress test
- Can AI-Generated Text be Reliably Detected?, Sadasivan et al., arXiv, 2023
- recursive paraphrasing âcan significantly reduce detection ratesâ while âonly slightlyâ lowering text quality
- attacked watermarks, neural detectors, zero-shot and retrieval-based detectors
- they give a theory: as AI text gets closer to human text, even the best detectorâs AUROC is bounded
- I read the abstract only, and it predates the current commercial models Russell et al. on people
- see Pangram section for the detector table
- 300 non-fiction articles, annotators who âfrequently use LLMs for writing tasksâ
- âThe majority vote among five such âexpertâ annotators misclassified only 1 of 300 articlesâ
- per-annotator results were uneven (one expert got 59.3 percent TPR overall) Pangram-style measurement studies
- Pangram 4 report intro claims: â9% of news articles (Russell et al., 2026a)â and â21% of machine-learning conference reviews (Pangram Labs, 2025)â were largely or fully AI generated
- I did not open the news study or the Pangram review study, so I only report the claim as Pangramâs summary
- GPTZero used the same way
- Latona et al. (above), 15.8 percent of ICLR 2024 reviews
- not Pangram or GPTZero
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews, Liang et al., arXiv, 2024
- uses a corpus-level estimate, not a per-document detector: âbetween 6.5% and 16.9% of text submitted as peer reviews ⊠could have been substantially modified by LLMsâ
- so it avoids the false-positive problem for individuals by only estimating a share
cross-claims and surprises
- every vendor wins its own test
- Pangramâs paper beats GPTZero and Originality; GPTZeroâs paper beats Pangram and Originality; Copyleaks and Originality each publish studies where they rank first
- GPTZero versus Pangram on humanizers is a flip
- Jabarian and Imas (independent): GPTZero FNR around 0.50 or more, Pangram almost zero
- GPTZeroâs paper: GPTZero 93.5 percent recall, Pangram 49.7 percent
- the GPTZero paper used Pangram 3.1; Pangram 4 was released after
- Pangram edits its own claims
- original report: â9 times lower error ratesâ, later arXiv version: â38 timesâ
- âAI-assistedâ detection is still weak
- Pangram 4 itself flags only about 55 to 65 percent of heavily AI-edited text as Mixed
- the DeGenTWeb link
- these detectors score a single text, they are built for 50+ words of prose
- none of them publishes how a site-level aggregate should be built
- Pangram 4âs own per-token output (fraction of AI) could replace Binoculars page scores, but needs price check at about $0.05 per 100 words
what I could not open
- Turnitin white paper text and FAQ page (gated or HTTP 403)
- the VUB paper itself, Perkins et al., Epoch AI style imitation, the Pangram news and ICLR-review measurement papers
- Copyleaks pricing and model description, GPTZero pricing numbers
- web search ran out of quota near the end, so Winston and Copyleaks independent tests are thin
Last edited: