Benchmark

How often our AI detector is wrong about a human

Every AI detector returns a confident percentage. None of them tells you how often that percentage is wrong about someone who wrote their own work. That second number is the one that matters if you are the student being accused, so it is the one we publish.

Last measured 24 August 2026. Every figure below comes from a script in the repository that can be re-run against the same labelled corpora.

Measured false-positive rates

The rate at which verified human-written academic text receives our strongest verdict. Accuracy is not one number: it changes with language and with document length, so we report four profiles rather than an average that would hide the weak ones. Every rate is a 95% upper bound rather than the count we observed — the bound is what we are entitled to claim, and it is always the worse of the two. The thresholds in these tables are on our internal measurement scales; the score a customer sees is the same information on one fixed display axis - amber begins at 20, red at 50 - with each measured threshold mapped exactly onto those anchors.

ProfileStrongest verdictFalse positivesObservedMeasured on
English · over 400 wordsStrong AI fingerprintsat most 3.13%2 of 199199 validation-half EN documents (blended scale, w_style=0.5)
German · over 400 wordsStrong AI fingerprintsat most 1.63%10 of 1,0351,035 validation-half German theses (blended scale, w_style=0.5)
English · under 400 wordsStrong AI fingerprintsnot published—no relevant human population held
German · under 400 wordsno “strong AI fingerprints” verdict is issued———

Short German text never receives a “strong AI fingerprints” verdict, and never will on the current evidence.

On German text under 400 words, no threshold strong enough to catch AI reliably costs fewer than about 7 false accusations in 100. One innocent student in fifteen is not a trade we are willing to make, so the German short-form profile stops at “some AI fingerprints”. This is a measured limit, not a feature we have yet to build.

The weaker verdicts, which are noisier

Most readers never see the strongest verdict. They see one of these, and the error rates here are the ones that will actually apply to them. The English lower amber band (“some AI fingerprints”) is the worst of them by a wide margin.

ProfileVerdictThresholdFalse positivesObserved
English · over 400 wordsSome AI fingerprints≥ 96at most 5.86%6 of 199
English · over 400 wordsSome AI fingerprints (lower amber band)≥ 91at most 8.37%10 of 199
German · over 400 wordsSome AI fingerprints≥ 88at most 4.02%31 of 1,035
German · over 400 wordsSome AI fingerprints (lower amber band)≥ 85at most 6.18%51 of 1,035

What these numbers are not

What we measured against

Every human document below was published before ChatGPT was available, so its authorship is not a matter of judgement.

German student theses, 2010-2021

n = 866

Whole documents, in the language and the genre we actually serve, human by construction because they predate ChatGPT. This is the corpus the German thresholds are set on.

scripts/fetch-thesis-negatives.mjs

Blend validation half: German student theses, 2008-2021, whole documents

n = 1,035

The held-out half of 2,079 real German theses (MOnAMi, Hochschule Mittweida, 2008-2021) scored by BOTH halves of the check - whole documents, in the genre we serve, never used to train the German model. The blended German cutoffs are cut on it, so its rates are the ones the German verdicts publish. Not customer-era writing: a polished 2026 thesis is a population nobody has measured, and that caveat is part of the number.

scripts/calibrate-components.mjs

Open-access papers, pre-2022

n = 1,412

Full-length academic prose. Journal articles are edited and written by teams, so they are not student writing — which is why the English thresholds are not moved on them.

scripts/fetch-longform-corpus.mjs

Blend validation half: papers + arXiv, 1,500-word slices

n = 199

The held-out half of 400 English documents (arXiv and pre-2022 open-access papers) scored by BOTH halves of the check. The blended cutoffs are cut on it, so its rates are the ones the blended verdicts publish. Journal prose, not student writing - the population is named because that caveat is part of the number.

scripts/calibrate-components.mjs

arXiv abstracts, pre-2022

n = 236

Short-form English. Retained for AUC work; no false-positive rate is published from it, because a 300-400 word verdict window is not what an abstract corpus measures.

scripts/fetch-human-corpus.mjs

German academic abstracts, pre-2022

n = 60

OpenAlex, filtered on German function words because the language field is unreliable. Behind the decision to issue no “strong AI fingerprints” verdict on short German at all.

scripts/fetch-german-corpus.mjs

We once shipped a detector that was worse than a coin flip

In an early version, measurement against 236 pre-2022 arXiv abstracts returned an AUC of 0.184. A coin flip is 0.50. The detector was ranking machine-written text as more human than human-written text; inverting its output would have improved it. It had been returning confident percentages the entire time.

The cause was a family of lexical-richness features — Yule’s K, MTLD, the Zipf exponent, hapax legomena — all running backwards with large effect sizes. They score dense human academic prose as machine-like and lexically varied machine prose as human. That is the same mechanism behind the published finding that detectors falsely flag a majority of non-native English writers.

Those features now carry zero weight. They are still computed and still shown as evidence; they simply do not move the score. We publish this because it is the strongest argument we have for measuring: the failure was invisible from the output and obvious from the corpus, and any detector that has never run this test has no idea whether it has the same problem.

What no detector can do

No AI detector, including this one, is proof of authorship. These tools measure statistical properties of text, and human writing that is formal, structured, or written by a non-native speaker shares many of those properties. A score is a reason to ask a question, never a reason to conclude an answer. If a detector has been used against your work, our guide on what to do when a detector flags your thesis sets out how to respond.

Check your own work against these numbers

The first 1,500 words are free, with no account. Every verdict has a measured false-positive rate behind it, and this page is where all of them are published.

Run the free AI check

Researchers and journalists: the corpora, the scoring code and the calibration harnesses are all in the repository behind this site. If you want to reproduce or challenge any figure here, write to contact@thesisdraft.com and we will help.