← Back to blog

How Accurate Are AI Detectors? We Measured Ours and Published the Numbers

|8 min read

Check your writing for AI text — free

The first 1,500 words are free, with no sign-up. Every verdict shows how often it is wrong about verified human writing — a figure no other detector publishes.

Need a researched draft to work from instead? From €4.99, anonymous, one-time payment — see pricing.

The Number Nobody Publishes

Run your thesis through any AI detector and you get a confident number. 78% AI. What you do not get is the one figure that tells you whether to believe it: how often that tool says 78% about writing it knows a human produced.

That figure is the false positive rate, and the industry does not publish it. Not GPTZero, not Originality.ai, not Copyleaks, not Winston. You are shown a percentage with two decimal places and no error bar.

We decided to publish ours. Below are the actual measurements, the method, and the places our detector is weakest. If you are staring at a score right now and trying to work out what it means, the last section is the one you want.

How We Measured It

The hard part of testing a detector is getting text you are certain a human wrote. You cannot use a second detector to label your test set, because then you are measuring agreement between two tools rather than truth.

So we used the calendar. Every human sample in our test set was published before ChatGPT existed. Text written and published in 2018–2021 cannot have been generated by a model that launched in November 2022.

  • 236 English abstracts from arXiv, submitted 2018–2021, across economics, biology, statistics, computer science and social physics.
  • 40 full-length English documents (1,500 words each) from open-access papers in Europe PMC, body text only, with abstracts, references, tables and captions stripped out.
  • 60 German abstracts from OpenAlex, filtered to genuinely German-language text.
  • 38 full-length German documents from Wikipedia article revisions dated before June 2021.

Against those we tested machine-written text in the same registers and languages, generated to the same lengths. Then we counted how often each verdict was wrong.

What We Found

Two findings mattered more than the headline accuracy.

Length changes everything. A detector tuned on short abstracts performs badly on a real thesis, and not only because the thresholds shift. Entirely different features carry the signal. Paragraph-level consistency says nothing about a 180-word abstract and is one of the strongest signals across 1,500 words. Anyone quoting a single accuracy number for a tool has not tested it at the length you care about.

Language changes everything too. The features that work in English are not the features that work in German. We run separate models for each, calibrated separately, because using English settings on German text throws away most of the available signal.

Text typeHuman writing correctly identifiedMachine writing detected
English, full document100%100%
English, short abstract84%95%
German, full document100%100%
German, short abstract90%65%

Read those 100% figures with suspicion. They come from test sets of 40 and 38 documents. Zero errors in 40 samples does not mean zero percent, it means the true rate is probably under about 7%. That is why the figure shown next to your verdict in our tool is 3%, not 0%. Publishing 0% would be exactly the overclaiming this article is complaining about.

Why Detectors Flag Human Writing

Most detectors, including earlier versions of ours, lean on lexical variety: how wide and unusual your vocabulary is. Measures with names like Yule's K, MTLD and the Zipf exponent all capture roughly this.

The assumption is that machines write with less variety than people. In academic writing, the opposite is true. A real thesis repeats its key terms relentlessly, because precision demands it. You do not elegantly vary between "participants", "subjects" and "respondents" when they mean different things. A language model, optimising for readable prose, varies its wording far more.

So a lexical-variety detector systematically flags exactly the wrong people:

  • Writers working in a second or third language, whose vocabulary range is narrower
  • Technical and STEM writers, whose terminology is fixed by the field
  • Anyone who edits carefully toward a consistent, disciplined style

This is the mechanism behind the Stanford finding that seven major detectors flagged 61.3% of TOEFL essays by non-native English writers as machine-generated. It is not a bug in one product. It is what happens when you measure vocabulary range and call it authorship.

When we measured our own feature set, that family scored backwards — it rated dense human academic prose as more machine-like than actual machine output. We removed it from scoring entirely. Our detector now reads sentence structure and rhythm instead: length variation, paragraph consistency, how often consecutive sentences resemble one another.

Where Ours Is Weakest

An accuracy claim without a weakness section is marketing. Ours:

  • Short German text is our weakest case. On German abstracts we catch only 65% of machine writing. So German documents under roughly 400 words never receive a "high probability" verdict from us, no matter the score. The measurement does not support that claim, so we do not make it.
  • Our machine-written test sets are small — 20 to 25 documents per language, from a limited number of models. The human side is large and its label is beyond dispute; the machine side is thinner.
  • Our German long-form test set is encyclopedia prose, not thesis prose. It is formal, German, long and verifiably human, but it is not identical to how a student writes.
  • Heavily edited machine text is harder. If someone rewrites AI output substantially, it becomes partly their own writing, and no detector resolves that cleanly.

What a Score Actually Means for You

If a detector has flagged your work and you wrote it yourself, the useful facts are these.

A score is not evidence of anything. It is a statistical observation about text. Ours included. We say so next to every verdict we produce, because a student handed a number by a supervisor deserves to know how often that number is wrong.

Ask which population the tool was tested on. If a detector cannot tell you its false positive rate on writing like yours, its confident percentage is not a measurement.

Keep your version history. The strongest defense is not a counter-score, it is a timeline showing your thesis grew over weeks. Google Docs and OneDrive keep this automatically. We wrote a full walkthrough in what to do when an AI detector flags your thesis, including the email to send your supervisor.

You can run our free AI checker on the first 1,500 words of your document. It reports the verdict, the false positive rate for that verdict, and which sentences drove it — so you can see the reasoning rather than just the number.

Frequently Asked Questions

How accurate are AI detectors in general?

It depends entirely on who wrote the text. On native-speaker English, leading detectors perform reasonably. On non-native English writing, a Stanford study (Liang et al., 2023, published in Patterns) found seven major detectors misclassified 61.3% of TOEFL essays as AI-generated. Accuracy claims that do not state which population they were measured on are close to meaningless.

What is a false positive rate and why does it matter more than accuracy?

The false positive rate is how often a tool calls human writing machine-written. It matters more than headline accuracy because it is the error that damages a student. A detector can be 95% accurate overall and still wrongly flag one honest thesis in ten, and that one student is the only one who cares.

Can an AI detector prove I used ChatGPT?

No. No detector can prove authorship, ours included. These tools measure statistical properties of text and report a probability. In Germany and most of the EU, examination boards require a formal procedure and concrete evidence beyond a single tool score.

Why do you publish your error rate when competitors do not?

Because a score without an error rate is not a measurement, it is a claim. If we tell you a passage reads as machine-written, you deserve to know how often we say that about writing we know was human. We also publish where our detector is weakest, which is the part that is genuinely useful when you have to defend your work.

Check your writing for AI text — free

The first 1,500 words are free, with no sign-up. Every verdict shows how often it is wrong about verified human writing — a figure no other detector publishes.

Need a researched draft to work from instead? From €4.99, anonymous, one-time payment — see pricing.