← Back to blog

Why AI Detectors Flag Non-Native English Writers More Often

|10 min read

Check your writing for AI text — free

The first 1,500 words are free, with no sign-up. Every verdict shows how often it is wrong about verified human writing — a figure no other detector publishes.

We are building a writing workspace: your Word or LaTeX document, your PDFs beside it, and an assistant that can only cite what is actually in them — see it and get notified.

The short answer

AI detectors flag second-language writers more often because they do not detect AI — they measure predictability. Writing in a language that is not your first tends to use a smaller vocabulary and safer, more conventional constructions, which makes the next word easier to guess, which is precisely the property these tools score. It is the best-documented bias in the category: a 2023 Stanford study published in Patterns found seven detectors misclassified a large share of TOEFL essays by non-native writers as AI-generated while classifying essays by US school students almost perfectly. The flag is a statement about the statistical shape of your prose, not about how it was produced.

If this has already happened to you, the two things that help are naming the effect as a documented phenomenon rather than a personal plea, and producing evidence of your writing process. Neither requires you to change how you write.

What do AI detectors actually measure?

Almost every detector in use is built on the same idea. Take a language model, feed it your text, and at every position ask: how surprised was the model by the word that actually came next? Aggregate that across the document and you get a measure of how predictable the writing is, along with a measure of how much that predictability varies from sentence to sentence.

The reason this works at all is that generated text is produced by a process that selects likely continuations. Given a choice between the ordinary phrasing and the unexpected one, it takes the ordinary phrasing more consistently than a person does. Across a few thousand words that produces prose which is smoother, more even, and less surprising than most human writing.

Two things follow, and they explain the whole rest of this post. First, a detector never observes the act of writing — only the finished text. It cannot distinguish a document that was generated from a document that merely resembles one. Second, “resembles one” is a property some humans have and others do not. It is not randomly distributed across the population. It concentrates.

Why does second-language writing score as predictable?

Because writing in a second language and generating text are both, in different ways, exercises in staying close to the safe option.

  • Smaller active vocabulary. Not smaller knowledge — smaller confident range. You reach for the word you are certain of rather than the one that is slightly better, which is usually the more common word, which is the more predictable one.
  • Learned constructions over improvised ones. Second-language writers rely on patterns they have been taught are correct: standard subordinate clause structures, textbook transitions, conventional sentence openings. These are correct. They are also exactly what a model predicts.
  • Fewer idioms, less register-play, no wordplay. The features that make first-language prose unpredictable are the ones acquired last, and often deliberately avoided in formal academic writing anyway.
  • More even sentence rhythm. A confident first-language writer varies sentence length instinctively, including fragments and long trailing clauses. Careful second-language writing tends toward a steadier, more uniform rhythm — which reads as low variation, and low variation is itself a signal these tools use.
  • Heavier editing. Non-native writers use grammar and style tools more, and those tools smooth away precisely the irregularities that read as human.

The uncomfortable conclusion is that these are all marks of doing it well. A student who has drilled academic English until their prose is clean, correct and conventional has produced exactly the text most likely to be flagged. The bias does not punish poor English. It punishes careful English.

How well documented is this bias?

Better than almost anything else in this field. The reference point is “GPT detectors are biased against non-native English writers” by Weixin Liang and colleagues at Stanford, published in Patterns in 2023. The design is simple enough to describe in a sentence: run seven widely used detectors over essays written by non-native English speakers for the TOEFL exam, and over essays written by US school students, and compare. The TOEFL essays were misclassified as AI-generated at a strikingly high rate; the US student essays were classified close to perfectly.

The study also did something more revealing than the headline. When the researchers took the TOEFL essays and enriched the vocabulary, the false positives largely disappeared — which is the demonstration that the detectors were keying on linguistic sophistication rather than on anything to do with machine authorship.

Quote it accurately, including its limits, because a claim stated precisely survives challenge and an inflated one does not. It tested particular detectors as they existed in 2023, on exam essays rather than theses, in English. Detectors have changed since. What has not changed is the mechanism, because predictability is still what the category measures, and that is why the finding continues to be cited.

What makes it worse in a thesis specifically?

A thesis stacks several of these effects at once, which is why the accusation cluster is so concentrated among international students:

  • Academic style is a predictability requirement. Formal, hedged, impersonal, consistently signposted — universities teach the register that scores worst.
  • Methods and theory chapters are formulaic by necessity. There are only so many ways to describe a standard procedure, and originality of phrasing is not wanted there.
  • Terminology must be repeated exactly. Varying your key terms for style is a defect in a thesis. Repetition raises predictability.
  • Proofreading is expected. Many programmes explicitly encourage language support, then run a detector over the smoothed result.

A carefully proofread methods chapter written in a second language is close to the worst-case input for a detector, and there is nothing dishonest anywhere in that sentence.

Why “write less predictably” is not the answer

The advice you will find elsewhere is to vary your sentence length, insert unusual words, and break up your rhythm. Do not take it, for three reasons.

It costs you marks. Clarity and precision are assessed; deliberately roughened prose is worse prose, and your grade is a certainty while the flag is a probability. It does not reliably work, because thresholds differ by tool and by version and you have no feedback loop. And it changes your position: if it later emerges that you edited your thesis specifically to influence a detector, you have converted a defensible situation into one that requires explaining.

There is a narrow, legitimate version of this. If a check shows one chapter reading as markedly more machine-like than the rest, that is worth looking at — not to disguise it, but because it is often genuinely your weakest, most padded, most formulaic writing, and improving it improves the thesis. That is editing for quality, and a supervisor would give you the same note.

What does an honest error rate look like?

Here is the practical consequence of everything above: a detector score is uninterpretable without a false-positive rate for your conditions. A tool telling you 80% could mean it is confident, or it could mean it hands out 80% to a quarter of all human theses. Without the error rate you cannot tell those apart, and most vendors do not publish one broken down by language and length.

Ours, measured on documents published before ChatGPT existed, and re-derivable from the scripts behind our accuracy page: roughly 3 false positives in 100 for English over 400 words, 2 in 100 for English under 400, and 4 in 100 for German over 400 words.

The limit belongs next to the number, so: those English figures are measured on English-language human documents, and we have not separately measured English written by non-native speakers. We would expect our rate to be worse on that population than on our corpus, for exactly the reasons this post describes. Anyone claiming otherwise about their own tool should be asked which corpus they measured it on.

When should a detector refuse to answer?

When no threshold is defensible. That is not a hypothetical for us: on German text under 400 words we issue no high-probability verdict at all. At that length, every cutoff strong enough to catch generated text reliably costs somewhere around 7 false positives in 100 — one innocent student in fifteen — and there is no version of this business in which that is an acceptable trade for a confident-looking number.

We are naming it here because it is the same problem this whole post is about, met with the only honest response available. Short text and second-language text are both cases where the signal is genuinely weaker, and the choice is between saying so and shipping a number that will be wrong about people who did nothing. Every detector faces that choice. Most resolve it by displaying the number.

What can you actually do about it?

  1. Keep your version history from day one. Google Docs retains it indefinitely; Word does through OneDrive or SharePoint. A timeline showing a document growing over weeks is the strongest available counter to a score, because the score describes the finished text and says nothing about how it came to exist. Our guide to proving you wrote it yourself covers what to keep.
  2. Record your language support. Which tool, which chapters, what for. Most such use is permitted, and a precise account of permitted use is far stronger than a discovered omission.
  3. Ask what the rules are, in writing, before you submit — and keep the reply. Where the line falls between assistance and authoring is covered in what counts as AI use.
  4. If you are accused, name the effect and cite the study. Ask which tool was used, and what false-positive rate it publishes for your language and text length. Then offer your process evidence. Our walkthrough is what to do when you are accused of using AI.

And if you want to know what a detector sees in your writing before your university runs one, our check is free for the first 1,500 words, with no account and no name attached. It will not prove you wrote your thesis — nothing can, and a tool that claims to is selling you a false comfort. What it gives you is the specific passages that read as statistically machine-like, and the measured rate at which that same verdict is wrong about a human writing in your language.

Frequently Asked Questions

Are AI detectors biased against non-native English speakers?

Yes, and it is the best-documented weakness in the category. A 2023 study by Liang and colleagues at Stanford, published in Patterns, ran seven detectors over TOEFL essays written by non-native speakers and over essays by US school students: the TOEFL essays were misclassified as AI-generated at a strikingly high rate while the US student essays were classified near-perfectly. The mechanism is not prejudice in any human sense — detectors score predictability, and second-language writing is measurably more predictable.

Why does writing in a second language look like AI?

Because both stay close to the safest available option. A second-language writer draws on a smaller working vocabulary and a repertoire of constructions they are confident are correct, which produces text where the next word is easier to guess. That is the exact quantity a detector scores, so competent, careful second-language prose lands in the same statistical region as generated text without sharing anything about how it was produced.

Does this affect German writing too, or only English?

The published evidence concerns English, because that is where the research and most detectors are focused. The mechanism is not language-specific, though, and shorter or non-idiomatic text is harder to judge in any language. We measure our own error rate separately for English and German precisely because assuming one transfers to the other is how a detector ends up confidently wrong.

Should I make my writing less polished so it does not get flagged?

No. Deliberately degrading your writing costs you marks in an assessment where clarity is graded, and it does not reliably change a score anyway. It also puts you in the position of having altered your work in response to a detector, which is difficult to explain later. The thing that protects you is a record of your writing process, not a change to your prose.

What should I say if I am accused and English is not my first language?

Say it explicitly and cite the evidence, because it is a documented and reviewable phenomenon rather than a personal excuse. Ask which tool was used and what false-positive rate it publishes for writers in your situation, and offer your version history and drafts. A committee can dismiss an assertion about your innocence; a named study plus a timeline of incremental drafting is a different kind of argument.

Is there a detector that is fair to non-native speakers?

None of them are free of the effect, ours included, because it follows from what the whole category measures. What differs is whether a tool tells you its error rate for your language and your text length so you can judge the verdict, and whether it declines to give a verdict when no threshold is defensible. Treat a confident percentage with no published error rate as an unfinished measurement.

Check your writing for AI text — free

The first 1,500 words are free, with no sign-up. Every verdict shows how often it is wrong about verified human writing — a figure no other detector publishes.

We are building a writing workspace: your Word or LaTeX document, your PDFs beside it, and an assistant that can only cite what is actually in them — see it and get notified.