Why AI text detectors are unreliable
Updated September 2026
Text detectors measure statistical regularity, not authorship. They cannot distinguish a person who writes plainly from a model, they misjudge non-native English writers more often, and light editing defeats them — so their output should never be used as evidence of misconduct.
Every few months a story surfaces about a student failed over a detector score, or a writer accused of using AI for work they laboured over. These are not freak accidents. They are the predictable consequence of treating a statistical estimate as a finding of fact.
This page explains what these tools actually measure, why the errors fall unevenly on particular groups of people, and what to do instead if you are responsible for making a judgement.
What they actually measure
AI text detectors do not recognise authorship. They measure how statistically predictable a passage is, usually through two proxies. Perplexity asks how surprised a language model is by each next word — generated text tends to be less surprising, because models select likely continuations. Burstiness measures variation in sentence length and structure — human writing tends to lurch between long and short sentences, while generated prose trends uniform.
Both are real, measurable properties. Neither is authorship. A person who writes in clear, even, formal sentences produces low perplexity and low burstiness, and the detector cannot tell them apart from a model. That is not a bug in a particular product; it is what the measurement is.
The errors are not evenly distributed
This is the part that matters most and gets discussed least. Writing in a second language tends to use more common vocabulary and simpler sentence construction — precisely the features detectors read as machine-generated. Research published in 2023 found detectors classified the large majority of TOEFL essays by non-native English speakers as AI-generated, while classifying essays by native speakers correctly.
The same skew catches anyone writing in a plain register: technical documentation, legal drafting, non-native academic writing, and people who have been coached to write simply and consistently. A false positive is not a random misfire. It lands disproportionately on people who already face more scrutiny and have less standing to contest it.
Editing defeats them, which inverts the incentive
Paraphrasing generated text, or editing it by hand, disrupts exactly the patterns detectors rely on. A few minutes of rewriting typically drops a score substantially, and purpose-built "humaniser" tools automate it.
Consider what that means in practice. Someone who generates an essay and edits it carefully is unlikely to be flagged. Someone who writes honestly in a plain, consistent style may well be. The tool systematically fails against the behaviour it is meant to catch, while penalising a writing style that is not misconduct at all.
Vendor accuracy claims do not survive contact with reality
Detectors are evaluated on test sets that resemble their training data: text from specific models, of a certain length, unedited. Real submissions are shorter, mixed, edited, and produced by models released after the detector was trained.
OpenAI is the useful data point here. It launched its own AI text classifier in early 2023 and withdrew it within six months, citing a low rate of accuracy. The organisation with the deepest possible knowledge of how its models write concluded it could not reliably detect their output.
What to do instead
If you are an educator or a reviewer, treat a score as a prompt to start a conversation, never as its conclusion. Ask the person about their process — what sources they used, why they structured an argument a certain way, what they discarded. Someone who did the work can discuss it; someone who did not, generally cannot.
Where it is practical, design the assessment so the question does not arise: drafts and version history, in-class or supervised components, oral defence, work that requires specific personal or local context. These are more effort than running a checker, and they are the only approaches that actually hold.
And if you are the person who was flagged: this is common, it is not a statement about the quality of your writing, and the reasonable response is to ask what specific evidence beyond the score is being relied upon. Usually there is none.
Try it on a file
Runs every check described above, in that order, and shows you which one produced the answer.
Drag a file here, or
JPEG, PNG, WebP, AVIF, PDF or .txt · up to 10 MB · you can also paste a screenshot
By uploading you agree to our Terms and Privacy Policy. Your file, IP address, approximate location and device details are stored.
Frequently asked questions
How accurate are AI text detectors?
Far less accurate than marketing suggests, and least accurate exactly where it matters — on short, edited or mixed text. Accuracy also degrades on models released after the detector was trained, with no indication to the user that it has.
Why do detectors flag non-native English writers?
Because writing in a second language typically uses more common vocabulary and simpler sentence structures, which are the same features detectors treat as signs of machine generation. The tools measure statistical regularity, and that correlates with second-language writing as strongly as it does with generated text.
Can a school rely on a detector score?
No. Several universities have withdrawn detector-based processes after false positives, and major vendors carry explicit warnings against using scores as sole evidence. A score is not a finding.
Is detecting AI images more reliable than text?
Considerably, when the file retains its original metadata. Images can carry Content Credentials or embedded generation parameters — evidence the file provides about itself. Text has no equivalent, which is why text detection rests entirely on statistical inference.