---
title: "The detectors flagged human writers as machines."
url: https://deadwritingtheory.com/cases/the-detector-that-convicts-the-innocent/
summary: "Seven AI detectors labelled more than half of 91 human-written TOEFL essays as AI (average false-positive rate 61.3%). Any exhibit that rests on a detector needs that caveat."
published: 2026-10-09
updated: 2026-10-09
author: "Jack Stovell"
publisher: "Adapt Progress Evolve Limited"
language: en-GB
---

# The detectors flagged human writers as machines.

Seven AI detectors labelled more than half of 91 human-written TOEFL essays as AI (average false-positive rate 61.3%). Any exhibit that rests on a detector needs that caveat.

**Case No. DWT-009: The detector that convicts the innocent.** Verdict for the defence. Filed 9 October 2026.

- **The question:** Do widely used GPT detectors misclassify writing by non-native English speakers as AI-generated?
- **The sample:** Seven widely used GPT detectors, run in 2023 on 91 human-written TOEFL essays from a Chinese forum and 88 US eighth-grade essays from the Hewlett Foundation's ASAP dataset
- **The study:** Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou (Stanford University), "GPT detectors are biased against non-native English writers", Patterns 4(7), 2023. https://doi.org/10.1016/j.patter.2023.100779
- **Whose study:** Third-party study. Not our study. Entered into evidence from Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou, Stanford University.
- **Peer review:** Preprint, not yet peer reviewed

## Findings of fact

1. The detectors labelled more than half the TOEFL essays "AI-generated": an average false-positive rate of 61.3%. The US essays were classified accurately.
2. All seven detectors agreed on 19.8% of the TOEFL essays, and at least one detector flagged 97.8% of them.
3. When ChatGPT enriched the TOEFL essays' vocabulary, the average false-positive rate fell from 61.3% to 11.6%.
4. Asking ChatGPT to rewrite its own essays "employing literary language" let them slip past the detectors.

## Caveats for the jury

- Published as an Opinion article in Patterns; the full method is in the arXiv preprint, which reports slightly different figures (for example 61.22% and 11.77%). This case quotes the published version.
- 2023 detectors and small samples (91 and 88 essays). Newer detectors claim to have fixed this bias; this study cannot say whether they have.
- It tests how far a detector can be trusted, not how much of the web is machine-written.

## The court reporter's account

Several exhibits elsewhere on this site rest on AI detectors. So when a study puts the detectors themselves on trial, I sit up. (Court reporters don't sit up, officially. I did anyway.)

The authors are Liang, Yuksekgonul, Mao, Wu and Zou, all of Stanford University. The paper is "GPT detectors are biased against non-native English writers", in Patterns, 2023. A caution: it's an Opinion article, and the full method sits in the arXiv preprint. The numbers here are from the published version.

The method is simple. Take seven widely used GPT detectors. Feed them 91 TOEFL essays written by non-native English speakers, collected from a Chinese forum, and 88 US eighth-grade essays from the Hewlett Foundation ASAP data. Every one of those essays was written by a human. Every single one.

So the detectors should say "human" every time. They didn't.

They classified the US essays accurately. But they labelled more than half the TOEFL essays "AI-generated", an average false-positive rate of 61.3%. All seven detectors agreed on 19.8% of the TOEFL essays, and at least one detector flagged 97.8% of them. Read that last one again. Nearly every non-native essay got accused by somebody.

Why? Most detectors rely on text perplexity, which is a measure of how predictable the words are. A writer with a more limited range of expression produces more predictable text, and the detector reads that as machine. The penalty lands on the people with the smaller toolkit, which is not what anyone wants a detector to do.

Then the authors tried a fix. When ChatGPT was used to enrich the TOEFL essays' vocabulary, the average false-positive rate fell from 61.3% to 11.6%. Same ideas, richer words, far fewer accusations. And separately, asking ChatGPT to rewrite its own essays "employing literary language" let them slip past the detectors. So the instrument convicts plain human writing and acquits dressed-up machine writing. Not a great combination.

What does this mean for the court? Plain, predictable human writing can look like a machine to a detector. Any exhibit that leans on a detector has to be read with that in mind.

Caveats, though. These were 2023 detectors and small samples. Newer detectors claim to have fixed this problem, and I can't tell you whether they have. And the study says nothing about how much of the web is machine-written. It only says how far a detector can be trusted.

That's a narrow finding, but it's a sharp one. The defence didn't need to prove the web is human. It only needed to show the witness is unreliable, and it did.

Verdict: for the defence.

## The exhibits entered in this case

### For the prosecution

- No exhibit entered for the prosecution in this case.

### For the defence

- **Exhibit D-12**: Detectors called most non-native writers' essays AI. **61.3%** average false-positive rate on 91 human TOEFL essays. Seven widely used detectors, 2023. At least one flagged 97.8% of the essays; all seven agreed on 19.8%. US eighth-grade essays were classified accurately.
- **Exhibit D-13**: Richer wording cleared the same essays. **61.3% → 11.6%** average false-positive rate after ChatGPT enriched the vocabulary. Detectors read predictable wording as machine-made, so plain human writing is penalised.

## The finding

Seven AI detectors labelled more than half of 91 human-written TOEFL essays as AI (average false-positive rate 61.3%). Any exhibit that rests on a detector needs that caveat.

## Related

- [More sites with little human input. More of them selling.](https://deadwritingtheory.com/cases/the-websites-nobody-wrote/)
- [Half the articles. Then the rise stopped.](https://deadwritingtheory.com/cases/half-the-articles-and-a-plateau/)
- [Three quarters of new pages had some AI in them. 2.5% were all AI.](https://deadwritingtheory.com/cases/three-quarters-touched/)

## Sources

- [Liang et al., "GPT detectors are biased against non-native English writers", Patterns 4(7), 2023 (the original study)](https://doi.org/10.1016/j.patter.2023.100779)
- [Liang et al., the preprint with the full method on arXiv](https://arxiv.org/abs/2304.02819)
- [Liang et al., the published article on PubMed Central (open access)](https://pmc.ncbi.nlm.nih.gov/articles/PMC10382961/)

## Questions people ask

### Are AI detectors reliable?

Not for everyone. In Liang et al. (Stanford, Patterns, 2023), seven widely used detectors labelled more than half of 91 human-written TOEFL essays as AI-generated, an average false-positive rate of 61.3%, while classifying US eighth-grade essays accurately.

### Why do detectors flag non-native writers?

Most rely on text perplexity, how predictable the words are. Writers with a more limited range of expression produce more predictable text. When ChatGPT enriched the same essays' vocabulary, the false-positive rate fell from 61.3% to 11.6%.

### What does this mean for estimates of AI text on the web?

Counts built on detectors carry a false-positive caveat. Several cases on this docket rely on detectors, and each says so in its caveats.

### Is this study out of date?

It tested 2023 detectors. Newer detectors claim to have fixed the bias; this study cannot tell the court whether they have.
