Several exhibits elsewhere on this site rest on AI detectors. So when a study puts the detectors themselves on trial, I sit up. (Court reporters don’t sit up, officially. I did anyway.)
The authors are Liang, Yuksekgonul, Mao, Wu and Zou, all of Stanford University. The paper is “GPT detectors are biased against non-native English writers”, in Patterns, 2023. A caution: it’s an Opinion article, and the full method sits in the arXiv preprint. The numbers here are from the published version.
The method is simple. Take seven widely used GPT detectors. Feed them 91 TOEFL essays written by non-native English speakers, collected from a Chinese forum, and 88 US eighth-grade essays from the Hewlett Foundation ASAP data. Every one of those essays was written by a human. Every single one.
So the detectors should say “human” every time. They didn’t.
They classified the US essays accurately. But they labelled more than half the TOEFL essays “AI-generated”, an average false-positive rate of 61.3%. All seven detectors agreed on 19.8% of the TOEFL essays, and at least one detector flagged 97.8% of them. Read that last one again. Nearly every non-native essay got accused by somebody.
Why? Most detectors rely on text perplexity, which is a measure of how predictable the words are. A writer with a more limited range of expression produces more predictable text, and the detector reads that as machine. The penalty lands on the people with the smaller toolkit, which is not what anyone wants a detector to do.
Then the authors tried a fix. When ChatGPT was used to enrich the TOEFL essays’ vocabulary, the average false-positive rate fell from 61.3% to 11.6%. Same ideas, richer words, far fewer accusations. And separately, asking ChatGPT to rewrite its own essays “employing literary language” let them slip past the detectors. So the instrument convicts plain human writing and acquits dressed-up machine writing. Not a great combination.
What does this mean for the court? Plain, predictable human writing can look like a machine to a detector. Any exhibit that leans on a detector has to be read with that in mind.
Caveats, though. These were 2023 detectors and small samples. Newer detectors claim to have fixed this problem, and I can’t tell you whether they have. And the study says nothing about how much of the web is machine-written. It only says how far a detector can be trusted.
That’s a narrow finding, but it’s a sharp one. The defence didn’t need to prove the web is human. It only needed to show the witness is unreliable, and it did.
Verdict: for the defence.