Skip to the content

The Court of Public Prose

In re Dead Writing Theory

Case No.
DWT-009
Matter
The detector that convicts the innocent
Filed
Verdict
Verdict for the defence

The detectors flagged human writers as machines.

For the defence every exhibit in this case · D-12 to D-13
The question
Do widely used GPT detectors misclassify writing by non-native English speakers as AI-generated?
The sample
Seven widely used GPT detectors, run in 2023 on 91 human-written TOEFL essays from a Chinese forum and 88 US eighth-grade essays from the Hewlett Foundation's ASAP dataset
Filed
The study
GPT detectors are biased against non-native English writers Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou. Stanford University. Patterns 4(7), 2023.

§ IFindings of fact

What the study found.

  1. The detectors labelled more than half the TOEFL essays "AI-generated": an average false-positive rate of 61.3%. The US essays were classified accurately.
  2. All seven detectors agreed on 19.8% of the TOEFL essays, and at least one detector flagged 97.8% of them.
  3. When ChatGPT enriched the TOEFL essays' vocabulary, the average false-positive rate fell from 61.3% to 11.6%.
  4. Asking ChatGPT to rewrite its own essays "employing literary language" let them slip past the detectors.

§ IICaveats for the jury

What this study cannot tell the court.

  • Published as an Opinion article in Patterns; the full method is in the arXiv preprint, which reports slightly different figures (for example 61.22% and 11.77%). This case quotes the published version.
  • 2023 detectors and small samples (91 and 88 essays). Newer detectors claim to have fixed this bias; this study cannot say whether they have.
  • It tests how far a detector can be trusted, not how much of the web is machine-written.

§ IVThe exhibits entered in this case

Prosecution v. Defence

Any tally counts the evidence items below; it is not a measure of truth.

Evidence for the theory

For the prosecution

0 exhibits

No exhibit entered for the prosecution in this case.

Evidence against the theory

For the defence

2 exhibits
  1. Exhibit D-12Case No. DWT-009

    61.3%

    average false-positive rate on 91 human TOEFL essays

    Detectors called most non-native writers' essays AI

    Seven widely used detectors, 2023. At least one flagged 97.8% of the essays; all seven agreed on 19.8%. US eighth-grade essays were classified accurately.

    Entered for the defence · Liang et al., Patterns 4(7), 2023 · third-party study

  2. Exhibit D-13Case No. DWT-009

    61.3% → 11.6%

    average false-positive rate after ChatGPT enriched the vocabulary

    Richer wording cleared the same essays

    Detectors read predictable wording as machine-made, so plain human writing is penalised.

    Entered for the defence · Liang et al., Patterns 4(7), 2023 · third-party study

§ VThe finding

The finding

Seven AI detectors labelled more than half of 91 human-written TOEFL essays as AI (average false-positive rate 61.3%). Any exhibit that rests on a detector needs that caveat.
Verdict for the defenceCase No. DWT-009 · The detector that convicts the innocent
Read the original studyLiang et al., Patterns 4(7), 2023. Not our study.

§ VIQuestions for the court

Asked and answered.

Are AI detectors reliable?

Not for everyone. In Liang et al. (Stanford, Patterns, 2023), seven widely used detectors labelled more than half of 91 human-written TOEFL essays as AI-generated, an average false-positive rate of 61.3%, while classifying US eighth-grade essays accurately.

Why do detectors flag non-native writers?

Most rely on text perplexity, how predictable the words are. Writers with a more limited range of expression produce more predictable text. When ChatGPT enriched the same essays' vocabulary, the false-positive rate fell from 61.3% to 11.6%.

What does this mean for estimates of AI text on the web?

Counts built on detectors carry a false-positive caveat. Several cases on this docket rely on detectors, and each says so in its caveats.

Is this study out of date?

It tested 2023 detectors. Newer detectors claim to have fixed the bias; this study cannot tell the court whether they have.