Skip to the content

The Court of Public Prose

In re Dead Writing Theory

Case No.
DWT-012
Matter
The scientists got there first
Filed
Verdict
Verdict for the prosecution

Computer scientists took up the machine's words first. Mathematicians, much less.

For the prosecution every exhibit in this case · P-15For the defence every exhibit in this case · D-16
The question
How much of the text in scientific papers was modified by large language models between January 2020 and September 2024, and in which fields?
The sample
1,121,912 papers from January 2020 to September 2024 (861,253 from arXiv, 205,094 from bioRxiv, 55,565 from Nature portfolio journals); abstracts and introductions, estimated with a population-level word-frequency model
Filed
The study
Quantifying large language model usage in scientific papers Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D. Manning, James Zou. Stanford University and others. Nature Human Behaviour 9(12), 2025.

§ IFindings of fact

What the study found.

  1. By September 2024 the estimated share of LLM-modified sentences in computer science was 22.5% for abstracts and 19.6% for introductions.
  2. Mathematics reached 7.7% for abstracts and 4.1% for introductions; the Nature portfolio 8.9% and 9.4%.
  3. In November 2022, before ChatGPT, the computer science estimate was 2.4%, consistent with the method's false-positive rate.
  4. Shorter computer science papers (below 5,000 words) reached 22.0% for abstracts, against 19.3% for longer ones.

§ IICaveats for the jury

What this study cannot tell the court.

  • It measures adoption (prevalence), not whether papers became more alike.
  • The authors say the method overestimates LLM use at the low end and underestimates it at the high end; it finds statistical patterns consistent with LLM text, not proven use.
  • Abstracts and introductions only, and the links to paper length and crowded fields are correlations.

§ IVThe exhibits entered in this case

Prosecution v. Defence

Any tally counts the evidence items below; it is not a measure of truth.

Evidence for the theory

For the prosecution

1 exhibit
  1. Exhibit P-15Case No. DWT-012

    22.5%

    of computer science abstract sentences LLM-modified, September 2024

    Computer science abstracts took on the machine's words

    2.4% in November 2022, before ChatGPT. Introductions 19.6%. 1,121,912 papers; a population-level estimate, not a detector.

    Entered for the prosecution · Liang et al., Nature Human Behaviour 9(12), 2025 · third-party study

Evidence against the theory

For the defence

1 exhibit
  1. Exhibit D-16Case No. DWT-012

    7.7%

    of mathematics abstract sentences LLM-modified, September 2024

    Mathematics lagged far behind

    4.1% for introductions. The Nature portfolio reached 8.9% and 9.4%.

    Entered for the defence · Liang et al., Nature Human Behaviour 9(12), 2025 · third-party study

§ VThe finding

The finding

By September 2024, 22.5% of computer science abstract sentences were LLM-modified, against 7.7% in mathematics. The machine is in science, but not evenly.
Verdict for the prosecutionCase No. DWT-012 · The scientists got there first
Read the original studyLiang et al., Nature Human Behaviour 9(12), 2025. Not our study.

§ VIQuestions for the court

Asked and answered.

How much of scientific writing is done with AI?

It depends on the field. Liang et al. (Stanford, Nature Human Behaviour, 2025) estimated that by September 2024, 22.5% of sentences in computer science abstracts were LLM-modified, against 7.7% in mathematics and 8.9% in Nature portfolio journals.

Is this the same study as case DWT-003?

Same Stanford group and method, different corpus. Case DWT-003 looked at press releases, complaints, job ads and UN releases; this one looks at 1,121,912 scientific papers.

Does it show that science writing is converging?

No. It measures how much text was LLM-modified, not whether papers became more alike.

Why is the verdict for the prosecution, if mathematics lags?

Because the estimates rose from their pre-ChatGPT levels: from 2.4% to 22.5% in computer science abstracts, and from 2.5% to 7.7% even in mathematics. The field gap is a point against "everywhere", entered as a defence exhibit.