---
title: "Computer scientists took up the machine's words first. Mathematicians, much less."
url: https://deadwritingtheory.com/cases/the-scientists-got-there-first/
summary: "By September 2024, 22.5% of computer science abstract sentences were LLM-modified, against 7.7% in mathematics. The machine is in science, but not evenly."
published: 2026-10-09
updated: 2026-10-09
author: "Jack Stovell"
publisher: "Adapt Progress Evolve Limited"
language: en-GB
---

# Computer scientists took up the machine's words first. Mathematicians, much less.

By September 2024, 22.5% of computer science abstract sentences were LLM-modified, against 7.7% in mathematics. The machine is in science, but not evenly.

**Case No. DWT-012: The scientists got there first.** Verdict for the prosecution. Filed 9 October 2026.

- **The question:** How much of the text in scientific papers was modified by large language models between January 2020 and September 2024, and in which fields?
- **The sample:** 1,121,912 papers from January 2020 to September 2024 (861,253 from arXiv, 205,094 from bioRxiv, 55,565 from Nature portfolio journals); abstracts and introductions, estimated with a population-level word-frequency model
- **The study:** Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D. Manning, James Zou (Stanford University and others), "Quantifying large language model usage in scientific papers", Nature Human Behaviour 9(12), 2025. https://doi.org/10.1038/s41562-025-02273-8
- **Whose study:** Third-party study. Not our study. Entered into evidence from Weixin Liang, Yaohui Zhang, Zhengxuan Wu, Haley Lepp, Wenlong Ji, Xuandong Zhao, Hancheng Cao, Sheng Liu, Siyu He, Zhi Huang, Diyi Yang, Christopher Potts, Christopher D. Manning, James Zou, Stanford University and others.
- **Peer review:** Peer reviewed

## Findings of fact

1. By September 2024 the estimated share of LLM-modified sentences in computer science was 22.5% for abstracts and 19.6% for introductions.
2. Mathematics reached 7.7% for abstracts and 4.1% for introductions; the Nature portfolio 8.9% and 9.4%.
3. In November 2022, before ChatGPT, the computer science estimate was 2.4%, consistent with the method's false-positive rate.
4. Shorter computer science papers (below 5,000 words) reached 22.0% for abstracts, against 19.3% for longer ones.

## Caveats for the jury

- It measures adoption (prevalence), not whether papers became more alike.
- The authors say the method overestimates LLM use at the low end and underestimates it at the high end; it finds statistical patterns consistent with LLM text, not proven use.
- Abstracts and introductions only, and the links to paper length and crowded fields are correlations.

## The court reporter's account

The court will recall that the same Stanford group appeared in case DWT-003, the one about press releases. Different corpus this time, same people, so I'll try not to look surprised.

The paper is "Quantifying large language model usage in scientific papers", in Nature Human Behaviour, 2025, and it's peer reviewed. The authors are Liang, Zhang, Wu, Lepp, Ji, Zhao, Cao, Liu, He, Huang, Yang, Potts, Manning and Zou, at Stanford University and others. That's a lot of names for one docket entry.

The method matters, so listen. It isn't a per-paper detector. It's a population-level statistical model of word frequencies, which means it asks how much of a whole body of text looks machine-modified, not whether paper X was written by a bot. The corpus: 1,121,912 papers from January 2020 to September 2024. Of those, 861,253 came from arXiv, 205,094 from bioRxiv and 55,565 from Nature portfolio journals. Abstracts and introductions only.

Now the numbers. By September 2024 the estimated share of LLM-modified sentences in computer science was 22.5% for abstracts and 19.6% for introductions. Mathematics reached 7.7% (abstracts) and 4.1% (introductions). The Nature portfolio came in at 8.9% and 9.4%. For comparison, in November 2022, before ChatGPT, the computer science estimate was 2.4%, which is consistent with the method's false-positive rate. So that's the baseline noise.

There's more. The share was higher in papers whose first authors post preprints often, in more crowded research areas, and in shorter papers. Below 5,000 words, 22.0% of abstract sentences; above, 19.3%.

The point for the court: the theory says "everywhere". In science the machine's words did spread, but unevenly: 22.5% of computer science abstracts against 7.7% in mathematics. Same tool, very different uptake. That's not nothing, and it isn't uniform either.

The caveats, and they're real. This measures prevalence, meaning how much text involved AI, not convergence, meaning whether writing became more alike. Different questions. The authors say the method overestimates at the low end and underestimates at the high end. It identifies statistical patterns consistent with LLM text, not proven use. And the associations are correlations.

Even so, a peer-reviewed measurement of 1,121,912 papers is hard to wave away.

Verdict: for the prosecution.

## The exhibits entered in this case

### For the prosecution

- **Exhibit P-15**: Computer science abstracts took on the machine's words. **22.5%** of computer science abstract sentences LLM-modified, September 2024. 2.4% in November 2022, before ChatGPT. Introductions 19.6%. 1,121,912 papers; a population-level estimate, not a detector.

### For the defence

- **Exhibit D-16**: Mathematics lagged far behind. **7.7%** of mathematics abstract sentences LLM-modified, September 2024. 4.1% for introductions. The Nature portfolio reached 8.9% and 9.4%.

## The finding

By September 2024, 22.5% of computer science abstract sentences were LLM-modified, against 7.7% in mathematics. The machine is in science, but not evenly.

## Related

- [The machine is in the room. In a lot of rooms.](https://deadwritingtheory.com/cases/a-quarter-of-the-press-release/)
- [The machine's words spread through medicine, measured without a detector.](https://deadwritingtheory.com/cases/the-delve-study-done-properly/)

## Sources

- [Liang et al., "Quantifying large language model usage in scientific papers", Nature Human Behaviour 9(12), 2025 (the original study)](https://doi.org/10.1038/s41562-025-02273-8)
- [Liang et al., the earlier preprint on arXiv](https://arxiv.org/abs/2404.01268)

## Questions people ask

### How much of scientific writing is done with AI?

It depends on the field. Liang et al. (Stanford, Nature Human Behaviour, 2025) estimated that by September 2024, 22.5% of sentences in computer science abstracts were LLM-modified, against 7.7% in mathematics and 8.9% in Nature portfolio journals.

### Is this the same study as case DWT-003?

Same Stanford group and method, different corpus. Case DWT-003 looked at press releases, complaints, job ads and UN releases; this one looks at 1,121,912 scientific papers.

### Does it show that science writing is converging?

No. It measures how much text was LLM-modified, not whether papers became more alike.

### Why is the verdict for the prosecution, if mathematics lags?

Because the estimates rose from their pre-ChatGPT levels: from 2.4% to 22.5% in computer science abstracts, and from 2.5% to 7.7% even in mathematics. The field gap is a point against "everywhere", entered as a defence exhibit.
