The court will recall that the same Stanford group appeared in case DWT-003, the one about press releases. Different corpus this time, same people, so I’ll try not to look surprised.
The paper is “Quantifying large language model usage in scientific papers”, in Nature Human Behaviour, 2025, and it’s peer reviewed. The authors are Liang, Zhang, Wu, Lepp, Ji, Zhao, Cao, Liu, He, Huang, Yang, Potts, Manning and Zou, at Stanford University and others. That’s a lot of names for one docket entry.
The method matters, so listen. It isn’t a per-paper detector. It’s a population-level statistical model of word frequencies, which means it asks how much of a whole body of text looks machine-modified, not whether paper X was written by a bot. The corpus: 1,121,912 papers from January 2020 to September 2024. Of those, 861,253 came from arXiv, 205,094 from bioRxiv and 55,565 from Nature portfolio journals. Abstracts and introductions only.
Now the numbers. By September 2024 the estimated share of LLM-modified sentences in computer science was 22.5% for abstracts and 19.6% for introductions. Mathematics reached 7.7% (abstracts) and 4.1% (introductions). The Nature portfolio came in at 8.9% and 9.4%. For comparison, in November 2022, before ChatGPT, the computer science estimate was 2.4%, which is consistent with the method’s false-positive rate. So that’s the baseline noise.
There’s more. The share was higher in papers whose first authors post preprints often, in more crowded research areas, and in shorter papers. Below 5,000 words, 22.0% of abstract sentences; above, 19.3%.
The point for the court: the theory says “everywhere”. In science the machine’s words did spread, but unevenly: 22.5% of computer science abstracts against 7.7% in mathematics. Same tool, very different uptake. That’s not nothing, and it isn’t uniform either.
The caveats, and they’re real. This measures prevalence, meaning how much text involved AI, not convergence, meaning whether writing became more alike. Different questions. The authors say the method overestimates at the low end and underestimates at the high end. It identifies statistical patterns consistent with LLM text, not proven use. And the associations are correlations.
Even so, a peer-reviewed measurement of 1,121,912 papers is hard to wave away.
Verdict: for the prosecution.