---
title: "The machine's words spread through medicine, measured without a detector."
url: https://deadwritingtheory.com/cases/the-delve-study-done-properly/
summary: "Excess words in 15.1 million PubMed abstracts: \"delves\" ran 28.0 times its expected rate in 2024, and at least 13.5% of 2024 abstracts were processed with LLMs."
published: 2026-10-09
updated: 2026-10-09
author: "Jack Stovell"
publisher: "Adapt Progress Evolve Limited"
language: en-GB
---

# The machine's words spread through medicine, measured without a detector.

Excess words in 15.1 million PubMed abstracts: "delves" ran 28.0 times its expected rate in 2024, and at least 13.5% of 2024 abstracts were processed with LLMs.

**Case No. DWT-008: The delve study, done properly.** Verdict for the prosecution. Filed 9 October 2026.

- **The question:** How widely were large language models used to write biomedical research abstracts by 2024, judged from the words that rose above their expected frequency?
- **The sample:** 15.1 million English-language PubMed abstracts from 2010 to 2024; each word's 2024 frequency compared with a projection from its 2021 to 2022 trend
- **The study:** Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, Jan Lause (University of Tübingen and Northwestern University), "Delving into LLM-assisted writing in biomedical publications through excess vocabulary", Science Advances 11(27), 2025. https://doi.org/10.1126/sciadv.adt3813
- **Whose study:** Third-party study. Not our study. Entered into evidence from Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, Jan Lause, University of Tübingen and Northwestern University.
- **Peer review:** Peer reviewed

## Findings of fact

1. "Delves" was used 28.0 times as often as expected in 2024; "underscores" 13.8 times and "showcasing" 10.7 times.
2. 379 style words showed excess use in 2024: 66% were verbs and 14% adjectives. In earlier years the excess words were content words, such as Covid terms.
3. At least 13.5% of 2024 abstracts were processed with LLMs: at least 200,000 papers a year.
4. The lower bound ranged from below 5% to over 40% across fields, countries and journals.
5. The LLM words gave a frequency gap of 0.135 against 0.069 for the four Covid words of 2021: LLM usage in 2024 was at least two times the size of the Covid-related literature in 2021.

## Caveats for the jury

- A lower bound: abstracts that went through an LLM without using any of the marker words are not counted, so the true share is likely higher.
- Abstracts only, a short and formulaic genre.
- It measures uptake of the machine's words, not rhythm, sentence shape or punctuation.

## The court reporter's account

This one I actually like. No detector, no made-up AI texts, which is the part that matters.

The authors are Kobak, González-Márquez, Horvát and Lause, of the University of Tübingen and Northwestern University. The paper is "Delving into LLM-assisted writing in biomedical publications through excess vocabulary", published in Science Advances in 2025, and it was peer reviewed. So far, so respectable.

The method, plainly. They borrowed the idea of "excess mortality". For each word, they predicted its 2024 frequency from the 2021 and 2022 trend, then compared that with what actually happened. The gap is the excess. The sample was 15.1 million English PubMed abstracts from 2010 to 2024.

The headline words first. "Delves" was used 28.0 times as often as expected in 2024. "Underscores" came in at 13.8 times, and "showcasing" at 10.7 times. Altogether, 379 style words showed excess use in 2024; 66% of them were verbs and 14% were adjectives.

Now the bit I found sharpest. In earlier years the excess words were content words. In the Covid years that meant coronavirus, pandemic, lockdown: nouns about the news, which is what you'd expect, because the world changed and the vocabulary followed. In 2024 the excess words were almost all style words. The excess was in how things were said, not in what they were about.

Then the number the court will care about. Combining marker words gives a lower bound: at least 13.5% of 2024 abstracts were processed with LLMs, meaning at least 200,000 papers a year. That lower bound ranged from below 5% to over 40% across fields, countries and journals. The four Covid words of 2021 gave a gap of 0.069, against 0.135 for the LLM words. In the authors' reading, LLM usage in 2024 was at least two times the size of the Covid-related literature in 2021. They call it an "unprecedented impact on scientific writing".

Caveats, and there are real ones. It's a lower bound, so the true share is likely higher. It covers abstracts only, and an abstract is a formulaic genre. And this is the point I'd underline: the study measures prevalence plus word convergence, the uptake of the machine's words. It doesn't touch rhythm or punctuation. Those are different questions, and this study doesn't answer them.

Still, the method is clean and the numbers are large. The defence has no obvious line of attack on how the evidence was gathered, only on how far it reaches.

Verdict: for the prosecution.

## The exhibits entered in this case

### For the prosecution

- **Exhibit P-11**: "Delves" turned up 28 times as often as expected. **28.0×** expected use of "delves" in 2024 PubMed abstracts. Expected from the 2021 to 2022 trend, across 15.1 million abstracts. "Underscores" 13.8×, "showcasing" 10.7×. 379 style words showed excess use in 2024.
- **Exhibit P-12**: At least 13.5% of 2024 biomedical abstracts were processed with LLMs. **13.5%** of 2024 PubMed abstracts, a lower bound. From excess style words, not a detector. Below 5% to over 40% across fields, countries and journals; at least two times the size of the 2021 Covid literature.

### For the defence

- No exhibit entered for the defence in this case.

## The finding

Excess words in 15.1 million PubMed abstracts: "delves" ran 28.0 times its expected rate in 2024, and at least 13.5% of 2024 abstracts were processed with LLMs.

## Related

- [Not an invention. An inheritance.](https://deadwritingtheory.com/cases/where-did-ai-isms-come-from/)
- [People started saying "delve". Then they stopped.](https://deadwritingtheory.com/cases/delve-out-loud/)
- [Computer scientists took up the machine's words first. Mathematicians, much less.](https://deadwritingtheory.com/cases/the-scientists-got-there-first/)

## Sources

- [Kobak et al., "Delving into LLM-assisted writing in biomedical publications through excess vocabulary", Science Advances 11(27), 2025 (the original study)](https://doi.org/10.1126/sciadv.adt3813)
- [Kobak et al., the preprint with the full method on arXiv](https://arxiv.org/abs/2406.07016)

## Questions people ask

### How much biomedical writing involves AI?

At least 13.5% of 2024 PubMed abstracts were processed with LLMs, according to Kobak, González-Márquez, Horvát and Lause (Science Advances, 2025). That is a lower bound, and it ranged from below 5% to over 40% across fields, countries and journals.

### How did they measure it without an AI detector?

Like excess deaths in a pandemic. They projected each word's 2024 frequency from its 2021 to 2022 trend across 15.1 million abstracts, then counted the words that overshot. "Delves" ran 28.0 times its expected rate.

### Is this evidence that writing is converging?

On vocabulary, partly: 379 style words rose together in 2024. It says nothing about rhythm or punctuation, and abstracts are a narrow genre.

### Is this a ScriptGrain study?

No. It is a third-party, peer-reviewed study, credited to its authors and linked to the original.
