This one I actually like. No detector, no made-up AI texts, which is the part that matters.
The authors are Kobak, González-Márquez, Horvát and Lause, of the University of Tübingen and Northwestern University. The paper is “Delving into LLM-assisted writing in biomedical publications through excess vocabulary”, published in Science Advances in 2025, and it was peer reviewed. So far, so respectable.
The method, plainly. They borrowed the idea of “excess mortality”. For each word, they predicted its 2024 frequency from the 2021 and 2022 trend, then compared that with what actually happened. The gap is the excess. The sample was 15.1 million English PubMed abstracts from 2010 to 2024.
The headline words first. “Delves” was used 28.0 times as often as expected in 2024. “Underscores” came in at 13.8 times, and “showcasing” at 10.7 times. Altogether, 379 style words showed excess use in 2024; 66% of them were verbs and 14% were adjectives.
Now the bit I found sharpest. In earlier years the excess words were content words. In the Covid years that meant coronavirus, pandemic, lockdown: nouns about the news, which is what you’d expect, because the world changed and the vocabulary followed. In 2024 the excess words were almost all style words. The excess was in how things were said, not in what they were about.
Then the number the court will care about. Combining marker words gives a lower bound: at least 13.5% of 2024 abstracts were processed with LLMs, meaning at least 200,000 papers a year. That lower bound ranged from below 5% to over 40% across fields, countries and journals. The four Covid words of 2021 gave a gap of 0.069, against 0.135 for the LLM words. In the authors’ reading, LLM usage in 2024 was at least two times the size of the Covid-related literature in 2021. They call it an “unprecedented impact on scientific writing”.
Caveats, and there are real ones. It’s a lower bound, so the true share is likely higher. It covers abstracts only, and an abstract is a formulaic genre. And this is the point I’d underline: the study measures prevalence plus word convergence, the uptake of the machine’s words. It doesn’t touch rhythm or punctuation. Those are different questions, and this study doesn’t answer them.
Still, the method is clean and the numbers are large. The defence has no obvious line of attack on how the evidence was gathered, only on how far it reaches.
Verdict: for the prosecution.