Next up, the defence gets a turn, and honestly it’s a lovely little study. Fitterer, Gangl and Ulbrich, Technische Universität Berlin, presented at the ACL 2025 Student Research Workshop. Peer reviewed, but small, and they’d say so themselves.
The question was blunt. If everyone’s using LLMs, has online English gone lexically flat? To find out, they pulled two random samples of roughly 30,000 news articles each from the News on the Web corpus, one from 2018 and one from 2024. Then they ran three lexical-diversity measures plus a ratio of LLM-style words.
You know the ones. The words people point at and say “that’s the robot talking”. I won’t list them, because I’m not here to hand out a bingo card.
What they found cuts both ways, which is why it’s fun. The LLM-style-word ratio rose from 0.230% to 0.347% of words, about half as much again. So yes, the fingerprints are there. But lexical diversity did not fall. One measure, MTLD, went from 214.45 to 254.65, so on that measure the 2024 articles were more varied; the other two barely moved.
Their own conclusion is careful: an apparent influence of LLMs on written online English, but the homogenisation effects don’t show in the measurements. I think that’s about as honest as a sentence gets.
Now, the caveats, and the court should hear them. Lexical diversity is one narrow test. It tells you about vocabulary, and nothing about rhythm, structure or punctuation, which is where a lot of the machine sameness is supposed to live. And news articles come from edited newsrooms, which may be the least affected corner of the web. A sub-editor is a decent filter. Your average content farm doesn’t have one.
So the defence gets a real point on the board, just not a knockout.
Verdict: for the defence.