---
title: "The fingerprints are there. The flattening is not."
url: https://deadwritingtheory.com/cases/the-news-that-didnt-flatten/
summary: "Machine-style words rose in English news between 2018 and 2024, but lexical diversity did not fall. A real point for the defence on vocabulary, not a knockout."
published: 2026-10-04
updated: 2026-10-04
author: "Jack Stovell"
publisher: "Adapt Progress Evolve Limited"
language: en-GB
---

# The fingerprints are there. The flattening is not.

Machine-style words rose in English news between 2018 and 2024, but lexical diversity did not fall. A real point for the defence on vocabulary, not a knockout.

**Case No. DWT-004: The news that didn't flatten.** Verdict for the defence. Filed 4 October 2026.

- **The question:** Has the widespread use of large language models made English news articles lexically more alike?
- **The sample:** Two random samples of roughly 30,000 news articles each from the News on the Web (NOW) corpus, one from 2018 and one from 2024
- **The study:** Sarah Fitterer, Dominik Gangl, Jannes Ulbrich (Technische Universität Berlin), "Testing English News Articles for Lexical Homogenization Due to Widespread Use of Large Language Models", Proceedings of ACL, Student Research Workshop, 2025. https://doi.org/10.18653/v1/2025.acl-srw.95
- **Whose study:** Third-party study. Not our study. Entered into evidence from Sarah Fitterer, Dominik Gangl, Jannes Ulbrich, Technische Universität Berlin.
- **Peer review:** Peer reviewed

## Findings of fact

1. The LLM-style-word ratio rose from 0.230% to 0.347% of words, a difference of 0.117% against within-year variation of 0.016%.
2. MTLD lexical diversity went from 214.45 to 254.65: on that measure the 2024 articles were more varied, not less.
3. MATTR went from 0.88011 to 0.88121, a difference of 0.00110 against within-year variation of 0.00109.
4. Maas went from 0.01469 to 0.01482, a difference of 0.00013 against within-year variation of 0.00016.
5. The authors conclude: "while there is an apparent influence of LLMs on written online English, homogenization effects do not show in the measurements."

## Caveats for the jury

- A student research workshop paper: peer reviewed, but a small study.
- Lexical diversity is one narrow test of sameness. It says nothing about rhythm, structure or punctuation.
- News comes from edited newsrooms, which may be the least affected corner of the web, as the authors note.

## The court reporter's account

Next up, the defence gets a turn, and honestly it's a lovely little study. Fitterer, Gangl and Ulbrich, Technische Universität Berlin, presented at the ACL 2025 Student Research Workshop. Peer reviewed, but small, and they'd say so themselves.

The question was blunt. If everyone's using LLMs, has online English gone lexically flat? To find out, they pulled two random samples of roughly 30,000 news articles each from the News on the Web corpus, one from 2018 and one from 2024. Then they ran three lexical-diversity measures plus a ratio of LLM-style words.

You know the ones. The words people point at and say "that's the robot talking". I won't list them, because I'm not here to hand out a bingo card.

What they found cuts both ways, which is why it's fun. The LLM-style-word ratio rose from 0.230% to 0.347% of words, about half as much again. So yes, the fingerprints are there. But lexical diversity did not fall. One measure, MTLD, went from 214.45 to 254.65, so on that measure the 2024 articles were more varied; the other two barely moved.

Their own conclusion is careful: an apparent influence of LLMs on written online English, but the homogenisation effects don't show in the measurements. I think that's about as honest as a sentence gets.

Now, the caveats, and the court should hear them. Lexical diversity is one narrow test. It tells you about vocabulary, and nothing about rhythm, structure or punctuation, which is where a lot of the machine sameness is supposed to live. And news articles come from edited newsrooms, which may be the least affected corner of the web. A sub-editor is a decent filter. Your average content farm doesn't have one.

So the defence gets a real point on the board, just not a knockout.

Verdict: for the defence.

## The exhibits entered in this case

### For the prosecution

- No exhibit entered for the prosecution in this case.

### For the defence

- **Exhibit D-8**: Lexical diversity in news did not fall. **214.45 → 254.65** MTLD lexical diversity, 2018 → 2024. Two samples of roughly 30,000 news articles each. Machine-style words did rise, from 0.230% to 0.347% of words, but the homogenisation did not show.
- **Exhibit D-9**: Two other diversity measures barely moved. **0.00110** change in MATTR lexical diversity, 2018 → 2024. Against within-year variation of 0.00109 (MATTR 0.88011 → 0.88121). Maas moved from 0.01469 to 0.01482, a difference of 0.00013 against 0.00016.

## The finding

Machine-style words rose in English news between 2018 and 2024, but lexical diversity did not fall. A real point for the defence on vocabulary, not a knockout.

## Sources

- [Fitterer et al., "Testing English News Articles for Lexical Homogenization Due to Widespread Use of Large Language Models", Proceedings of ACL, Student Research Workshop, 2025 (the original study)](https://doi.org/10.18653/v1/2025.acl-srw.95)
- [Fitterer, Gangl and Ulbrich in the ACL Anthology](https://aclanthology.org/2025.acl-srw.95/)

## Questions people ask

### Has AI made news writing more alike?

Not on this evidence. Fitterer, Gangl and Ulbrich (TU Berlin, ACL 2025) compared roughly 30,000 news articles from 2018 with roughly 30,000 from 2024. Lexical diversity did not fall: MTLD went from 214.45 to 254.65, and the other two measures barely moved.

### Did AI words show up in the news?

Yes. The ratio of LLM-style words rose from 0.230% to 0.347% of words. The authors call it an apparent influence of LLMs on written online English, with no homogenisation in the measurements.

### Why is the verdict for the defence?

Because the study tested for homogenisation directly and did not find it. It only tests vocabulary, though, and only in edited news, so it is a point for the defence rather than a knockout.
