This is the big one, and I mean that in the dull sense: lots of texts, lots of datasets, lots of moving parts. Bear with me.
The authors are Sourati, Karimi-Malekabadi, Ozcan, McDaniel, Ziabari, Trager, Tak, Chen, Morstatter and Dehghani, of the University of Southern California. The paper is “The shrinking landscape of linguistic diversity in the age of large language models”, Nature Human Behaviour, 2026, peer reviewed. Three studies, seven datasets, over 880,000 texts.
Study 1a tracked the variance of writing complexity month by month, January 2018 to November 2024. The material was 80,238 arXiv computer science abstracts, 318,490 Reddit creative stories and 379,583 Patch community news articles (Patch only to November 2023). The authors set that variance against the share flagged as AI by a detector called Binoculars. Study 1b did something different: GPT-3.5, Llama 3 70B and Gemini Pro each polished 1,000 pre-ChatGPT human texts from each of Reddit and arXiv, using neutral prompts.
Findings first. The variance of writing complexity fell after ChatGPT’s release in all three datasets. On Reddit and arXiv, AI use predicted later drops in variance. On Patch news it did not, and the authors suggest editorial workflows may buffer the effect. Hold that thought.
In the polishing experiment, meaning was kept: 87% of similarity scores were above 0.95. But writing-complexity variance fell by a statistically significant 21 to 50% across datasets and models. So the machine leaves the message alone and quietly narrows the style. LLMs also amplified patterns associated with dominant characteristics and suppressed others, which is a worrying thing to find in a polishing tool.
Now a contrast for the court. The Fitterer case (DWT-004) found no lexical flattening in edited news. This study finds the range of styles narrowing, but most clearly where no editor stands in the way. Those two don’t quite collide. One counted words in edited news; the other measured a composite of complexity, and its weakest link to AI use was in the edited news. This is about convergence, whether writing became more alike, not just about how much text involved AI. That’s the question the theory really turns on.
Caveats, and they matter. The 21 to 50% comes from rewrite experiments, so it’s what an LLM does to a text, not what writers actually published. The real-world part is a correlation in time, nothing stronger. And it measures one composite of complexity, not words or punctuation.
So the court gets a strong finding with a clear boundary around it. The range is narrowing, in the places the study could see, and the cleanest numbers come from the lab rather than the wild. I’d call that persuasive, not conclusive.
Verdict: for the prosecution.