This one’s a bit different, and I’ll be upfront: it’s a preprint. He, Ardi, Govindan and Madhyastha, University of Southern California, April 2026, arXiv, not yet peer reviewed. The paper is called “DeGenTWeb: A First Look at LLM-dominant Websites”. Treat it as a witness who hasn’t been cross-examined yet.
The twist is that they don’t judge articles one by one. They classify whole websites as LLM-dominant, running a calibrated detector over several pages per site. The sample was roughly 100k sites archived by Common Crawl, plus roughly 20k sites from Bing results for 10k how-to searches.
And the numbers jump out. LLM-dominant sites rose from 2.1% in the second half of 2022 to 29.4% in the first half of 2025. That’s a lot of sites that, as far as the detector can tell, had very little human input.
Then the searches. For 46.6% of the how-to searches, at least one top-10 result was an LLM-dominant site. Widen it to the top 20 and it’s 65.7%. If you’ve ever searched for how to fix something and landed on a page that says a great deal and tells you nothing, well, you’ll feel seen.
Money turns up too. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites. Make of that what you will; I’d say it’s a motive, and every good trial needs one.
Fairness requires the caveats. It’s a preprint. It’s detector-based, though they validated on pre-ChatGPT sites and flagged only 0.29%, and they openly warn that AI-text detectors perform much worse than advertised when you try not to falsely accuse human writers. And it counts sites with little human input, so it misses human copy that was merely AI-assisted. The real number of machine-touched pages could be higher, or the labelling could wobble. Probably some of each.
Verdict: for the prosecution.