---
title: "More sites with little human input. More of them selling."
url: https://deadwritingtheory.com/cases/the-websites-nobody-wrote/
summary: "LLM-dominant websites rose from 2.1% to 29.4% of sites sampled between late 2022 and early 2025, and they turn up in how-to search results. A preprint, detector-based."
published: 2026-10-04
updated: 2026-10-04
author: "Jack Stovell"
publisher: "Adapt Progress Evolve Limited"
language: en-GB
---

# More sites with little human input. More of them selling.

LLM-dominant websites rose from 2.1% to 29.4% of sites sampled between late 2022 and early 2025, and they turn up in how-to search results. A preprint, detector-based.

**Case No. DWT-005: The websites nobody wrote.** Verdict for the prosecution. Filed 4 October 2026.

- **The question:** How many websites are dominated by LLM-generated text, how has that changed since ChatGPT, and do such sites reach people through search?
- **The sample:** Roughly 100k sites archived by Common Crawl and roughly 20k sites found in Bing results for 10k how-to search queries, each site classified from several of its pages
- **The study:** Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha (University of Southern California), "DeGenTWeb: A First Look at LLM-dominant Websites", arXiv preprint 2605.00087, 2026. https://arxiv.org/abs/2605.00087
- **Whose study:** Third-party study. Not our study. Entered into evidence from Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha, University of Southern California.
- **Peer review:** Preprint, not yet peer reviewed

## Findings of fact

1. The share of LLM-dominant sites rose steadily from 2.1% in the second half of 2022 to 29.4% in the first half of 2025.
2. For 46.6% (4,664 of 10,000) of how-to search queries, at least one of the top 10 results pointed to an LLM-dominant site; for the top 20 results it was 65.7%.
3. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites.
4. On sites from before ChatGPT, the classifier flagged only 0.29% as LLM-dominant, an approximate upper bound on its false positive rate.
5. The authors warn that, when aiming not to falsely attribute human writing to LLMs, "detectors of LLM-generated text perform much worse than advertised."

## Caveats for the jury

- A preprint, in submission: not yet peer reviewed.
- It relies on an AI-text detector, though validated with a published false-positive bound (0.29%).
- It counts sites with little human input, so it misses human copy that was only AI-assisted.
- Site creation dates are estimated from the earliest archive date.

## The court reporter's account

This one's a bit different, and I'll be upfront: it's a preprint. He, Ardi, Govindan and Madhyastha, University of Southern California, April 2026, arXiv, not yet peer reviewed. The paper is called "DeGenTWeb: A First Look at LLM-dominant Websites". Treat it as a witness who hasn't been cross-examined yet.

The twist is that they don't judge articles one by one. They classify whole websites as LLM-dominant, running a calibrated detector over several pages per site. The sample was roughly 100k sites archived by Common Crawl, plus roughly 20k sites from Bing results for 10k how-to searches.

And the numbers jump out. LLM-dominant sites rose from 2.1% in the second half of 2022 to 29.4% in the first half of 2025. That's a lot of sites that, as far as the detector can tell, had very little human input.

Then the searches. For 46.6% of the how-to searches, at least one top-10 result was an LLM-dominant site. Widen it to the top 20 and it's 65.7%. If you've ever searched for how to fix something and landed on a page that says a great deal and tells you nothing, well, you'll feel seen.

Money turns up too. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites. Make of that what you will; I'd say it's a motive, and every good trial needs one.

Fairness requires the caveats. It's a preprint. It's detector-based, though they validated on pre-ChatGPT sites and flagged only 0.29%, and they openly warn that AI-text detectors perform much worse than advertised when you try not to falsely accuse human writers. And it counts sites with little human input, so it misses human copy that was merely AI-assisted. The real number of machine-touched pages could be higher, or the labelling could wobble. Probably some of each.

Verdict: for the prosecution.

## The exhibits entered in this case

### For the prosecution

- **Exhibit P-7**: LLM-dominant websites rose steadily. **2.1% → 29.4%** of sites, second half of 2022 → first half of 2025. Roughly 100k sites archived by Common Crawl, classified site by site. A preprint; on pre-ChatGPT sites the classifier flagged only 0.29%.
- **Exhibit P-8**: Machine-written sites reach the top of how-to searches. **46.6%** of 10,000 how-to searches had one in the top 10. 65.7% for the top 20 results. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites.

### For the defence

- No exhibit entered for the defence in this case.

## The finding

LLM-dominant websites rose from 2.1% to 29.4% of sites sampled between late 2022 and early 2025, and they turn up in how-to search results. A preprint, detector-based.

## Sources

- [He et al., "DeGenTWeb: A First Look at LLM-dominant Websites", arXiv preprint 2605.00087, 2026 (the original study)](https://arxiv.org/abs/2605.00087)

## Questions people ask

### How many websites are written mostly by AI?

In the DeGenTWeb preprint (He et al., University of Southern California, 2026), the share of LLM-dominant sites in its Common Crawl sample rose from 2.1% in the second half of 2022 to 29.4% in the first half of 2025.

### Do AI-written sites show up in search?

Yes. For 46.6% of 10,000 how-to searches, at least one of the top 10 Bing results was an LLM-dominant site; for the top 20 results it was 65.7%.

### Can AI detectors be trusted?

The authors warn that detectors "perform much worse than advertised" when you try not to falsely accuse human writers. Their own site-level classifier flagged only 0.29% of pre-ChatGPT sites, but the study is a preprint and not yet peer reviewed.
