Skip to the content

The Court of Public Prose

In re Dead Writing Theory

Case No.
DWT-005
Matter
The websites nobody wrote
Filed
Verdict
Verdict for the prosecution

More sites with little human input. More of them selling.

For the prosecution every exhibit in this case · P-7 to P-8
The question
How many websites are dominated by LLM-generated text, how has that changed since ChatGPT, and do such sites reach people through search?
The sample
Roughly 100k sites archived by Common Crawl and roughly 20k sites found in Bing results for 10k how-to search queries, each site classified from several of its pages
Filed
The study
DeGenTWeb: A First Look at LLM-dominant Websites Sichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha. University of Southern California. arXiv preprint 2605.00087, 2026.

§ IFindings of fact

What the study found.

  1. The share of LLM-dominant sites rose steadily from 2.1% in the second half of 2022 to 29.4% in the first half of 2025.
  2. For 46.6% (4,664 of 10,000) of how-to search queries, at least one of the top 10 results pointed to an LLM-dominant site; for the top 20 results it was 65.7%.
  3. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites.
  4. On sites from before ChatGPT, the classifier flagged only 0.29% as LLM-dominant, an approximate upper bound on its false positive rate.
  5. The authors warn that, when aiming not to falsely attribute human writing to LLMs, "detectors of LLM-generated text perform much worse than advertised."

§ IICaveats for the jury

What this study cannot tell the court.

  • A preprint, in submission: not yet peer reviewed.
  • It relies on an AI-text detector, though validated with a published false-positive bound (0.29%).
  • It counts sites with little human input, so it misses human copy that was only AI-assisted.
  • Site creation dates are estimated from the earliest archive date.

§ IVThe exhibits entered in this case

Prosecution v. Defence

Any tally counts the evidence items below; it is not a measure of truth.

Evidence for the theory

For the prosecution

2 exhibits
  1. Exhibit P-7Case No. DWT-005

    2.1% → 29.4%

    of sites, second half of 2022 → first half of 2025

    LLM-dominant websites rose steadily

    Roughly 100k sites archived by Common Crawl, classified site by site. A preprint; on pre-ChatGPT sites the classifier flagged only 0.29%.

    Entered for the prosecution · He et al., arXiv preprint 2605.00087, 2026 · third-party study

  2. Exhibit P-8Case No. DWT-005

    46.6%

    of 10,000 how-to searches had one in the top 10

    Machine-written sites reach the top of how-to searches

    65.7% for the top 20 results. 78.8% of LLM-dominant sites have a clear financial incentive, against 55.8% of other sites.

    Entered for the prosecution · He et al., arXiv preprint 2605.00087, 2026 · third-party study

Evidence against the theory

For the defence

0 exhibits

No exhibit entered for the defence in this case.

§ VThe finding

The finding

LLM-dominant websites rose from 2.1% to 29.4% of sites sampled between late 2022 and early 2025, and they turn up in how-to search results. A preprint, detector-based.
Verdict for the prosecutionCase No. DWT-005 · The websites nobody wrote
Read the original studyHe et al., arXiv preprint 2605.00087, 2026. Not our study.

§ VIQuestions for the court

Asked and answered.

How many websites are written mostly by AI?

In the DeGenTWeb preprint (He et al., University of Southern California, 2026), the share of LLM-dominant sites in its Common Crawl sample rose from 2.1% in the second half of 2022 to 29.4% in the first half of 2025.

Do AI-written sites show up in search?

Yes. For 46.6% of 10,000 how-to searches, at least one of the top 10 Bing results was an LLM-dominant site; for the top 20 results it was 65.7%.

Can AI detectors be trusted?

The authors warn that detectors "perform much worse than advertised" when you try not to falsely accuse human writers. Their own site-level classifier flagged only 0.29% of pre-ChatGPT sites, but the study is a preprint and not yet peer reviewed.