Skip to content

A Third of New Web Pages Are Written by a Machine: on .gov and .edu It Is One Percent, on .com Ten Times More

1 min read
Share
A Third of New Web Pages Are Written by a Machine: on .gov and .edu It Is One Percent, on .com Ten Times More

A study by the American centre Pew Research measured what everyone assumed but nobody counted: how much of the internet is no longer written by a human. The answer, for pages published after November 2022, when ChatGPT came out, is 35 percent. Every third new page on the web carries clear traces of having been assembled or thoroughly reworked by a machine.

The method is simple enough to check. Around half a million English-language pages from the last five years were pulled from the Common Crawl archive, deliberately including the period before ChatGPT existed. The Open Pangram tool, which estimates whether a text is machine-made, was then run over them. In a random sample of 10,000 pages collected in July this year, 10 percent showed significant signs of artificial authorship. But that includes old pages, written when such tools did not exist. When Pew removed the older ones and kept only those published after ChatGPT, the figure jumped to the 35 percent mentioned above.

Where machine text lives

The breakdown by domain is the part that says the most. Commercial .com addresses show artificial authorship around ten times more often than .edu and .gov, which hold at about one percent. Organisations on .org sit at 4.6 percent. In other words: where text is written to sell something or catch a click, the machine has already taken over the shift. Where somebody signs with their own name and institution, it has not yet.

The report comes weeks after Cloudflare announced that bot traffic had overtaken human traffic, and faster than the company itself had predicted. Put the two measurements together and you get a picture heavier than it sounds: bots reading pages written by other bots. The human in that exchange is increasingly optional.

Pew itself admits the measurement is not perfect. Pangram, like every detection tool, can wrongly flag human text as machine-made. But on a scale of half a million pages, the direction is hard to dispute, even if the exact figure shifts by a few percent.

The list of small habits that have grown over the years is also interesting: longer dashes, the Oxford comma, the „not X, but Y” construction. None of that is proof in itself. But when an entire web starts writing with the same rhythm and the same mannerisms, the difference between ten sources and one source becomes cosmetic.

The question that remains is not technological. If a third of what is new on the internet was assembled by a machine trained on what was older on the internet, what exactly will the next models learn? And when somebody looks for an answer to something banal next year, will they find what somebody actually knew, or only what rhymes nicely with everything else written in the meantime?