Almost a third of the text in a filtered sample of the web from August 2026 was written by AI, according to a new preprint. The detector Pangram labeled 31.1 percent of the tokens, the small pieces of text models read, as AI-generated, up from 27.5 percent in June.

That matters because the web is where language models learn. New models are trained on text that older models wrote. The question is whether that makes them better or worse.

800 models

The team pretrained 800 language models. Each got a different mix of human and AI text. Then they measured how well each one handled text it had not seen.

The answer depends on the amount. For a model short on data, some AI text helps at first. The gain then levels off as more is added, and soon turns into harm, the authors report.

Existing rules for predicting how models improve with more data did not capture this, so the team fit new ones. They released the 83-billion-token dataset they built, called WildAI, along with all 800 models and the code.

The paper has not yet been peer reviewed. The share comes from a single detector, and detectors make mistakes.