Scaling Laws for Wild AI-Generated Web Text in Pretraining
October 1, 2026
Analysis of 800 language models shows that unlabeled AI-generated text from web scrapes affects pretraining differently based on model scale. For data-starved models, AI tokens initially lower human-text loss, but this benefit saturates and reverses into harm as AI token ratios increase. High-budget human-text models experience immediate loss increases when exposed to AI tokens.
HOW THIS AFFECTS YOU
●
builderThis affects your data procurement strategy as the volume of AI-generated web content grows.
●
researcherYou should account for the specific ratio of synthetic-to-human data when designing pretraining curricula.