Researchers have raised concerns over the potential collapse of generative AI models when they are excessively trained on AI-generated content rather than original data. This warning highlights the ongoing debate surrounding the use of synthetic data in AI training, as it can still be beneficial if combined with real data and properly filtered. To combat quality degradation, experts advocate for maintaining data provenance, ensuring the origin of the training data is documented and that human-created sources remain accessible.

Reuters: Reuters is a global news organization that publishes financial, political, and business reporting and commentary. It is relevant here because the Breakingviews item appears under Reuters’ editorial umbrella and frames the discussion of AI model competition and model quality.
AI models: AI models are software systems trained on data to generate text, images, code, or other outputs. The news item points to a broader debate over what happens when generative models increasingly consume content produced by other models, a dynamic associated with model collapse concerns.
Breakingviews: Reuters Breakingviews is a commentary and analysis service that publishes short-form opinion on markets, companies, and macroeconomic developments. In this item, it is the source of a weekly column about competing AI business models and the risks created when AI systems train on AI-generated output.

Model collapse: Researchers have warned that generative AI systems can degrade when they are trained too heavily on AI-generated content rather than original data.
Synthetic data: Recent work suggests synthetic data can still be useful if it is carefully filtered or combined with real data instead of fully replacing it.
Data provenance: A common mitigation for AI quality degradation is to preserve the origin of training data and keep access to human-created sources.