Beware of AI 'model collapse': How training on synthetic data pollutes the next generation

Recent research highlights a critical vulnerability in the training of generative AI models known as "model collapse." When models are trained on synthetic data generated by other AI systems rather than human-created content, their accuracy degrades significantly over successive generations. This recursive training loop causes the models to lose track of rare or complex facts, resulting in generic, irrelevant, and eventually nonsensical outputs that fail to reflect reality accurately. The phenomenon occurs because synthetic data often contains subtle errors and biases that compound with each iteration, distorting the probability distributions the model learns. As these "polluted" datasets are fed into new models, the diversity of responses shrinks, and the models begin to hallucinate improbable sequences. This degenerative process suggests that relying heavily on AI-generated content for future training could lead to a systemic failure in model quality, rendering them useless for complex tasks. This finding is highly relevant to open data because it underscores the urgent need to preserve authentic, human-generated datasets for public access and model training. As the internet becomes saturated with AI-generated content, distinguishing high-quality open data from synthetic noise becomes increasingly difficult. Without clear mechanisms to identify and retain genuine human data, the open data ecosystem risks being overwhelmed by low-quality synthetic outputs, hindering the development of reliable and transparent AI systems.

Source: zdnet.com
Published on 2024-07-31