AI-Generated Data Can Poison Future AI Models

The rapid proliferation of generative AI poses a critical threat to the quality of open data used for training future machine learning models. As synthetic content floods the internet and crowdsourced platforms, it increasingly contaminates the datasets required to develop new technologies. This influx of AI-generated material creates a feedback loop where subsequent models are trained on outputs from previous models, rather than original human-created data. This phenomenon, termed "model collapse," leads to the gradual accumulation of errors and a significant loss of data diversity. The impact is particularly severe on less common information, causing models to lose nuanced details and exacerbate existing biases against marginalized groups. Consequently, the unique variability inherent in human data diminishes, rendering AI outputs less meaningful, accurate, and representative of reality over successive generations. This development is highly relevant to open data because it signals a potential crisis in data integrity for public research and development. Without deliberate efforts to curate and preserve datasets consisting solely of human-created content, the open data ecosystem risks becoming polluted with synthetic noise. Addressing this requires new standards for data provenance and rigorous filtering mechanisms to ensure that open resources remain reliable, unbiased, and reflective of genuine human experience.

Source: scientificamerican.com
Published on 2024-03-10