AI-Generated Data Can Poison Future AI Models
The rapid proliferation of generative AI poses a critical threat to the quality of open data used for training future machine learning models. As synthetic content floods the internet and crowdsourced platforms, it increasingly contaminates the datasets required to develop new technologies. This influx of AI-generated material creates a feedback loop where subsequent models are trained on outputs from previous models, rather than original human-created data. This phenomenon, termed "model collapse," leads to the gradual accumulation of errors and a significant loss of data diversity. The impact is particularly severe on less common information, causing models to lose nuanced details and exacerbate existing biases against marginalized groups. Consequently, the unique variability inherent in human data diminishes, rendering AI outputs less meaningful, accurate, and representative of reality over successive generations. This development is highly relevant to open data because it signals a potential crisis in data integrity for public research and development. Without deliberate efforts to curate and preserve datasets consisting solely of human-created content, the open data ecosystem risks becoming polluted with synthetic noise. Addressing this requires new standards for data provenance and rigorous filtering mechanisms to ensure that open resources remain reliable, unbiased, and reflective of genuine human experience.
Source: scientificamerican.comPublished on 2024-03-10
Related news
- Comienza el Censo de Población y Vivienda 2024 con múltiples medidas de seguridad | Puranoticia.cl
- Cooling performance of NT government shade structure worsening as maintenance costs soar
- Biden a los líderes legislativos: es responsabilidad del Congreso mantener el gobierno abierto
- Le intentan robar el móvil y explica el método para que no te pase: "Llevaba el número en un papel"