The article warns that training artificial intelligence models using data generated by other AIs can lead to a “model collapse,” a degenerative process that irreversibly distorts the reality perceived by these systems. As successive generations learn from contaminated content, the quality of information degrades rapidly, replacing original knowledge with nonsensical hallucinations. This trend poses a critical threat to the reliability of current technological tools, which increasingly depend on recursive loops of automated data. The relevance to open data lies in the critical need to ensure the integrity and provenance of the datasets used for AI development. If training sources are not carefully filtered and validated, the open data ecosystem can become contaminated, propagating biases and systemic errors. This underscores the importance of maintaining high-quality human-generated data as a foundational base, preventing mass automation from compromising the veracity of information available to the public and researchers. Although some experts question the inevitability of collapse in mixed scenarios that include human data, the study serves as a fundamental warning about the risks of algorithmic self-sufficiency. Transparency in data provenance and rigor in data curation are essential to preserve the utility of artificial intelligence. Without strict oversight of information sources, there is a danger that digital tools may become ineffective or even harmful to evidence-based decision-making.
Source:Published on 2024-07-25
Related news
- El fascinante mundo de la Ciencia de Datos: Un viaje desde el preprocesamiento hasta el despliegue
- Las pernoctaciones hoteleras en Castilla-La Mancha aumentaron un 5,1% interanual en el mes de junio
- Los precios hoteleros aumentaron en España casi un 8% en junio, llegando la subida hasta el 18% en Madrid
- La Nación / Medios “independientes” atacan para evitar ley de transparencia
- The tech industry can’t agree on what open-source AI means. That’s a problem.