'Model Collapse': Future AI generations may start speaking gibberish you don't understand
This research reveals that AI systems degrade over successive training cycles, producing lower-quality outputs that pollute the digital ecosystem. This "cognitive pollution" limits the diversity of data available for future models, creating barriers for new entrants and entrenching advantages for those with existing large-scale data access. The findings are critical for open data advocates as they highlight a potential crisis in data sustainability. The influx of synthetic content threatens to contaminate public datasets, making it increasingly difficult to train accurate and diverse AI models using openly available web scrapes. Consequently, the reliance on public, open data sources may become obsolete or unreliable. This underscores the urgent need for transparent data provenance and robust filtering mechanisms within the open data community to preserve the integrity of training materials against AI-generated noise.
Source: businesstoday.inPublished on 2023-06-21
Related news
- The open-source AI boom is built on Big Tech’s handouts. How long will it last?
- Seeing Machines unveils collaboration with Swedish startup Devant
- Generative AI Is at the Heart of the Ongoing Reddit Protest—Here’s Why
- Caso Cecilia Strzyzowski: el Gobierno ofrece 5 millones de pesos a quien aporte datos
- Parents of Nashville shooting victims urge court to deny release of shooter’s writings in new declarations | KRDO