Uso de datos sintéticos a futuro no servirán a Inteligencia artificial

Recent research highlights a critical vulnerability in the development of artificial intelligence: the reliance on synthetic data for training models leads to rapid degradation and nonsensical outputs. As major tech companies face diminishing supplies of human-generated content, they increasingly turn to AI-created data. However, experiments demonstrate that training models on their own previous outputs causes errors to amplify, resulting in a swift loss of utility where factual information devolves into incoherent narratives. This "model collapse" occurs because the systems become overwhelmed by the inaccuracies accumulated from successive generations of automated training. This phenomenon underscores the inherent risks of the current trajectory in AI development, suggesting that without careful management, these technologies may fail to maintain accuracy over time. The study implies that the advantage lies with entities that can secure high-quality human data before resources are exhausted. Consequently, the industry is urged to reconsider strategies that prioritize synthetic data loops, as these threaten the long-term viability and reliability of large language models. The findings serve as a warning that technical shortcuts may compromise the fundamental integrity of AI systems. This article is highly relevant to open data because it exposes the dangers of closed, proprietary data silos and the potential contamination of information ecosystems. It highlights the necessity for transparency in data sources and the need for open, verifiable datasets to maintain model integrity. By illustrating how synthetic data corrupts training sets, it advocates for open data practices that ensure authenticity and prevent the systemic errors that arise from unregulated, closed-loop AI training processes, which are essential for trustworthy AI development.

Source: milenio.com
Published on 2024-07-27