Experts warn AI is running out of training data

Experts predict that the rapid exhaustion of high-quality human-generated data for training large language models could severely constrain AI development by 2026. This scarcity threatens to stall the growth of sophisticated AI systems, potentially impacting billions of dollars in projected economic benefits and limiting the capabilities of tools that are increasingly integrated into daily life and industry. The urgency of this issue highlights a critical dependency in open data ecosystems. Much of the available digital training material consists of publicly shared or scraped information; as this reservoir depletes, the open data community faces a bottleneck that affects not just proprietary models but the broader field of accessible AI research. The depletion of diverse, high-fidelity text and image data underscores the need for sustainable data sources beyond the current internet corpus. However, the industry is expected to adapt through algorithmic innovations rather than sheer data accumulation. Techniques like Mixture of Experts and improved embeddings aim to extract more value from existing datasets, allowing models to achieve higher performance with less raw material. For open data advocates, this shift emphasizes the importance of data quality, metadata, and efficiency in training pipelines over volume, ensuring that future AI advancements remain viable even as easily accessible information becomes scarce.

Source: technology.inquirer.net
Published on 2023-11-21