AI 'gold rush' for chatbot training data could run out of human-written text - ET Telecom
Recent research indicates that the exponential growth of artificial intelligence faces an impending bottleneck as high-quality, human-generated public text data is projected to be exhausted by the early 2030s. This "data famine" threatens the current paradigm of scaling language models by simply adding more information, forcing tech companies to compete fiercely for remaining valuable sources like news outlets and social media platforms. Without fresh human input, the industry may struggle to maintain its rapid pace of innovation, prompting a shift toward purchasing exclusive data access or exploring alternatives to traditional training methods. In the absence of new human content, developers are increasingly tempted to utilize synthetic data generated by existing AI models. However, experts warn that this approach risks "model collapse," where iterative training on AI-generated outputs degrades quality and amplifies existing biases and errors. This reliance on imperfect synthetic data highlights a critical limitation in current technological strategies, suggesting that merely increasing computational power or reusing existing data is insufficient for sustaining long-term progress in model capabilities and reliability. This narrative is highly relevant to the open data community as it underscores the existential value of freely accessible, high-quality public information. If the internet’s open ecosystem is depleted for AI training, it could lead to increased enclosure of knowledge, where essential data becomes proprietary or paywalled. Furthermore, the potential pollution of the web with low-quality synthetic content threatens the integrity of open repositories like Wikipedia. Open data advocates must therefore emphasize the preservation of authentic human contributions and the importance of ethical data sourcing to prevent the collapse of the public knowledge commons that underpins modern AI development.
Source: telecom.economictimes.indiatimes.comPublished on 2024-06-08
Related news
- Zyphra debuts Zyda LLM training dataset with 1.3T tokens
- Meta's plan to train its AI on all your old Facebook data is raising eyebrows among privacy advocates
- Critican a Meta por dificultar a los usuarios oponerse al uso de sus datos para entrenar la IA
- AI ‘gold rush’ for chatbot training data could run out of human-written text