AI 'gold rush' for chatbot training data could run out of human-written text

Research indicates that AI systems will likely exhaust the supply of high-quality, publicly available human-generated text by the early 2030s. This depletion threatens the primary method for scaling current language models, as the abundance of internet data has fueled rapid improvements in AI capabilities. Without this resource, developers face a significant bottleneck in maintaining the trajectory of technological advancement, forcing a strategic pivot toward new data sourcing methods. Consequently, the industry is compelled to consider controversial alternatives, such as accessing private user data or relying heavily on synthetic content generated by AI itself. However, these solutions present serious risks, including "model collapse" where systems degrade by learning from their own imperfect outputs, thereby amplifying biases and errors. This shift challenges the sustainability of current training paradigms, highlighting the limitations of simply overtraining on recycled or artificially created information. This issue is critically relevant to the open_data movement, which advocates for the preservation and accessibility of public knowledge. As AI companies compete for valuable digital assets, platforms like Wikipedia and news outlets must navigate complex ethical and economic questions regarding their content usage. Ensuring that human-created data remains available and incentivized is essential to prevent the pollution of the information ecosystem, thereby safeguarding the quality and integrity of future AI development against a backdrop of resource scarcity.

Source: voanews.com
Published on 2024-06-09