Researchers warn we could run out of data to train AI by 2026.

The rapid expansion of AI capabilities faces a critical bottleneck: the depletion of high-quality training data. While the volume of online content appears vast, the supply of clean, reliable text and images is growing much slower than the industry's demand for large datasets. This scarcity threatens to stall the progress of current AI models, which rely on massive amounts of nuanced information to function accurately and avoid propagating biases or inaccuracies found in low-quality sources like social media. To mitigate this risk, the industry is shifting toward strategies that maximize existing resources rather than simply accumulating more. Developers are focusing on creating more efficient algorithms that require less data and exploring the generation of synthetic data to supplement real-world inputs. Additionally, there is a significant move toward licensing content from publishers and digitizing offline repositories, signaling a transition from open scraping to structured, compensated data sourcing to sustain model development. This situation is highly relevant to open_data because it highlights the fragility of relying on freely accessible internet content for foundational technology. The push to monetize and restrict data access undermines the principle of open availability, potentially creating proprietary data silos that hinder transparency and independent research. Understanding this shift is crucial for advocating sustainable, ethical, and accessible data ecosystems that balance innovation with the rights of content creators and the public interest.

Source: thehindu.com
Published on 2023-11-10