OpenAI, Google y Meta se quedan sin datos para entrenar sus modelos de IA, ¿cómo aprenderán estas herramientas?

The rapid advancement of artificial intelligence faces a critical bottleneck: the looming exhaustion of available high-quality training data on the internet. Major tech companies are struggling to find sufficient information to improve their models, prompting them to explore unconventional sources. This scarcity threatens the continuous improvement of AI systems that rely on vast datasets for accurate interpretation and decision-making. To address this challenge, tech giants are considering diverse strategies, such as leveraging internal user data, generating synthetic information, or acquiring media libraries. However, these approaches raise significant ethical and practical concerns. Using proprietary user data conflicts with privacy policies, while relying on synthetic data risks reinforcing existing errors. Similarly, purchasing media companies introduces complex legal and licensing hurdles that complicate data integration. This situation is highly relevant to open data advocates because it highlights the tension between proprietary corporate interests and the collective resource of digital information. As companies scramble for exclusive or restricted datasets, the principle of open, accessible data may be undermined. Ensuring transparency and equitable access to data remains crucial to prevent a future where AI development is solely driven by closed, controlled resources rather than open collaboration.

Source: 20minutos.es
Published on 2024-04-12