La desaparición de los datos que alimentan la IA: Un problema en auge

Access to open data for training artificial intelligence is facing a significant crisis, marked by the mass withdrawal of web content due to concerns about consent. Researchers have identified a sharp decline in the availability of essential public information, forcing a rethinking of traditional data collection methods and raising ethical questions about the appropriation of digital resources without explicit authorization. This restriction impacts not only major technology corporations but also the academic community and nonprofit organizations that rely on accessible datasets for research. The difficulty in obtaining high-quality information, combined with the rise of alternative solutions such as synthetic data, suggests that the current model of AI development is unsustainable without a fundamental shift in how computing resources are managed and shared. This article is relevant to open data because it highlights the urgency of establishing clear governance and consent frameworks. The current situation underscores the need for tools that enable content creators to specifically control how their data is used, promoting reciprocity and legal respect. Without these measures, the availability of open data will continue to decline, threatening transparency and collaborative innovation in the global technology ecosystem.

Source: wwwhatsnew.com
Published on 2024-07-21