La gigantesca biblioteca digital Internet Archive sufre una caída mundial porque usaban sus datos para entrenar una IA

The article highlights a critical vulnerability in accessing open data at scale, where AI-driven mass downloading overwhelmed the Internet Archive’s infrastructure. This incident demonstrates the fragility of open repositories when faced with uncoordinated automated access, jeopardizing service availability for all users rather than just the offending entities. The collapse underscores the urgent need for sustainable access policies and direct communication between data consumers and providers. It suggests that responsible stewardship of public domain materials requires users to manage their request volumes carefully, preventing technical bottlenecks that hinder the broader mission of open information preservation. This case is relevant to open data as it illustrates the tension between unrestricted access and system stability. It advocates for structured protocols in open data ecosystems to ensure that high-demand users, such as AI developers, do not inadvertently exclude the general public from utilizing these vital cultural and informational resources.

Source: 20minutos.es
Published on 2023-06-02