Descubren miles de imágenes de abuso sexual a menores en las librerías con las que se entrenan las inteligencias artificiales

A recent investigation revealed that the LAION-5B dataset, a massive open collection of billions of images used to train generative artificial intelligence, contains over a thousand confirmed instances of child sexual abuse material. This discovery highlights a critical vulnerability in the foundational data sources of popular AI models, such as Stable Diffusion, which rely on scraped web content to learn visual concepts. Although filters and blocked word lists are employed by developers to mitigate harm, the sheer scale of these public datasets makes comprehensive manual review impossible, allowing illegal content to persist within the training data. The relevance to open data is profound, as it exposes the tension between the open-source ethos of shared, accessible datasets and the ethical necessity of data quality and safety. Open data initiatives typically prioritize volume and accessibility, but this incident demonstrates that without rigorous curation and verification mechanisms, such repositories can inadvertently propagate harmful material. The ability of AI systems trained on these unvetted collections to potentially reproduce similar imagery underscores the urgent need for transparent accountability in data provenance. Ultimately, this case forces a reevaluation of how open datasets are managed and maintained. It suggests that the open data community must develop more robust standards for ethical filtering and collaborative reporting tools to identify illegal content automatically. The temporary withdrawal of these catalogs by LAION signifies a critical moment for the industry, emphasizing that openness cannot come at the expense of safety, and that sustainable open data requires active, ongoing stewardship to prevent the reinforcement of societal harms through algorithmic learning.

Source: elmundo.es
Published on 2023-12-21