AI text-to-image generators being trained on images of child sexual abuse: Study

A recent study by the Stanford Internet Observatory reveals that thousands of images depicting child sexual abuse were present in the LAION dataset, a widely used open-source resource for training artificial intelligence image-generation models. This discovery highlights a critical failure in current safety filtering mechanisms for web-scale datasets, demonstrating that illicit and harmful content can persist even in systems designed with rigorous screening protocols. The findings challenge the assumption that AI-generated abuse is merely a synthesis of benign concepts, instead showing that models are directly trained on actual illegal material. This contamination significantly impacts the potential for AI tools to generate explicit content involving minors, including deepfakes and non-consensual imagery. The presence of such material raises serious ethical and legal concerns regarding the commercial exploitation of victims and the perpetuation of harm. The researchers argue that the risks associated with uncurated, open datasets are too severe for public deployment, as they inevitably contain not only illegal content but also non-consensual intimate imagery and copyrighted material that violate privacy and intellectual property rights. The study urges the open_data community to reconsider the practice of releasing massive, unvetted datasets to the public. Instead, it advocates for restricting such data to controlled research environments and developing more curated, well-sourced alternatives for public models. While LAION has responded by temporarily withdrawing its datasets for cleaning and collaboration with safety organizations, the incident underscores the urgent need for stricter governance and ethical standards in how open data is collected, filtered, and distributed to prevent further exploitation and harm.

Source: courthousenews.com
Published on 2023-12-21