La IA genera audios plagados de machismo, racismo e infracciones de derechos de autor

A comprehensive study reveals that the audio datasets used to train generative artificial intelligence are rife with social biases, offensive language, and copyrighted material. The research found that many datasets reflect gender stereotypes and racial discrimination, perpetuating prejudices rather than representing authentic human diversity. This implies that the resulting models may amplify existing discriminations, negatively affecting underrepresented or marginalized groups. This finding is crucial for the open data community because it underscores the urgent need to audit and clean databases before their publication or commercial use. Unlike text, audio files require greater computational resources for analysis, which hinders transparency and equitable access. The lack of data cleaning in these open datasets not only compromises technical quality but also raises significant legal and ethical risks, particularly when treating voice as protected biometric data. The relevance to the open data movement lies in the urgent need to develop ethical standards that balance information availability with the protection of individual rights and cultural diversity. If these biases and intellectual property violations are not addressed, open data ecosystems could become vehicles for exclusion and legal abuse. Therefore, shared responsibility demands that data providers prioritize ethical curation over mere volume accumulation, ensuring that technological innovation is not built on non-consensual exploitation or the reproduction of social prejudices.

Source: elpais.com
Published on 2024-12-10