La lista secreta de sitios web que hacen que una IA como ChatGPT parezca inteligente

The research reveals that AI chatbots rely on massive datasets scraped from the web without adequate oversight, raising serious ethical and legal concerns regarding intellectual property, privacy, and security. By incorporating content from pirate sites, personal blogs, and public databases without consent or compensation, these technologies expose creators and users to significant risks, highlighting the lack of transparency in the practices of major technology companies. Analysis of training data shows the presence of dangerous biases, hate speech, and misinformation, while safety filters prove insufficient for removing extremist or illegal content. This demonstrates that the “black box” of AI training not only replicates human knowledge but also internalizes prejudices and misinformation, which can lead to the spread of false or discriminatory information. The opacity of these data sources hinders auditing and accountability, directly undermining public trust in these tools. This report is crucial for the open data movement, as it underscores the urgent need for transparency and responsible governance in AI development. It advocates for the principle that data used to train global systems must be auditable, ethically sourced, and publicly documented. Without clear mechanisms for openness and oversight, technological progress is built on unstable foundations that may violate fundamental rights, making transparency an essential demand for fair and sustainable AI.

Source: infobae.com
Published on 2023-04-20