¿De dónde sale la información que da ChatGPT?
The research reveals that artificial intelligence models are trained on vast amounts of web data containing racial and religious biases, as well as copyrighted content, posing significant ethical and legal risks. When inadequately filtered datasets are used, these technologies can perpetuate prejudices and facilitate the spread of propaganda or misinformation, while exposing creators and users to violations of privacy and intellectual property rights. A critical aspect is the lack of transparency regarding the origin of the information, since a considerable portion of the websites used does not appear publicly on the current web, making traceability and accountability difficult. Furthermore, the safety filters applied by companies show inconsistencies, sometimes allowing the inclusion of problematic symbolic content or unclassified hate speech, which compromises the neutrality and security of virtual assistants. This article is relevant to the open data community because it highlights the urgent need for open, clean, and well-documented datasets that respect the lawfulness of the content. It underscores that the quality of artificial intelligence depends directly on the integrity of training data, driving the demand for clear ethical standards, transparency in data usage, and robust governance mechanisms to ensure that technological development does not reinforce social inequalities or infringe upon fundamental rights.
Source: semana.comPublished on 2023-04-22
Related news
- La primera alternativa de código abierto a ChatGPT ya está lista: así es StableLM
- Las páginas web que utiliza ChatGPT para generar sus respuestas
- WHO releases largest global collection of health inequality data
- Federal agency in charge of Haskell takes nearly a year to fulfill FOIA request — and redacts most of it
- Stack Overflow joins Reddit and Twitter in charging AI companies for training data