¿De dónde sale la información que da ChatGPT?

The research reveals that artificial intelligence models are trained on vast amounts of web data containing racial and religious biases, as well as copyrighted content, posing significant ethical and legal risks. When inadequately filtered datasets are used, these technologies can perpetuate prejudices and facilitate the spread of propaganda or misinformation, while exposing creators and users to violations of privacy and intellectual property rights. A critical aspect is the lack of transparency regarding the origin of the information, since a considerable portion of the websites used does not appear publicly on the current web, making traceability and accountability difficult. Furthermore, the safety filters applied by companies show inconsistencies, sometimes allowing the inclusion of problematic symbolic content or unclassified hate speech, which compromises the neutrality and security of virtual assistants. This article is relevant to the open data community because it highlights the urgent need for open, clean, and well-documented datasets that respect the lawfulness of the content. It underscores that the quality of artificial intelligence depends directly on the integrity of training data, driving the demand for clear ethical standards, transparency in data usage, and robust governance mechanisms to ensure that technological development does not reinforce social inequalities or infringe upon fundamental rights.

Source: semana.com
Published on 2023-04-22