OpenAI lanza GPTBot: el rastreador que recopilará datos públicos de Internet para entrenar modelos de inteligencia artificial

OpenAI has introduced GPTBot, a web crawler designed to access public content in order to improve the accuracy and safety of its future artificial intelligence models. Although the company guarantees that it will filter out confidential or copyrighted information, this capacity for mass data collection has sparked controversy among developers and users concerned about the non-consensual use of creative materials. The deployment of this system poses significant challenges for web data governance. Operating under a logic of implicit permission, where access is granted unless explicitly refused, GPTBot challenges traditional standards of privacy and intellectual property. This reflects the growing tension between the need for large volumes of data for AI training and the respect for the rights of original content creators. This case is crucial for the open data movement, as it illustrates the ethical and technical limits of automated crawling. It highlights the urgency of establishing clear frameworks that balance technological innovation with the protection of information, underscoring how AI tools can redefine access to and the utility of public data without compromising digital ethics.

Source: 20minutos.es
Published on 2023-08-10