Google confirma usar datos públicos en la web para entrenar sus modelos de IA

Google has officially updated its privacy policies to acknowledge that it uses publicly available information on the Internet to train its artificial intelligence models. This move aims to provide transparency about how products such as Bard and Google Translate are developed, aligning its practices with the realities of generative AI. The disclosure directly impacts the open data ecosystem, as it legitimizes the use of public content as raw material for advanced algorithms, raising ethical debates about intellectual property and consent in the digital age. However, the company faces significant criticism for not explicitly addressing discrimination against copyrighted materials or the anti-scraping policies of other websites. Unlike competitors such as Adobe, which prioritize respect for creators, Google uses data without clear distinctions, leading to tensions with publishers and platforms. This stance has resulted in lawsuits accusing Google of acting as a plagiarism engine and engaging in monopolistic practices, highlighting the urgent need for clearer boundaries in how open data is leveraged for commercial AI development. This article is relevant to the open data movement because it exposes the friction between the public availability of information and the corporate interests of major technology companies. By normalizing the use of open data for AI training without robust compensation or regulatory mechanisms, the sustainability of the concept of free information is called into question. The case underscores the need for legal and ethical frameworks that protect both transparency and the rights of original creators, which is essential for maintaining integrity and fairness in open data initiatives.

Source: eju.tv
Published on 2023-07-07