Google confirma usar datos públicos en la web para entrenar sus modelos de IA

Google has officially updated its Privacy Policy to include the use of publicly available information on the Internet for training its artificial intelligence models. This transparency marks a turning point by explicitly validating practices that previously operated in a regulatory vacuum, directly impacting how current generative AI tools are built and establishing a corporate precedent regarding the exploitation of open data. However, this practice generates significant tensions concerning copyright and the protection of digital creativity. Unlike more respectful approaches that prioritize content with explicit licenses, Google’s strategy faces accusations of being a “plagiarism engine” that diverts traffic away from original creators. The lack of clarity about how protected content or content subject to anti-scraping policies is handled raises urgent ethical and legal dilemmas for platforms that rely on public resources. This case is relevant to the open data community because it exposes the conflict between the technical availability of data and the ethical sustainability of its use. By legitimizing the scanning of the public web for training purposes, Google normalizes the mass extraction of freely available knowledge, which could discourage open publishing if creators are neither compensated nor protected. The resolution of these disputes will determine whether open data remains a common resource or becomes raw material exploited without clear consent.

Source: wwwhatsnew.com
Published on 2023-07-06