Google will let publishers hide their content from its insatiable AI
Google has introduced a new robots.txt directive allowing publishers to opt out of training generative AI models like Bard. This development highlights the growing tension between AI developers and content creators regarding the ethical use of web data. By granting explicit control over data ingestion, Google acknowledges the need for transparency in how emerging technologies leverage intellectual property, setting a potential industry standard for AI providers. The initiative addresses publisher concerns that AI aggregation threatens business models by diverting traffic and ad revenue. Although AI tools cite sources, they summarize information, potentially reducing direct visits to original websites. This control mechanism offers a critical safeguard for publishers, ensuring they can protect their economic interests while participating in the digital ecosystem, thereby preserving the viability of content creation amidst rapid technological shifts. This update is vital for open data discussions as it demonstrates a shift toward consent-based data usage in AI development. It suggests that future open data frameworks may require granular user and publisher controls rather than blanket access. By enabling selective participation, Google is helping balance the need for robust AI training data with the rights of content owners, influencing how open information and commercial interests coexist in the age of generative AI.
Source: engadget.comPublished on 2023-09-29
Related news
- Challenges raised by AI regarding the right to information - WAN-IFRA
- OpenAI's GPTBot and other AI web crawlers are being blocked by even more companies now
- La Comisión de Transparencia recibe 794 reclamaciones, el mayor número desde su creación hace siete años
- Exposing the secretive company at the forefront of facial recognition technology
- A plan to codify FOIA in Arkansas