Google’s updated privacy policy explicitly allows the use of publicly available information to train AI models like Bard and Cloud AI. This marks a significant shift from previous language model restrictions, expanding the scope of data ingestion to include any online content. Consequently, users’ public posts can now be harvested to develop new products, broadening the dataset beyond direct platform interactions to the entire public web. This change implies that information previously protected by platform-specific privacy boundaries is now accessible for training generative AI. It highlights a growing tension between service improvement and user privacy, as data sharing extends beyond the platform to the wider internet. Users must recognize that public digital footprints are no longer isolated, affecting how personal data is perceived and managed in an increasingly AI-driven landscape. This update is relevant to open data because it demonstrates how corporations leverage openly accessible information to build proprietary technologies. It raises critical questions about consent, attribution, and the ethical boundaries of training data when vast amounts of public knowledge are ingested. Understanding this shift helps researchers and users navigate the complex intersection between open information ecosystems and commercial AI development.
Source: trak.inPublished on 2023-07-12