LinkedIn has triggered significant backlash by utilizing user-generated content to train its generative AI features without obtaining explicit prior consent. This practice, characterized as scraping first and addressing permissions later, has led to widespread accusations of betraying user trust. The company updated its privacy policies to allow this data collection, offering only a retrospective opt-out mechanism rather than seeking initial permission, a strategy that mirrors broader industry trends in big tech’s development of large language models. The implications for data privacy are profound, as personal information, posts, and feedback are now processed to improve AI capabilities. While the platform employs privacy-enhancing technologies to redact personal data, there remains a risk that private inputs could inadvertently appear in AI outputs. Notably, users in the European Union, the UK, Iceland, Norway, Liechtenstein, and Switzerland are exempt from this training, highlighting how regional regulations significantly influence corporate data handling practices and protect specific populations from such exploitation. This incident is highly relevant to the open_data movement because it underscores the critical tension between corporate AI ambitions and individual data sovereignty. It illustrates the necessity for transparent data governance, where users retain control over how their contributions are utilized. The subsequent regulatory intervention in the UK, which forced a pause on training with British user data, further emphasizes that voluntary corporate compliance is insufficient without robust legal frameworks ensuring that data sharing respects user rights and ethical standards.
Source:Published on 2024-09-21