Meta has confirmed that it utilized publicly available posts from Facebook and Instagram to train its new AI virtual assistant, while deliberately excluding private messages and personal datasets to protect user privacy. By filtering out content heavily laden with personal information and opting out of using platforms like LinkedIn, the company aims to balance innovation with consumer trust. This selective approach highlights the ongoing industry effort to distinguish between open, publicly shared data and sensitive private information during model development. The reliance on public data underscores a significant ethical and legal tension within the tech sector, where major companies face lawsuits over copyright infringement for using scraped internet content without permission. Meta acknowledges that litigation is likely as stakeholders debate whether training AI on copyrighted creative works constitutes fair use. Unlike competitors who have signed licensing deals for specific content libraries, Meta relies on existing terms of service and safety restrictions to mitigate intellectual property violations, reflecting the unresolved nature of data ownership in the age of generative AI. This development is crucial to the open data discourse because it illustrates how tech giants navigate the complex boundary between using publicly accessible data for commercial AI training and respecting individual privacy rights. It demonstrates that while data for training may be technically "open," the methods used to curate, filter, and utilize that data have profound implications for digital rights and legal frameworks. Understanding these practices is essential for policymakers and advocates working to define transparent and equitable standards for data usage in artificial intelligence development.

Source:
Published on 2023-09-30