Column: It's not just Zoom: How websites and apps harvest your data to build AI

The recent controversy surrounding Zoom’s terms of service highlights a critical shift in how digital platforms handle user data, revealing that personal information from intimate communications is increasingly being considered for artificial intelligence training. This incident underscores a broader, often invisible, reality: nearly all public online content, from social media posts to professional blogs, is being scraped and utilized to train large language models. The backlash demonstrates that users are no longer passive recipients of data extraction but are actively demanding transparency regarding how their digital footprints are commodified and repurposed by tech giants. This landscape presents a pivotal moment for the open data community to renegotiate the social and economic contracts governing data usage. As AI companies aggressively ingest vast amounts of existing online material, the default assumption that public visibility implies consent for commercial exploitation is being challenged. The article argues that this disruption offers a unique opportunity to establish new standards for consent, forcing a reevaluation of what is socially and ethically acceptable in data harvesting. It suggests that without proactive resistance and clear regulations, individual autonomy will continue to erode under the weight of automated data processing. To safeguard privacy, the article recommends moving away from platforms with opaque data policies toward services built with encryption and explicit user consent, such as Signal or Proton Mail. For those maintaining public websites, technical safeguards like robots.txt can help restrict unauthorized scraping. Ultimately, this narrative is vital to open data because it frames the current AI boom not merely as a technological advancement, but as a civil liberties issue. It calls for a fundamental shift toward a "doctrine of consent," urging developers and users alike to prioritize privacy-by-design and to actively resist the unchecked accumulation of personal data for AI development.

Source: latimes.com
Published on 2023-08-17