Google Wants AI Scraping to Be 'Fair Use.' Will That Fly in Court?

The article highlights a critical tension between major tech companies and copyright holders regarding the use of protected data to train AI models. Google is actively seeking to codify the right to scrape vast amounts of copyrighted content for training, arguing that such use constitutes fair use through an opt-out mechanism. This approach shifts the burden onto creators to identify and block their data, effectively challenging established property norms where protection exists unless explicitly waived. From an open data perspective, this debate defines the boundaries of the "training data" commons. If courts rule that ingestion for machine learning is fair use, it legitimizes the extraction of public and protected information without compensation or permission, potentially starving original creators of revenue while empowering AI aggregators. The outcome will determine whether the open web remains a shared resource for innovation or becomes a proprietary feed for commercial AI products, fundamentally altering how data flows and is valued online. The implications extend beyond legal precedent to the sustainability of digital publishing. If AI models can generate high-quality answers directly from scraped content, they may cannibalize traffic to source websites, undermining the economic model of the open web. This scenario poses a threat to the diversity and availability of high-quality information, as creators may withdraw from the digital space or restrict access, thereby shrinking the pool of open data available for future research and development.

Source: tomshardware.com
Published on 2023-08-12