The New York Times is the latest to go to battle against AI scrapers
The era of unrestricted access to open internet data for training generative AI is rapidly coming to an end. Major publishers like The New York Times are explicitly banning their content from being scraped to train future AI models, shifting the paradigm from free, open access to controlled, often compensated licensing. This closure of data sources signals a critical turning point where the foundational resources for AI development are becoming proprietary rather than universally available, forcing technology companies to seek new, legal pathways to obtain high-quality information. This shift highlights a growing conflict between the AI industry’s data hunger and content creators’ demands for ownership and revenue. As seen with other platforms restricting API access, media organizations are asserting control over their intellectual property to ensure they are compensated for the value their work provides to AI developers. The focus of the debate is moving away from retrofitting existing models, which already contain scraped data, toward securing rights for future iterations, thereby establishing a market where data is treated as a licensable asset rather than a public good. For the open data community, this trend poses a significant threat to the ethos of free and accessible information. If the primary sources of high-quality text and media become gated behind paywalls or exclusive licensing agreements, the ability to develop transparent and unbiased AI systems diminishes. This development underscores the urgent need for clear legal frameworks that balance innovation with creator rights, as the current uncertainty could stifle the collaborative spirit essential to the open data movement and lead to a fragmented digital landscape.
Source: popsci.comPublished on 2023-08-17
Related news
- Anti-Piracy Group Takes Prominent AI Training Dataset ''Books3' Offline * TorrentFreak
- Tech Thoughts: Media needs a united front against data scraping to train AI
- Column: It's not just Zoom: How websites and apps harvest your data to build AI
- UMG, Sony, and Capitol launch copyright lawsuit against Internet Archive