Big Tech in ‘underground’ race to license archives that will train Artificial Intelligence

Photobucket is pivoting to license its vast archive of billions of photos and videos to generative AI developers, signaling a shift from free web scraping to paid data acquisition. This move highlights the urgent demand for high-quality, copyright-compliant datasets as tech giants seek to train AI models without facing litigation. The article illustrates a burgeoning market where private data holders can monetize legacy content. By negotiating significant fees for access to their archives, these entities are capitalizing on the scarcity of non-scrapable data, demonstrating how historical digital assets are gaining value in the current AI landscape. This case is relevant to open data because it underscores the tension between accessible information and proprietary control. As the AI industry increasingly relies on licensed, closed datasets, it challenges the principles of openness by prioritizing exclusive access over free sharing, potentially limiting the availability of training data for open-source and community-driven AI initiatives.

Source: thehindu.com
Published on 2024-04-07