Major sites opt out of Apple’s AI scraping

Major publishers are excluding their data from Apple’s AI training, signaling a critical shift in how web content is harvested. By updating robots.txt files, organizations assert control over their intellectual property against automated crawlers. This action highlights the growing conflict between tech giants and content creators regarding AI training rights. The robots exclusion protocol, a long-standing web standard, is now the battleground for defining boundaries in digital data usage. This relevance to open data lies in the challenge of maintaining transparent, accessible information sources. As more entities restrict bot access, the comprehensiveness of publicly available datasets may diminish, impacting research and innovation reliant on open web data.

Source: macdailynews.com
Published on 2024-08-31