Major sites opt out of Apple’s AI scraping
Major publishers are excluding their data from Apple’s AI training, signaling a critical shift in how web content is harvested. By updating robots.txt files, organizations assert control over their intellectual property against automated crawlers. This action highlights the growing conflict between tech giants and content creators regarding AI training rights. The robots exclusion protocol, a long-standing web standard, is now the battleground for defining boundaries in digital data usage. This relevance to open data lies in the challenge of maintaining transparent, accessible information sources. As more entities restrict bot access, the comprehensiveness of publicly available datasets may diminish, impacting research and innovation reliant on open web data.
Source: macdailynews.comPublished on 2024-08-31
Related news
- Facebook, Instagram opt out of allowing Apple AI to scrape their data for training
- Bolivia supera los 11 millones de habitantes y Santa Cruz es el departamento más poblado
- Study: Transparency is often lacking in datasets used to train large language models
- Child abuse images removed from AI image-generator training source, researchers say