AI companies are reportedly still scraping websites despite protocols meant to block them
Major AI companies, including Perplexity, OpenAI, and Anthropic, are reportedly bypassing robots.txt protocols to scrape content for training, raising significant ethical concerns. Despite public claims of respecting these exclusion signals, investigations reveal that these firms and their third-party crawlers ignore webmasters’ instructions to restrict access. This deliberate disregard undermines the voluntary standards developers have used since 1994 to control data ingestion, highlighting a gap between corporate public relations and actual operational practices. The implications for content creators are severe, as their intellectual property is being harvested with minimal attribution and occasional inaccuracies. Perplexity’s leadership argues that the protocol lacks legal force, suggesting a need for new licensing frameworks. However, this stance effectively permits unchecked data extraction, leaving publishers without recourse. The situation exposes a power imbalance where AI developers prioritize access over consent, potentially devaluing original journalism and creative works that fuel these technologies. This incident is crucial for open data discourse because it illustrates the conflict between open web principles and proprietary data monopolies. While open data advocates support free information flow, this case demonstrates how "open" scraping can violate explicit site constraints, challenging the notion of a truly open internet. It forces a reevaluation of how digital rights are managed, emphasizing that technical accessibility does not equate to ethical or legal permission. Ultimately, it underscores the urgent need for clear, enforceable standards regarding data usage in the age of generative AI.
Source: engadget.comPublished on 2024-06-23