Amazon Is Investigating Perplexity Over Claims of Scraping Abuse

Amazon Web Services has launched an investigation into Perplexity AI, probing whether the startup violates cloud hosting rules by ignoring web standards that restrict data scraping. This inquiry stems from reports that Perplexity’s systems accessed content on major news sites, such as WIRED and The New York Times, despite these publishers explicitly forbidding automated access through the Robots Exclusion Protocol. While this protocol is not legally binding, AWS mandates its adherence to prevent abusive activities, making this a critical test of how cloud providers enforce ethical data usage standards among their customers. The situation highlights the growing tension between AI development and content ownership, as Perplexity claims its infrastructure respects these protocols but was found bypassing them via undisclosed servers. The startup attributes some of this activity to third-party crawlers and argues that ignoring restrictions is sometimes necessary for user queries, a stance that challenges traditional web norms. This conflict underscores the ambiguity in regulating AI data collection, where technical capabilities often outpace established legal and contractual frameworks for protecting publisher rights and server integrity. This case is highly relevant to open data because it exposes the fragile balance between free information access and the controlled sharing of digital assets. It raises questions about who controls data availability when private entities use powerful infrastructure to bypass public restrictions. The outcome will likely influence future policies on data scraping, potentially impacting how open datasets are sourced, preserved, and accessed in an era where AI companies increasingly rely on aggregated public information to train their models.

Source: wired.com
Published on 2024-06-28