The Guardian, New York Times, CNN figure in growing list of sites blocking OpenAI crawler
Major publishers are increasingly blocking AI crawlers to protect their intellectual property, asserting that unauthorized scraping violates terms of service. This collective action highlights a growing tension between media entities seeking commercial licensing partnerships and technology companies aiming to freely gather data for training large language models. The shift suggests a reevaluation of how digital content is valued and monetized in the age of artificial intelligence. The widespread adoption of these blocks indicates that nearly one-fifth of the world’s top websites are resisting unregulated data extraction. While some platforms continue to support open data use for AI development, the resistance from diverse sectors demonstrates a clear demand for consent and compensation. This trend challenges the assumption that publicly available web content is inherently free for commercial AI training without permission. This development is crucial for open data advocates as it signals a potential retreat from the unrestricted data harvesting that fueled recent AI advancements. It emphasizes the necessity for transparent data sourcing policies and raises questions about the legal frameworks governing digital information. Understanding this pushback is essential for navigating the future of open data, where access must balance innovation with the rights of content creators.
Source: rappler.comPublished on 2023-09-07
Related news
- OpenAI and Microsoft accused of stealing data to train ChatGPT in new class-action suit
- Technology Innovation Institute Introduces World’s Most Powerful Open LLM: Falcon 180B
- Right to Information Month launched to increase awareness
- Maya PH’s open-source LLM, Godzilla 2, surpasses ChatGPT in truthfulness
- ¿Cómo saber si estafadores lo tienen en la mira?