Websites can now block OpenAI's web crawling bot
OpenAI has responded to widespread criticism regarding its use of publicly available data for training ChatGPT by providing webmasters with clear mechanisms to opt out. The company acknowledges its reliance on scraped internet content but emphasizes that voluntary participation enhances model accuracy and safety. This approach balances the need for comprehensive training data with respect for individual and corporate preferences regarding their digital assets. By issuing specific guidance on utilizing the robots.txt protocol, OpenAI empowers website owners to block its crawler, GPTBot, from accessing their sites. This move offers greater control compared to existing alternatives, provided that the company strictly adheres to the exclusion rules it promotes. Transparency regarding the crawler’s identity and technical specifications is presented as a key step toward ethical data collection practices in the rapidly evolving AI landscape. This development is highly relevant to open data as it highlights the tension between the open nature of the web and the proprietary demands of commercial AI models. It underscores the growing necessity for standardized, transparent frameworks that allow content creators to maintain sovereignty over their contributions. Furthermore, OpenAI’s alignment with voluntary safety commitments suggests a shift toward more responsible engagement with the open data ecosystem, potentially influencing how future AI systems interact with publicly accessible information.
Source: techspot.comPublished on 2023-08-10
Related news
- How to block OpenAI's new AI-training web crawler from ingesting your data
- OpenAI lets websites block GPTBot, a web crawler that looks for AI training data
- Organizaciones de medios de comunicación piden negociar uso de contenidos para inteligencia artificial
- Google Takes Wildly Different Stances on AI "Deepfakes" and Web Scraping