How to block OpenAI's new AI-training web crawler from ingesting your data
OpenAI has introduced GPTBot, a dedicated web crawler designed to systematically gather data for training its large language models. This release signifies a shift toward transparent data acquisition practices, offering developers clear instructions on how to permit or restrict access. By providing specific tools for website owners, OpenAI aims to balance the need for extensive training data with respect for publisher preferences, moving away from opaque scraping methods. The company emphasizes that allowing GPTBot access contributes to improving AI accuracy, safety, and general capabilities. OpenAI asserts that it employs filters to exclude paywalled content, personally identifiable information, and policy-violating text. This approach suggests an effort to legitimize data sourcing by filtering out sensitive or restricted materials, thereby addressing some ethical concerns while maintaining the volume of data necessary for model improvement. This development is highly relevant to open data because it highlights the growing tension between public information availability and proprietary AI training needs. As lawsuits concerning copyright and data theft increase, OpenAI’s move to offer blocking mechanisms via robots.txt represents a step toward negotiated data access rather than unilateral extraction. It underscores the critical importance of standardizing web crawling protocols to protect open sources while enabling technological advancement in an era where data scarcity and consent are major legal and ethical challenges.
Source: zdnet.comPublished on 2023-08-10
Related news
- Websites can now block OpenAI's web crawling bot
- OpenAI lets websites block GPTBot, a web crawler that looks for AI training data
- Organizaciones de medios de comunicación piden negociar uso de contenidos para inteligencia artificial
- Google Takes Wildly Different Stances on AI "Deepfakes" and Web Scraping