OpenAI's GPTBot and other AI web crawlers are being blocked by even more companies now

A significant shift is occurring in web data accessibility as major online platforms increasingly block AI web crawlers to protect their proprietary information. This resistance against automated data scraping reflects a growing corporate anxiety regarding the unauthorized use of web content for training generative artificial intelligence models. Prominent websites are actively deploying technical measures to restrict access to these bots, which feed into large language models. Consequently, the volume of available public data for AI development is shrinking, forcing developers to rely more heavily on licensed or proprietary datasets rather than freely scraped web content. This trend is highly relevant to open data because it challenges the assumption that the public web remains a shared, accessible resource for information reuse. As more sites impose restrictions, the ideal of universal data availability faces practical barriers, potentially limiting transparency and innovation that depend on unrestricted access to digital information.

Source: businessinsider.com
Published on 2023-09-29