Medium has adopted a strict default policy against allowing its content to be used for training artificial intelligence models without explicit consent. This shift aims to protect writers from the dual injustice of unpaid data harvesting and the potential displacement of their livelihoods by AI-generated text. By updating its robots.txt file, the platform instructs major AI providers like OpenAI to refrain from crawling its pages, signaling a broader industry movement to assert control over creative intellectual property in the age of large language models. However, the article highlights the significant limitations of technical blocking measures. While platforms like OpenAI and Google generally respect robots.txt instructions, smaller or less cooperative crawlers may ignore these signals, turning enforcement into a tedious game of "whack-a-mole." Consequently, Medium combines these digital barriers with legal threats, promising cease and desist letters to unauthorized scrapers. This approach underscores the fragility of relying solely on technical protocols to protect digital content, as many data collectors operate outside the established norms of these voluntary standards. This situation is highly relevant to open data because it illustrates the growing tension between the unrestricted access required for massive AI training and the legal rights of content creators. It challenges the prevailing assumption that publicly available web data is freely usable for any purpose, including commercial AI development. As Medium prepares to potentially re-opt into AI training once fair compensation protocols are established, the platform serves as a critical case study for how open data ecosystems might need to evolve to include attribution and financial reciprocity, rather than relying on the default "scrape everything" mentality.
Source:Published on 2023-09-30
Related news
- Google Will Enable Web Admins To Block Systems from Scraping Sites for AI Training
- Google Now Allowing Website Owners to Opt Out of Google Bard
- Books 3 has revealed thousands of pirated Australian books. In the age of AI, is copyright law still fit for purpose?
- Here's how web publishers can opt out of Google crawlers scraping website data to train AI models | MediaNama
- OpenAI offers a way for creators to opt out of AI training data. It's so onerous that one artist called it 'enraging.'