AI Chatbots are scraping news reporting and copyrighted content, News Media Alliance says - KTVZ
A major news media trade group has issued a stark warning to artificial intelligence developers, asserting that they have extensively scraped copyrighted news content to train their generative chatbots without authorization or compensation. The research highlights that these AI systems rely disproportionately on high-quality journalistic material, demonstrating an implicit recognition of its unique value while simultaneously undermining the economic sustainability of the publishers who create it. This dynamic threatens not only the news industry’s survival but also the long-term reliability of AI models that depend on trustworthy data sources. The industry argues that the common defense of AI “learning” facts like humans is flawed, as the models retain protected expressions rather than just underlying concepts. In response, many publishers have implemented technical barriers to prevent future scraping, but these measures do not address the massive amounts of content already used for training. Consequently, the trade group is urging policymakers to legally classify this unauthorized use as infringement, emphasizing that current voluntary norms are insufficient to protect intellectual property in the age of machine learning. This issue is critical to open data because it challenges the assumption that publicly available information is free to use for commercial AI development. It highlights the urgent need for clear ethical and legal frameworks that balance data accessibility with creators’ rights. For open data advocates, this case underscores the necessity of ensuring that datasets used for training AI are obtained through proper licensing, preserving both the integrity of human-generated knowledge and the sustainability of the media ecosystem that produces it.
Source: ktvz.comPublished on 2023-11-02