Microsoft, OpenAI, Sued Over Copyrighted Training Data For ChatGPT, AI Models

A proposed class action lawsuit alleges that OpenAI and Microsoft infringe copyright by training AI models like ChatGPT on nonfiction works without permission or compensation. The suit argues that this practice constitutes rampant theft, as the platforms profit from vast amounts of copyrighted data while denying authors licensing opportunities. This legal challenge highlights a critical tension in the development of generative AI, where technological advancement is perceived to come at the direct expense of intellectual property rights. Microsoft is specifically cited as a key defendant due to its provision of essential cloud infrastructure and its alleged deep involvement in the development and commercialization of these models. The lawsuit posits that Microsoft had knowledge of the unauthorized use of copyrighted material and benefits financially from the resulting technology. This connection underscores the shared liability of major tech investors and providers when supporting AI systems built on unlicensed data, raising questions about corporate accountability in the AI ecosystem. This case is highly relevant to the open data movement as it illustrates the ongoing conflict between the unrestricted use of data for training artificial intelligence and the legal protections afforded to creators. It challenges the assumption that publicly available data can be freely harvested for commercial AI purposes without addressing the rights of original authors. Ultimately, the litigation seeks to establish whether fair use applies to large-scale data scraping, potentially reshaping how open datasets are collected, used, and regulated in future AI development.

Source: techtimes.com
Published on 2023-11-23