Novelists Sue OpenAI for Copyright Infringement Over Books Used as Training Data

This article highlights a landmark legal challenge to generative AI, marking the first major class-action lawsuit concerning the use of text-based copyrighted works for training large language models. Novelists Paul Tremblay and Mona Awad have sued OpenAI, alleging that their books were used without consent to train GPT-3.5 and GPT-4. The plaintiffs argue that because the AI models can reproduce detailed plot summaries and specific information from their works, the systems themselves constitute infringing derivative works, thereby violating copyright law and the Digital Millennium Copyright Act. The significance of this case extends beyond individual compensation, as it challenges the foundational "fair use" doctrine currently relied upon by the AI industry. Unlike previous rulings regarding search engine caching, which was deemed transformative, this lawsuit contends that using creative texts to train AI harms the market for original works and lacks true transformation. The outcome will determine whether scraping publicly available text for model training is legally permissible, potentially forcing companies to alter their data collection and processing practices. This litigation is highly relevant to the open data movement because it scrutinizes the ethics and legality of using publicly accessible information to build commercial artificial intelligence. It underscores the tension between the unrestricted availability of digital data and the intellectual property rights of creators. The case may establish new precedents for how open data can be utilized in machine learning, potentially imposing stricter requirements for attribution, consent, or compensation when public domain or openly shared content is repurposed for commercial AI development.

Source: tomshardware.com
Published on 2023-07-01