'Millions' of NYT and NY Daily News stories taken by OpenAI for training data
The core issue in these ongoing copyright lawsuits is the fundamental disagreement over who bears the responsibility for identifying copyrighted material within OpenAI’s massive training datasets. While publishers have discovered millions of their works through their own searches, they argue that as the entity that collected the data, OpenAI should be compelled by court order to explicitly identify and admit which specific content was used to train its models. Publishers contend that this transparency is foundational to establishing the scope of infringement, asserting that the tech company is uniquely positioned to provide this information rather than forcing rights holders to conduct exhaustive, costly investigations. This dispute highlights the significant practical and financial burdens placed on plaintiffs navigating the discovery phase of artificial intelligence litigation. News publishers have spent substantial resources and time inspecting the data in a controlled environment, yet they report that the process remains inefficient, expensive, and technically fraught. By demanding that OpenAI take the lead in cataloging its own usage of copyrighted material, plaintiffs aim to reduce the disproportionate costs and delays associated with manually trawling through hundreds of terabytes of unstructured text, shifting the logistical burden back to the AI developer. This conflict is highly relevant to the open data community as it sets a critical precedent for data transparency and accountability in machine learning. It challenges the opacity of proprietary training processes and underscores the tension between the technical feasibility of large-scale data inspection and the legal necessity of proving copyright infringement. The outcome will likely influence how future disputes handle the intersection of intellectual property rights and AI development, potentially forcing greater openness regarding how models are built and what public or copyrighted data is consumed without explicit consent.
Source: pressgazette.co.ukPublished on 2024-11-07