Clearing Rights For A ‘Non-Infringing’ Collection Of AI Training Media Is Hard

The emergence of curated, "non-infringing" datasets for AI training highlights the significant practical challenges of verifying copyright status at scale. While aggregating millions of public domain and Creative Commons Zero (CC0) images seems like a straightforward solution to liability concerns, real-world implementation reveals deep complexities. Databases like Source.Plus demonstrate that even when metadata suggests an image is free to use, verifying the underlying license requires navigating conflicting sources, changing platform policies, and identifying outdated contributors. The fragility of these claims becomes apparent when tracing specific images through multiple platforms. An image might appear as CC0 on a repository due to historical upload rules that no longer apply, while current licensing terms restrict commercial use. This discrepancy illustrates that automated aggregation often fails to capture the dynamic nature of licensing, leaving developers with a false sense of security. The risk is not just technical but legal, as relying on potentially inaccurate metadata can create liability gaps that are difficult to resolve. This article is crucial to the open_data community because it underscores the necessity of rigorous data provenance and transparent licensing information. As AI developers seek to rely on openly licensed data, they must recognize that "open" does not automatically mean "safe" or "clearly licensed." The case emphasizes that sustainable open data practices require more than simple aggregation; they demand robust verification mechanisms and clear communication about license evolution to ensure ethical and legal compliance in large-scale AI training.

Source: techdirt.com
Published on 2024-06-01