An AI engine scans a book. Is that copyright infringement or fair use?

The central tension lies between AI developers claiming copyright fair use for training large language models and creators demanding consent and compensation. While tech firms argue that digesting data to learn language patterns is a transformative, non-expressive act similar to search indexing, authors contend that these tools create market substitutes for original works without attribution or payment. This conflict highlights a critical disconnect in how creative value is recognized in the digital age, threatening the sustainability of the creative economy if consent and credit are ignored. Relevance to open data is profound, as this legal debate dictates the accessibility and quality of the foundational datasets used to build public AI tools. If courts rule against fair use, the pool of legally scrapable public knowledge could shrink significantly, forcing reliance on proprietary or restricted datasets that favor large corporations. This risks creating a monopolized information ecosystem where only well-funded entities can afford to train models, undermining the open, democratic potential of data-driven innovation. Ultimately, the outcome will determine whether AI development remains rooted in open public resources or shifts toward closed, licensed ecosystems. Establishing a legal balance that protects creators’ rights while preserving access to diverse data sources is essential. Without this equilibrium, the open data movement may face severe restrictions, limiting transparency, reducing model diversity, and exacerbating inequalities in who can participate in the AI revolution.

Source: cjr.org
Published on 2023-10-27