Periódicos demandan a Microsoft y OpenAI por infringir derechos de autor con su IA

This case highlights the critical tension between AI development and journalistic copyright, directly impacting open data discourse by questioning the legality of scraping protected web content for training models. The lawsuit argues that major AI providers utilize unauthorized, massive collections of journalistic work, challenging the current open data paradigm that often relies on publicly accessible information without explicit licensing or compensation. The plaintiffs assert that generative systems not only reproduce copyrighted material verbatim but also fabricate damaging misinformation attributed to these news outlets. This behavior underscores the risks of unregulated data consumption, suggesting that relying on scraped public data without strict ethical boundaries can lead to harmful outputs and reputational damage for content creators, thereby necessitating stricter governance in open data usage. Ultimately, this legal battle is relevant to open data because it forces a re-evaluation of data provenance and consent. It emphasizes that "open" access does not equate to unrestricted commercial use, pushing the community toward more transparent, licensed, and ethical frameworks for acquiring and utilizing data to train artificial intelligence systems responsibly.

Source: larepublica.co
Published on 2024-05-03