Meta admite haber usado libros con copyright para entrenar a su IA

The article highlights a significant ethical and legal controversy involving Meta’s use of the Books3 dataset, which contains thousands of copyrighted books, to train its Llama models. This admission has sparked intense debate regarding the rights of original authors and the practices of major tech firms. The core issue lies in whether using such extensive proprietary material without compensation or explicit authorization constitutes fair use or infringes on intellectual property rights, thereby challenging the current legal frameworks governing artificial intelligence development. This situation reflects a broader industry pattern, with companies like OpenAI and Microsoft facing similar accusations for relying on copyrighted content to advance their models. The tension arises because the rapid evolution of AI technology often outpaces existing legislation, creating a gray area where the necessity of vast training data conflicts with copyright protections. The refusal by these corporations to compensate authors underscores a critical disconnect between technological innovation and legal responsibility, prompting urgent questions about the sustainability of current training methodologies. This case is highly relevant to open data because it challenges the fundamental assumption that data for AI development must remain closed or unauthorized. It urges the open data community to advocate for transparent, legally compliant, and ethically sound data sourcing practices. Establishing clear precedents for author compensation and data consent is essential to balance innovation with respect for creators, ensuring that the future of AI respects intellectual property rights rather than exploiting them.

Source: wwwhatsnew.com
Published on 2024-01-14