OpenAI Conceals Training Data Sources, Including J.K. Rowling's Harry Potter Series, for ChatGPT
Recent research highlights a growing tension between artificial intelligence development and intellectual property rights. Studies show that major language models, including ChatGPT, inadvertently reproduce copyrighted text from their extensive training datasets. This "copyright leakage" has sparked legal challenges from authors and prompted tech companies to adopt less transparent practices regarding their training data sources to mitigate infringement risks. The findings reveal that while developers are actively implementing safeguards to prevent verbatim output, the underlying issue persists. Despite efforts to align model outputs and mask training origins, these AI systems still generate text closely resembling protected works. The research indicates that leakage is fundamentally linked to the presence of copyrighted material in the training data itself, suggesting that superficial output adjustments cannot fully eliminate the problem. This article is highly relevant to the open data community as it underscores the critical importance of data provenance and licensing in AI ethics. It demonstrates that relying on opaque, unlicensed internet text creates significant legal and ethical vulnerabilities for both developers and content creators. The study advocates for more robust training methodologies and greater transparency, encouraging the open data movement to prioritize legally compliant and ethically sourced datasets to foster sustainable innovation.
Source: techstory.inPublished on 2023-08-23
Related news
- Scraping or Stealing? A Legal Reckoning Over AI Looms
- NVIDIA DLSS 3.5 uses a mountain of data and AI to improve ray tracing on RTX GPUs
- Rechazan petición de proteger con derechos de autor obra de arte hecha con IA en EE. UU. - Pulzo
- La Comunitat Valenciana se posiciona como la región con más viviendas vacías de España
- Drug deaths in Scotland down for second year in a row - but rich-poor gap widens