Las páginas web que utiliza ChatGPT para generar sus respuestas
The article reveals that ChatGPT relies on the C4 dataset, which aggregates content from millions of websites. Crucially, it highlights that a vast amount of this information remains protected by copyright, raising significant legal concerns about the models' training processes and potential infringement issues. This situation underscores the ethical and legal challenges inherent in using openly accessible web data for artificial intelligence. It questions the transparency of data sourcing and the responsibility of developers regarding intellectual property rights when building commercial and educational tools. This is highly relevant to open data because it exposes the gap between data availability and usage rights. It emphasizes that open access does not equate to free use, prompting the open data community to advocate for better licensing frameworks and ethical standards to protect creators while enabling innovation.
Source: unocero.comPublished on 2023-04-22
Related news
- Feds Still Fighting Release of J6 Tapes Despite Mounting Legal Pressure A consortium of major media companies is suing the Justice Department and the FBI for ignoring Freedom of Information Act requests to obtain the still-secret recordings of January 6. By Julie Kelly
- Censo 2023: preocupación en organizaciones por el pedido de cédula de identidad y la “protección de datos personales”
- Stack Overflow joins Reddit and Twitter in charging AI companies for training data
- Empresas: pilas con la privacidad de los datos personales