‘Data poisoning’: How artists are fighting back against Artificial Intelligence image generators
This article explains how "data poisoning" serves as a defensive mechanism for artists against unauthorized AI scraping, primarily through tools like Nightshade. By subtly altering images to confuse machine vision while remaining visually normal to humans, creators can corrupt the training datasets of generative AI models. This results in unpredictable outputs, effectively forcing tech companies to respect copyright and intellectual property rights by making unethical data collection less viable. The relevance to open data lies in the critical tension between the open-scraping culture and data integrity. Open data initiatives often rely on vast, uncurated public datasets, but this article demonstrates how such practices can introduce significant quality and ethical issues. It highlights that open access does not equate to open consent, showing how unchecked data harvesting undermines the trust and reliability of open-source machine learning resources. Consequently, the piece argues for stricter governance regarding data provenance rather than relying solely on technical fixes like ensemble modeling or audits. It suggests that data poisoning is a legitimate response to moral intrusions, challenging the assumption that online data is free for unrestricted use. This underscores the need for open data frameworks to incorporate ethical sourcing standards and user rights, ensuring that openness does not come at the expense of creator rights or system stability.
Source: scroll.inPublished on 2023-12-28
Related news
- The New York Times wants OpenAI and Microsoft to pay for training data | TechCrunch
- The NYT demanda a OpenAI y Microsoft por infracción a derechos de autor, al utilizar sus artículos para entrenar chatbots
- 'The New York Times' demanda a OpenAI y Microsoft por infracción de derechos de autor para entrenar la IA