The article highlights a contentious debate surrounding data consent in AI development, centered on the tool img2dataset. Its creator argues that website owners must actively opt out of having their images scraped for training generative models, claiming that defaulting to consent is necessary to avoid blocking widespread technological benefits. This stance creates a significant burden on individual site operators, who must implement specific headers to protect their content, rather than developers seeking permission before harvesting data. This conflict underscores a critical ethical challenge for open data: the tension between the open nature of web data and the rights of creators. Traditional web standards, such as robots.txt, which allow site owners to control crawler access, are deliberately bypassed by this tool. Consequently, individuals bear the financial and technical costs of defending their sites against aggressive scraping, raising questions about whether non-consensual data harvesting constitutes ethical open-data practice or exploitative extraction. The relevance to open data lies in the precedent this sets for dataset creation. As major platforms restrict API access to prevent free scraping, the reliance on uncompensated, non-consensual web data becomes unsustainable and legally risky. This incident illustrates that true openness cannot ignore consent; it requires mechanisms that respect creator agency. Without shifting from extractive scraping to respectful, permission-based data sharing, the open data community risks fostering distrust and legal fragmentation, ultimately hindering the collaborative potential of shared knowledge.

Source:
Published on 2023-04-26