The GDPR significantly impacts AI development by establishing strict conditions under which personal data must be deleted upon request. Primarily, if a company relies solely on user consent or legitimate interest to train AI models, individuals can effectively revoke access to their data. This creates a complex challenge for developers, as publicly sourced data often lacks explicit consent, forcing organizations to demonstrate overriding legitimate interests to retain training datasets against erasure demands. The right to be forgotten is not absolute and allows for exceptions when data is needed for legal compliance or defense. However, framing deletion as an individual right exposes companies to civil liability, complicating the balance between data privacy and AI innovation. This legal ambiguity means that even if processing is technically lawful, the threat of lawsuits may pressure organizations to remove data, potentially undermining the integrity and volume of training sets required for robust machine learning systems. This article is relevant to open data because it highlights the tension between public data accessibility and individual privacy rights. As open data initiatives increasingly utilize publicly available information, they must navigate GDPR constraints that prioritize individual control over data utility. Understanding these erasure conditions is crucial for ensuring that open data ecosystems remain compliant while preserving the availability of information for scientific and commercial AI research.

Source:
Published on 2023-08-01