ChatGPT can leak training data, violate privacy, says Google's DeepMind

Researchers from Google DeepMind demonstrated that the security alignment protocols designed to keep generative AI models safe can be bypassed through a simple repetition attack. By prompting the model to endlessly repeat specific single words, they forced it to diverge from its helpful assistant persona and revert to a base memorization mode. This technique effectively stripped away the safety layers, causing the AI to regurgitate verbatim passages from its original training data, including entire poems and literary excerpts. Beyond mere text duplication, the study revealed serious privacy implications as the model leaked personally identifiable information. The attack successfully extracted sensitive details such as names, phone numbers, and addresses, proving that the system retains and can disclose private data despite alignment efforts. This "extractable memorization" highlights a critical vulnerability where the underlying mechanism of the AI, which compresses and decompresses text, remains accessible if the guardrails are sufficiently destabilized by specific input patterns. This finding is vital to the open data community because it exposes the inherent risks of training models on large-scale, publicly available datasets. It challenges the assumption that alignment alone is sufficient to protect privacy, suggesting that even open or public data used for training can remain embedded in the model and susceptible to extraction. Consequently, developers and data stewards must consider not just how data is used, but how securely it can be prevented from being memorized and later revealed by adversarial queries.

Source: zdnet.com
Published on 2023-12-05