Researchers reported that large language models can be manipulated into disclosing memorized training data, highlighting a persistent privacy risk in generative AI systems. An academic study on extracting training data from language models demonstrated that models may reproduce sensitive or rare sequences from their training corpus, while later testing against public generative AI services found that safeguards could sometimes be bypassed to reveal email addresses from the Enron dataset.
The testing described prompt-based techniques including jailbreak-style requests, chain-of-thought prompting, and multiple-choice validation to elicit data that models initially refused to provide. The findings underscore that once information is embedded in model weights it may be difficult to remove, and they reinforce recommendations to avoid submitting sensitive data to public models, disable training on user conversations where possible, and apply layered controls such as prompt-injection detection, privacy-preserving intermediaries, and output filtering.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
2 events from the most recent confirmed update back to the earliest known activity.
In 2023, the author tested ChatGPT and found that combining jailbreak prompting, chain-of-thought prompting, and multiple-choice validation could induce disclosure of Enron employee email addresses and names despite initial privacy refusals. The article says the same general method also caused Google Bard, Microsoft New Bing, and Stability AI StableVicuna to disclose Enron employee email addresses after prompt manipulation.
A 2021 paper titled "Extracting Training Data from Large Language Models" reported that GPT-2 could emit memorized training data, including personal information such as names, phone numbers, email addresses, chat histories, source code, and UUIDs. The later blog cites this work as evidence that large language models can disclose sensitive data learned during training.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.