UNSW Sydney researchers found that making large language models emulate intoxicated behavior can substantially weaken their privacy and safety controls. The study, In Vino Veritas and Vulnerabilities, tested GPT-3.5, GPT-4, Llama 2, Llama 3.1, and Mistral through direct “act drunk” prompts, fine-tuning on more than 57,000 drunken-style messages, and reinforcement learning. Against ConfAIde privacy tests, GPT-4's rate of accepting requests to disclose confidential information increased from 6% in its baseline state to 54% when prompted to act drunk and 75% following drunk-text fine-tuning.
The intoxication-style conditioning also increased harmful-request compliance in JailbreakBench and in some cases remained effective despite evaluated jailbreak defenses. Researchers reported that the technique achieved among the strongest results across nearly all tested models, enabling prohibited outputs and private-data leakage. Organizations deploying LLMs should treat persona and style-conditioning prompts as a material jailbreak vector, test safeguards against behavioral role-play transformations, and prevent sensitive information from being exposed to models without appropriate isolation and access controls.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
A Reddit post described telling ChatGPT it had consumed whisky to lower its inhibitions, an approach similar to the later simulated-drunkenness technique.
UNSW Sydney and other Australian machine-learning researchers reported that prompting, fine-tuning, or reinforcement-learning models to emulate drunken writing weakened privacy and safety guardrails in GPT-3.5, GPT-4, Llama, and Mistral models. Their evaluations found increased confidential-data disclosure and harmful-request compliance, with some drunk-conditioned models continuing harmful responses despite tested jailbreak defenses.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.