New research reports that LLM agents can deanonymize pseudonymous users by extracting biographical and behavioral clues from a small number of “anonymous” posts and then using web search and reasoning to identify real-world identities. The technique was evaluated across multiple sources of unstructured text (including Hacker News, Reddit, LinkedIn, and anonymized interview transcripts) and is described as achieving high precision while scaling to tens of thousands of candidates, making a task that previously required significant human effort increasingly practical.
Experimental results described in coverage indicate LLM-based approaches can outperform classical deanonymization baselines (including methods similar to the Netflix Prize-style linkage attack), with precision degrading more slowly as the attacker makes more guesses and with agentic steps (e.g., “Reason” and “Calibrate”) improving recall at high-precision thresholds. Proposed mitigations include rate limiting and anti-scraping controls on user-data APIs, restricting bulk exports, and LLM-provider monitoring/guardrails to detect and refuse deanonymization misuse; the research warns the capability could be abused by governments to unmask critics, by advertisers for hyper-targeted profiling, or by criminals to build target dossiers at scale.

Track how attackers are adapting to this technology.
3 events from the most recent confirmed update back to the earliest known activity.
Alongside the findings, the authors proposed defenses including rate limiting, anti-scraping detection, restricting bulk data exports, and LLM-provider guardrails to refuse deanonymization requests. They warned that improving LLM-based re-identification could aid government surveillance, hyper-targeted advertising, and personalized social engineering.
The research reported successful deanonymization across platforms including Hacker News, Reddit, LinkedIn, anonymized interview transcripts, and a Netflix-derived evaluation setup. In one reported test, the system achieved up to 67% accuracy at 90% precision on nearly 1,000 LinkedIn-to-Hacker News matches, and another dataset showed correct re-identification of 9 out of 125 candidates.
Researchers from ETH Zurich created an automated framework that uses large language models to infer identity-relevant details from pseudonymous online posts, search the web for matching profiles, and rank likely identities with confidence scoring. The work was designed to show that LLMs can automate deanonymization that previously required substantial manual effort.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
4 references tracked. Mallory keeps watching after this page renders.
cyberscoop.com
Open sourcetechxplore.com
Open sourcearstechnica.com
Open sourceschneier.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.