# AI-Powered Deanonymization: How Large Language Models Are Breaking Online Pseudonymity
## What the Research Reveals
A team of researchers from ETH Zurich, the Machine Alignment Training and Safety (MATS) program, and Anthropic has demonstrated that large language models can identify the real-world identities behind anonymous online accounts with striking accuracy — using nothing more than publicly available information, web search capabilities, and reasoning AI models.
The study, conducted through a formal ethics review process at ETH Zurich, found that an autonomous AI pipeline could match pseudonymous accounts to real professional profiles at approximately 90% precision, costing between one and four dollars per target. The researchers have chosen not to release their code, prompts, or any of the identities they uncovered, framing their work as a documentation of an emerging capability rather than a tool for harm.
“This represents a fundamental shift in what it means to be anonymous online,” the researchers wrote in their paper. “The practical obscurity protecting pseudonymous users no longer holds, and threat models for online privacy need to be reconsidered.”
## The Four-Step Method Behind the Findings
The researchers structured their attack methodology into four distinct phases they call **Extract, Search, Reason, and Calibrate**:
– **Extract:** The AI ingests a target’s written posts and produces a summary profile — noting geographic hints, occupational details, personal interests, and linguistic quirks like regional spelling patterns or references to pets and hobbies.
– **Search:** That summary is converted into a numerical representation called an embedding, which allows the system to scan thousands of candidate profiles and find close matches based on similarity.
– **Reason:** A more advanced reasoning model evaluates the top candidates, cross-referencing details the way a human investigator would, but at dramatically greater speed and scale.
– **Calibrate:** The AI assesses its own confidence level, only making a guess when it is sufficiently certain and remaining silent otherwise to avoid false accusations.
## Testing the Approach
In one major experiment, the researchers selected 338 users of a pseudonymous tech forum who had publicly linked their LinkedIn profiles in their bios. After stripping out all directly identifying information, they tasked an AI agent with discovering each person’s real identity using only web search capabilities. The system correctly identified 226 of the 338 individuals — roughly 67% — and among its guesses, approximately one in ten was incorrect.
In a separate test using interview transcripts from 125 scientists, the AI correctly identified at least nine individuals based solely on how they described their research and professional backgrounds.
## Why This Feels Different From Past Deanonymization
The concept of linking anonymous data to real identities is not entirely new. In 2008, researchers famously de-anonymized Netflix’s movie-rating dataset by matching it against public reviews on IMDb. The critical difference today is the nature of the data and the method.
Earlier deanonymization efforts relied on structured, clean datasets — spreadsheets of ratings or ratings paired with timestamps. Modern AI models can work with unstructured, messy human language: jokes, slang, offhand remarks, fragmented thoughts, and stylistic quirks that were previously too noisy to exploit. Moreover, the entire process is now automated and costs almost nothing to execute.
## The Cost and Accessibility Factor
Perhaps the most unsettling aspect of the research is the cost. Running a complete deanonymization query against a target costs between $1 and $4 in AI subscription fees. No hacking, no data breaches, no insider access — just the same web search and text analysis abilities that millions of people use daily for mundane tasks.
This low barrier means the technique does not require specialized technical skills or criminal infrastructure. The researchers argue this is precisely why traditional security approaches fall short: there is no single “deanonymize” switch to disable, only a chain of individually benign-looking operations that, when combined, become powerful.
## Important Caveats and Limitations
Several factors temper the most alarming interpretations of these findings:
**Controlled test conditions:** The researchers selected subjects whose real identities they already knew, using accounts that had voluntarily linked to LinkedIn or splitting a single person’s posting history between two profiles to test whether the AI could connect them. This represents a best-case scenario, not evidence that every anonymous account is immediately vulnerable.
**The haystack problem:** As the pool of potential candidates grows, finding the correct match becomes significantly harder. Against a candidate pool of 89,000 profiles, the most effective method still captured only about half of all correct matches, even at 90% precision. Precision — how often a guess is right — remains high, but recall — how many true matches are actually found — drops as the search space expands.
**No release of tools:** The researchers have not published their code, prompts, or any real identities they uncovered. They are not distributing a ready-made deanonymization toolkit.
## Who Is Most at Risk?
While anyone with an online pseudonym faces some degree of exposure, certain communities have stronger reasons for concern:
– **Activists and whistleblowers** who rely on anonymity to speak freely may find their real identities increasingly traceable through AI analysis of their posting history.
– **Victims of abuse** who use pseudonyms for safety could be targeted by adversaries with minimal resources.
– **Members of marginalized communities** — including those discussing sexuality, immigration, or political dissent — face heightened risks if their posting history contains enough contextual clues.
– **Professionals in sensitive roles** whose employers discourage public opinion-sharing may be exposed through hobbies, hometown references, or employer-specific jargon.
The cryptocurrency community has witnessed this threat firsthand. A wave of doxxings and coordinated harassment campaigns following data breaches in recent years demonstrated how quickly a leaked identity can translate into physical threats at someone’s home address. AI-driven deanonymization removes the need for a data breach entirely — it works from what people have already posted publicly.
## What Can Be Done?
The researchers and privacy advocates suggest several practical approaches:
– **Reduce identifying details:** Avoid linking a single pseudonym to too many specific, searchable facts — hometowns, employers, pet names, and niche interests accumulate into a recognizable fingerprint over time.
– **Separate identities:** Use different pseudonyms for different contexts so that no single account contains a comprehensive picture of your life.
– **Choose platforms carefully:** Some services are designed with data retention policies that minimize the information available for analysis.
– **Support AI safety research:** The gap between AI capabilities and the guardrails meant to contain them continues to widen, as highlighted by recent disclosures that state-backed actors have used AI tools for cyberespionage campaigns.
—
## Frequently Asked Questions
**Q: Is this just a theoretical paper, or has anyone actually used this technique in the real world?**
A: The paper documents a capability that exists in current AI models. The researchers have not released their tools, but the methods described use commercially available AI services and web search — meaning anyone with a credit card and internet access could potentially replicate the approach.
**Q: How accurate is the AI at identifying people?**
A: In controlled tests, the system achieved 90% precision — meaning 9 out of 10 guesses were correct. However, it only found about 67% of the true identities in the Hacker News test, and the success rate dropped significantly against larger candidate pools of tens of thousands of profiles.
**Q: Can this work on any anonymous account?**
A: Not necessarily. Accounts with very little posting history, no personal details, and minimal external web presence are much harder to identify. The attack works best when a person has accumulated years of posts containing contextual clues — hobbies, locations, workplace details, and writing style.
**Q: Why didn’t the researchers release their code and tools?**
A: The researchers withheld their code, prompts, and all real identities uncovered as an ethical precaution. The study was reviewed and approved by ETH Zurich’s ethics board before publication. Their goal was to document the risk, not to provide a weapon.
**Q: How much does it cost to deanonymize someone?**
A: According to the paper, running one complete search costs between $1 and $4 in AI service fees, making it accessible to virtually anyone.
**Q: Is this the first time AI has been used for deanonymization?**
A: No, but this research represents one of the most comprehensive demonstrations at scale. Earlier efforts like the Netflix-IMDb matching used structured data and manual analysis, whereas this approach uses AI to handle messy, unstructured human language automatically.
**Q: What should I do if I’m concerned about my own anonymity online?**
A: Privacy experts recommend minimizing the number of specific, identifying details tied to any single pseudonym, using separate identities for different online activities, and avoiding platforms with poor data retention policies for sensitive discussions.
—
## Conclusion
The research from ETH Zurich, MATS, and Anthropic serves as both a warning and a call to action. It demonstrates that the combination of current AI capabilities — web search, text summarization, embedding-based matching, and logical reasoning — can erode the practical anonymity that millions of internet users rely on for safety, freedom of expression, and privacy.
The findings do not suggest that anonymity is impossible, but they do indicate that the cost and effort required to break it have dropped to near-zero levels. For platforms, policymakers, and individual users alike, the message is clear: the assumptions underlying pseudonymous communication online need to be revisited and strengthened before these capabilities become widely weaponized.
The researchers chose to publish responsibly — documenting the problem without handing out the tools. The next step belongs to the broader community to develop better protections, smarter defaults, and a deeper understanding of what it truly means to be anonymous in an age of intelligent machines.
Thank you for reading



