Large language models (LLMs) present a dual-use dilemma in cybersecurity, as their capabilities can be leveraged for both defensive and offensive purposes. Security researchers have identified purpose-built malicious LLMs, such as WormGPT and KawaiiGPT, which are designed to facilitate cybercrime by generating convincing phishing content and rapidly producing or modifying malicious code. The thin line between beneficial and harmful use of LLMs is defined largely by developer intent and the presence or absence of ethical safeguards, raising concerns about the proliferation of offensive AI tools in the threat landscape.
In addition to malicious use, LLMs face significant challenges in maintaining privacy and security due to contextual integrity failures and regulatory-driven censorship. Research from Microsoft highlights the need for AI agents to respect contextual privacy norms, as current models may inadvertently leak sensitive information. Meanwhile, the DeepSeek-R1 model demonstrates how geopolitical censorship mechanisms can introduce security flaws, such as insecure code generation and broken authentication, especially when handling politically sensitive prompts. These issues underscore the urgent need for robust privacy controls and security-aware development practices in the deployment of LLM-powered systems.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
4 events from the most recent confirmed update back to the earliest known activity.
Palo Alto Networks Unit 42 published research detailing how malicious LLMs such as WormGPT 4 and KawaiiGPT can support phishing, malware scaffolding, reconnaissance, and ransomware workflows. The report argues that commercialization and democratization of these tools are making AI-enabled cybercrime more scalable.
After the original WormGPT's emergence, a successor branded 'WormGPT 4' was promoted as a subscription service through Telegram and underground forums. Unit 42 describes this as a sign of commercialization of malicious LLMs and broader access for less-skilled threat actors.
Unit 42 reports that KawaiiGPT version 2.5 was identified in July 2025 as a freely available tool on GitHub. The model was described as capable of generating spear-phishing lures, lateral-movement scripts, and data-exfiltration code.
Unit 42 says WormGPT first emerged in July 2023 as a purpose-built malicious large language model, reportedly based on GPT-J 6B and trained on malicious datasets. It was marketed for criminal use cases such as phishing, malware development, and other offensive tasks with safety guardrails removed.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
3 references tracked. Mallory keeps watching after this page renders.
unit42.paloaltonetworks.com
Open sourcemicrosoft.com
Open sourcescworld.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.