Multiple academic research reports describe new ways attackers can subvert or extract sensitive information from AI systems beyond traditional prompt-based abuse. University of Florida researchers presented Nullspace Steering (“Jailbreaking the Matrix”) as an internal, model-level technique to bypass AI guardrails by manipulating a model’s internal decision pathways, arguing that evaluating and hardening safety controls requires “popping the hood” rather than only testing external prompts. Separately, researchers from Shanghai Qi Zhi Institute, East China Normal University, Tsinghua University, and the Chinese Academy of Sciences reported that LLM model-editing workflows (e.g., “locate-then-edit” approaches intended to remove or correct memorized sensitive data) can leak the very data they aim to eliminate because parameter updates act as a side channel; they describe a two-stage reverse-engineering attack framework (KSTER) and a mitigation approach (subspace camouflage).
Georgia Tech researchers disclosed a backdoor-style vulnerability dubbed VillainNet affecting AI “super networks” used in autonomous driving, where a malicious modification can remain dormant across many configurations and then trigger under specific real-world conditions (e.g., a particular subnetwork selected during rain/road changes), enabling targeted control of vehicle behavior with high likelihood of success. In contrast, a Gopher Security blog post argues for lattice-based, zero-trust identity verification for agentic AI integrations (including Model Context Protocol tooling) and warns about quantum risks to RSA/ECC and token signatures; it is largely forward-looking guidance/advocacy rather than reporting a specific newly disclosed vulnerability or incident tied to the other research items.

Track how attackers are adapting to this technology.
5 events from the most recent confirmed update back to the earliest known activity.
The University of Florida team said its HMNS research was accepted to ICLR 2026 in Rio de Janeiro. The researchers also indicated plans to use UF’s HiPerGator supercomputer for larger-scale follow-on computations.
University of Florida researchers disclosed Head-Masked Nullspace Steering, an internal jailbreaking technique that suppresses and steers attention components to bypass LLM guardrails. They said the method outperformed prior approaches on four benchmarks and framed it as a defensive stress-testing tool.
Researchers from several Chinese institutions published an arXiv preprint describing KSTER, a reverse-engineering attack that can recover fingerprints of LLM edits and reconstruct sensitive edited content. They also proposed a mitigation called subspace camouflage and released code for both the attack and defense on GitHub.
The VillainNet paper, titled “VillainNet: Targeted Poisoning Attacks Against SuperNets Along the Accuracy-Latency Pareto Frontier,” was published in the 2025 ACM SIGSAC CCS proceedings. The research reported high attack reliability and argued that exhaustive detection would require roughly 66 times more compute and time.
Georgia Tech researchers presented VillainNet, a targeted poisoning/backdoor technique against AI SuperNet architectures used in autonomous driving, at ACM CCS in October 2025. The work showed a malicious behavior could be embedded in a single subnetwork and triggered under specific real-world conditions.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
3 references tracked. Mallory keeps watching after this page renders.
techxplore.com
Open sourcetechxplore.com
Open sourcetechxplore.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.