Recent research and expert commentary highlight the growing sophistication of attacks targeting artificial intelligence systems, particularly large language models (LLMs) and multimodal AI. Studies reveal that attackers can exploit weaknesses in open weight LLMs through multi-turn prompt injection, significantly increasing the success rate of bypassing safety guardrails compared to single-turn attacks. These findings underscore the importance of robust alignment strategies and continuous adversarial testing, as attackers adapt their tactics over extended interactions to elicit harmful or unauthorized outputs from AI models.
In addition to text-based prompt injection, multimodal AI systems face unique threats from hidden or invisible instructions embedded in both text and visual data. Techniques such as white-on-white text in resumes or hidden prompts in online profiles can manipulate AI-driven screening tools, allowing malicious actors to influence outcomes undetected by human reviewers. These developments demonstrate that as AI systems become more integrated into critical decision-making processes, their attack surface expands, necessitating new approaches to securing both the models and their input channels against increasingly subtle and complex adversarial techniques.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
A Help Net Security report examined how attackers can use gradual, patient interaction strategies to push past security guardrails in open-weight AI models, underscoring an emerging attack technique against AI systems.
A SecuritySenses report described how multimodal AI systems can already be manipulated with invisible or hidden instructions embedded in content, highlighting a real-world security weakness in these models.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.