Researchers at NeuralTrust disclosed a new multi-turn jailbreak technique dubbed semantic chaining that can bypass safety controls in multimodal/image-capable AI models including xAI’s Grok 4 and Google’s Gemini Nano Banana. The method splits a malicious objective into several benign-looking steps so the model’s intent tracking and safety checks evaluate each prompt in isolation and miss the cumulative, policy-violating end state; reporting characterized it as simple enough for non-technical users to execute.
The demonstrated attack flow uses staged image-editing prompts (e.g., generate a harmless base image, make minor substitutions, then pivot to sensitive content) to exploit “editing mode” blind spots and ultimately produce prohibited visuals. A key limitation noted is that the harmful payload is primarily realized via image rendering, including embedding disallowed text as pixel-level content (e.g., “educational posters”/diagrams) that may evade text-focused filters even when the model would refuse to output the same content directly as text, underscoring gaps in current guardrails for agentic, multi-step interactions and image-generation pipelines.

Track how attackers are adapting to this technology.
3 events from the most recent confirmed update back to the earliest known activity.
The disclosure highlighted that some models that block banned text directly may still render prohibited text inside images, exposing a text-safety loophole. Researchers argued that current reactive guardrails are inadequate for multi-turn or agentic use and called for intent-aware defenses embedded in the model's reasoning or editing process.
In testing, NeuralTrust reported that Semantic Chaining could bypass safety mechanisms in models including Grok 4 and Gemini Nano Banana Pro, as well as other prominent image models, to generate policy-violating visual outputs.
NeuralTrust researchers developed and documented a multimodal jailbreak method called 'Semantic Chaining' that distributes malicious intent across multiple benign-looking prompt or image-edit steps to evade model safety checks.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.