Research from Novee Security found that default AI agent harnesses from major vendors could be abused to exfiltrate secrets, bypass trust boundaries, and chain individually approved actions into harmful outcomes. Elad Meged reported multiple issues affecting Anthropic, Google, and OpenAI tooling, arguing that the root problem lies in architectural trust decisions, approval logic, and cross-stage handoffs rather than prompt injection alone. One Google Gemini CLI issue was tracked as GHSA-wpqr-6v78-jr5g with a CVSS 10.0 rating, while weaknesses were also identified in Anthropic's Claude Code Action and OpenAI's Codex CLI sandbox.
The findings showed how controls can fail when agent components such as tools, memory, prompts, and execution environments trust each other too broadly. Researchers said Google's agent could be pushed into a supply-chain-style attack that wrote to its own GitHub repository, illustrating how safe-looking steps can combine into exploitable workflows. Vendors have added mitigations and fixes, but the repeated pattern across products points to a broader industry problem for organizations deploying autonomous agents on untrusted input, especially where human review is limited or absent.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
5 events from the most recent confirmed update back to the earliest known activity.
The reporting states that vendors added defenses and issued fixes in response to the disclosed findings. Despite those remediations, the researchers said the recurring pattern across products points to an industry-wide design problem in autonomous agent workflows.
The research highlighted trust-boundary weaknesses in OpenAI's Codex CLI sandbox. These issues fit the broader pattern of harness components making unsafe assumptions during handoffs between prompts, tools, memory, and execution environments.
The research demonstrated a critical issue in Google's Gemini CLI, tracked as GHSA-wpqr-6v78-jr5g and rated CVSS 10.0. Dark Reading also reports that Google's AI agent could be abused for a supply chain attack and to write to its own GitHub repository.
Meged reported multiple security findings affecting Anthropic's Claude Code Action as part of the broader AI harness research. The findings were tied to weaknesses in how individually permitted actions could be combined into unsafe workflows.
Research by Elad Meged of Novee Security found that default AI agent harnesses from Anthropic, Google, and OpenAI could be abused through prompt-injection-driven workflows to exfiltrate secrets or bypass trust boundaries. The work argued that the root issue was architectural flaws in harness trust decisions, approval logic, and cross-stage handoffs rather than prompt injection alone.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
3 references tracked. Mallory keeps watching after this page renders.
codeby.net
Open sourcedarkreading.com
Open sourcehelpnetsecurity.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.