Researchers from Ben-Gurion University of the Negev unveiled GAVEL (Governance via Activation-based Verification and Extensible Logic), a research-stage AI safety framework that inspects the internal activation patterns of large language models instead of relying only on prompt and output moderation. The model-agnostic approach breaks unsafe behavior into granular cognitive elements and combines them into rules, aiming to identify policy-violating or dangerous activity inside the model's reasoning process.
The team said GAVEL could improve detection of jailbreaks, phishing-related behavior, and other malicious intent that can evade traditional filters through alternate languages or representation tricks. The work, presented at Black Hat USA 2026, is positioned as an additional defensive layer rather than a replacement for existing token-level controls, and researchers said broader adoption of shared tools, code, and rules will be needed before the approach becomes practical at scale.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
The researchers' GAVEL framework is being presented at Black Hat USA 2026 as a new AI security approach focused on inspecting model internals to detect malicious or policy-violating behavior. Coverage describes the project as still in the research stage and dependent on broader adoption of tools, code, and rules to become practical.
Researchers from Ben-Gurion University of the Negev developed GAVEL, a model-agnostic approach that analyzes internal neural activations of large language models and maps them to granular cognitive elements for detecting unsafe behavior such as jailbreaks and phishing-related activity. The work is described as an experimental defensive layer intended to complement token-based moderation rather than replace it.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.