xAI released Grok 4.7, a larger coding and professional-task model trained through extended reinforcement-learning runs on difficult, multi-hour workflows. The company says the approach improves self-verification, long-context management, and operation within its Grok Bot agent harness, aiming to reduce error accumulation during unattended tasks. Grok 4.7 scored 38.0% on Terminal-Bench 4.0, 46.3% on CursorBench 4.0, and 1,657 on AA Briefcase v1.1—improvements over Grok 4.6, though below Claude Fable 5.1's 57.9% Terminal-Bench result.
The release also adds stronger safety filtering for potentially dangerous cybersecurity requests. xAI reported a 3.3% acceptance rate for risky dual-use prompts on HackerBench v0.3 and is offering defensive red-team capabilities to selected invited partners. xAI has not disclosed technical evidence explaining the source or scope of the reported performance improvements, and long-running coding agents remain unreliable despite the gains.

Track how attackers are adapting to this technology.
1 event from the most recent confirmed update back to the earliest known activity.
xAI released Grok 4.7, a model trained with reinforcement learning on difficult multi-hour tasks and intended to improve self-verification and long-context handling. The release included availability through Cursor, Grok Build, the Grok API and other platforms, as well as new safety measures including stronger refusal behavior for dangerous cybersecurity requests and limited defensive red-team access for invited partners.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.