OpenAI said internal evaluations of its upcoming Astra model showed enough progress in agentic coding and cybersecurity that it cannot rule out the system meeting the Critical threshold in the company’s Preparedness Framework. OpenAI defines that threshold as the ability to autonomously identify and develop functional zero-day exploits against many hardened real-world critical systems, or to execute novel end-to-end cyberattacks against hardened targets from only a high-level goal. The company said Astra was not involved in the recent Hugging Face exploitation activity.
In response, OpenAI said it is slowing Astra’s path to release, pausing internal work that does not satisfy strengthened safeguards, and expanding robustness testing with government agencies and selected AI safety organizations. The tighter controls include isolated testing environments, restricted network and tool access, sandboxed execution, enhanced protection for model weights, and universal monitoring for risky actions; reports also said any eventual access to Astra’s strongest capabilities could be limited through vetting and oversight similar to OpenAI’s Trusted Access for Cyber program.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
14 events from the most recent confirmed update back to the earliest known activity.
On August 7, 2026, OpenAI published a security notice disclosing Astra's preliminary cyber-risk assessment and the company's response under its Preparedness Framework. OpenAI said it would also work with relevant government agencies and select AI safety organizations on further testing.
OpenAI explicitly stated that Astra was not involved in exploiting Hugging Face. The clarification accompanied its public discussion of Astra's cyber capabilities and new safeguards.
In response to Astra's evaluation results, OpenAI said it was increasing robustness testing, implementing stricter internal security controls, and pausing Astra-related internal activities that did not meet the strengthened requirements. The controls include isolated testing environments, restricted network and tool access, enhanced protections, sandboxed execution, and universal monitoring for risky actions.
Axios reported that Anthropic warned about models improving themselves in a company blog in June and called for a global pause in AI development. This was presented as part of the broader industry debate over frontier model risk.
Axios reported that Anthropic released a safer version of its most cyber-capable model, Mythos, in June. Anthropic said it was being deliberately more conservative with that release.
The New Stack reported that OpenAI operates a Trusted Access for Cyber program that gives approved security professionals access to tools not available to everyone. The reference to the program is supported by OpenAI's February 2026 blog post introducing it.
Axios reported that Anthropic rolled back its earlier commitment to pause training of powerful models if capabilities surpassed its ability to control them in an update to its Responsible Scaling Policy. The update occurred in February.
OpenAI said the Preparedness Framework previously guided its response to models approaching the high capability threshold for biology in June 2025. This was cited as an earlier use of the framework before Astra's cyber assessment.
OpenAI said it first published its Preparedness Framework in December 2023. The framework later governed how the company responds to models approaching higher-risk capability thresholds.
OpenAI announced it would expand and restructure its Daybreak program into Daybreak Blue and Daybreak Red to provide controlled access to security-focused AI models for trusted defenders. The company said Daybreak Red would include a new model called GPT-5.6-Cyber for flaw hunting and validation.
Axios reported that industry participants were briefed on an AI model evaluation framework during the week of the report. The framework is part of the Trump administration's effort to develop a pre-release review process for AI models.
Axios reported that a White House official said OpenAI voluntarily informed the administration of its plans to delay Astra's release. The report tied this to the company's decision to slow development while stronger safeguards are put in place.
OpenAI said recent internal evaluations of its upcoming Astra model showed enough advancement in agentic coding and cybersecurity that it could not rule out Astra reaching the Critical cybersecurity capability threshold. The company described the assessment as preliminary and said benchmarking and evaluation were continuing.
Anthropic said it is reducing how often its Fable model falls back or refuses prompts involving biology after earlier restrictions had made the model less useful for researchers. The change was described as a loosening of prior safeguards that were intended to prevent harmful chemical warfare guidance.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
15 references tracked. Mallory keeps watching after this page renders.
itpro.com
Open sourceinfosecurity-magazine.com
Open sourcesecurityaffairs.com
Open sourcezdnet.fr
Open sourceopenai.com
Open sourceopenai.com
Open sourceopenai.com
Open sourcecdn.openai.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.