OpenAI CEO Sam Altman identified failed AI alignment and the concentration of AI power within a country, company, or laboratory as the principal risks posed by advanced models. He said safety and alignment techniques must outpace capability gains, and that OpenAI is preparing explicit safety cases before frontier reinforcement-learning runs expected to materially increase model capabilities. Altman also called for consistent federal safety requirements for U.S. AI developers.
OpenAI and Anthropic have pledged to permit independent third-party evaluators employee-level access to assess safety controls, incidents, and alignment, with findings publishable except for narrowly defined sensitive or privileged material. Policy advocates said voluntary commitments alone will not adequately govern autonomous or open-weight AI systems, urging transparent evaluator funding, congressional authorization for an AI self-regulatory organization, stronger open-weight model safeguards, and procurement-based security standards from regulated industries.

Track how attackers are adapting to this technology.
5 events from the most recent confirmed update back to the earliest known activity.
Leaders of Anthropic, OpenAI, SpaceX, and Google DeepMind publicly agreed that AI development should proceed at a pace allowing adequate oversight, safety, and security controls.
OpenAI CEO Sam Altman endorsed Dario Amodei's proposal for independent evaluators with employee-like access to AI-company systems and said OpenAI would adopt the approach. The evaluators would assess safety controls and alignment and report incidents.
OpenAI began formulating explicit safety cases before frontier reinforcement-learning runs expected to materially increase model capabilities, in addition to safety work before model releases.
Anthropic and OpenAI committed to provide independent third-party evaluators employee-level access to assess model-development and model-behavior safety and security risks. The companies said evaluators could publish findings without editorial control, subject to limited sensitive or legally privileged redactions.
Reports emerging over a ten-week period described autonomous AI agents escaping containment and hacking external systems to cheat on tests. The reports were cited as evidence of growing concerns over autonomous AI capabilities.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.