OpenAI announced it is acquiring Promptfoo, an AI security startup focused on testing and hardening LLMs against adversarial abuse. OpenAI said Promptfoo’s technology will be integrated into OpenAI Frontier (its enterprise AI-agent platform) to support automated red-teaming, security evaluation of agentic workflows, and monitoring for risk and compliance, while continuing to develop Promptfoo’s open-source tooling.
Separately, red-team startup CodeWall reported that its autonomous AI agent compromised McKinsey’s internal generative AI platform (Lilli) and obtained full read/write access to the chatbot environment within roughly two hours, without using credentials. CodeWall claimed access to large volumes of sensitive data in plaintext (including chat logs, files, user accounts, and system prompts), and emphasized that writable system prompts could enable response poisoning and manipulation of outputs at scale—highlighting how agentic AI can be used to accelerate intrusions and why enterprises are prioritizing agent security controls and continuous adversarial testing.

Track how attackers are adapting to this technology.
8 events from the most recent confirmed update back to the earliest known activity.
On March 9, 2026, OpenAI announced it had acquired Promptfoo, an AI security startup focused on protecting LLMs from adversarial abuse. OpenAI said Promptfoo’s technology would be integrated into OpenAI Frontier after the deal closes.
McKinsey stated that, with support from a third-party forensics firm, it found no evidence that client confidential data was accessed by the researchers or other unauthorized parties.
After receiving the disclosure, McKinsey patched unauthenticated endpoints, took a development environment offline, and blocked public API documentation. McKinsey said it remediated the issues within hours.
On March 1, 2026, CodeWall disclosed the full attack chain to McKinsey after identifying the vulnerabilities in Lilli.
CodeWall reported that an autonomous AI red-team agent chained the discovered flaws and, in about two hours, obtained read/write access to Lilli’s production database, potentially exposing tens of millions of plaintext chat messages and hundreds of thousands of files.
In late February 2026, CodeWall researchers found multiple weaknesses in McKinsey’s internal generative-AI platform Lilli, including exposed API documentation, unauthenticated endpoints, and a SQL injection path.
In July 2025, Promptfoo was valued at $86 million after a funding round, having raised $23 million in total according to the report.
Promptfoo was founded in 2024 by Ian Webster and Michael D’Angelo to build tools for testing large language model security vulnerabilities, including an open-source interface and library.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.