The UK Government Cyber Coordination Centre, working with the UK AI Security Institute and the National Cyber Security Centre, used frontier AI models in a month-long series of in-person hackathons to test public government code repositories across nine organisations. The exercise uncovered 407 findings, including authentication bypass, data exposure, and remote code execution flaws; some were previously unknown zero-days, while others had already been identified and were protected by compensating controls. Officials said all critical and high-risk exploitable weaknesses were remediated and found no evidence that the vulnerabilities had been abused.
The pilot evaluated several workflows, including multi-stage agentic pipelines, deterministic scanners enhanced with AI triage, and reusable domain-specific Claude Skills, with humans kept in the loop for validation and escalation. One notable critical issue involved legacy GitHub Actions workflows that allowed an external user to trigger arbitrary code execution on a runner through a crafted comment on an open pull request. The government said the project used about £13,000 in model tokens and concluded that structured architecture, triage, and human oversight mattered more than the specific model, with a second phase planned to expand testing to more departments, more models, and closed-source environments.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
6 events from the most recent confirmed update back to the earliest known activity.
The UK government published a case study describing the pilot's methods, findings, and lessons learned, including that tightly scoped AI components in structured pipelines produced the strongest results. The publication also disclosed that the month-long effort cost £13,000 in model tokens.
Following the pilot, the Government Cyber Coordination Centre said it plans a second phase that will expand to more departments, include additional models, and extend testing from public repositories into closed-source environments. It also identified integrating prioritization, review, and patch generation as the next major task.
The UK government said all critical weaknesses identified in the pilot, including all critical and high-risk exploitable issues, were remediated. It also reported no evidence that any of the identified findings had been exploited.
One notable critical vulnerability affected legacy GitHub Actions workflows in a repository supporting a key government digital service. An external user could trigger arbitrary code execution on a runner by posting a specially crafted comment on an open pull request, potentially exposing secrets and enabling wider repository compromise.
During the exercise, participants identified 407 findings, including critical weaknesses that could lead to authentication bypass, data exposure, and remote code execution. Some findings were previously unknown, while others were already understood and mitigated by compensating controls.
The Government Cyber Coordination Centre, working with the UK AI Security Institute and NCSC, ran a month-long series of in-person hackathons using frontier AI models to test vulnerability discovery across public government code repositories in nine organisations. Teams used structured pipelines, scanner-plus-model workflows, and domain-specific Claude Skills with human validation throughout.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
2 references tracked. Mallory keeps watching after this page renders.
infosecurity-magazine.com
Open sourcegov.uk
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.