Microsoft released Agent Lightning v1.0, an MIT-licensed framework for training agentic systems with reinforcement learning without rebuilding the production agent loop inside the training stack. The framework leaves context construction, tool execution, and agent-environment interactions in the serving harness while optimizing observed LLM request-response calls across a service boundary, aiming to reduce train-serve mismatch.
Using 6,000 training examples and modest compute, Microsoft reported improving Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified, OpenAI's human-validated software-engineering benchmark subset. Agent Lightning includes reproducible scripts, open data and models, a data-cleaning pipeline, and reward-hacking safeguards; however, organizations deploying it still need to validate reward design, evaluation environments, safety and observability integration, and version control for the production harness.

Track how attackers are adapting to this technology.
4 events from the most recent confirmed update back to the earliest known activity.
Microsoft Research introduced Agent Lightning in August 2025 as an infrastructure concept for optimizing post-training LLM-based agents through reinforcement learning.
Microsoft reported that Agent Lightning reinforcement learning using 6,000 training examples raised Qwen3.5-9B's SWE-bench Verified score from 41.8% to 56.4%, a 14.6 percentage-point gain.
Microsoft released Agent Lightning v1.0 under the MIT license, including its complete workflow, data-cleaning pipeline, reward-hacking prevention, and reproducible GitHub training scripts. The release uses a production harness for context construction, tool execution, and agent-environment interaction while training optimizes LLM request-response pairs across a service boundary.
OpenAI introduced SWE-bench Verified, as indicated by the referenced OpenAI announcement.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.