Microsoft Releases Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Real-Harness Training
微軟開源 Agent Lightning v1.0:僅 3,500 行程式碼,直接用生產環境 Harness 訓練 Agent 的強化學習框架

Traditional agentic RL requires rebuilding complex agent loops inside training frameworks, causing high development costs and behavioral drift. Agent Lightning v1.0 solves this by introducing 'Harnessed Agentic RL,' which inserts an LLM proxy between the agent and the model, allowing the deployment harness to participate directly in training. Built on just 3,500 lines of code, the framework natively supports standard Kubernetes jobs to avoid costly commercial sandboxes. It features 'Collocated Async RL' for GPU sharing, doubling end-to-end training speed. Using only 6,000 samples, it boosted Qwen3.5-9B's SWE-bench Verified Pass@1 score from 41.8% to 56.4%.
Key points
Direct Training on Real Harnesses
Uses an LLM proxy to record prompts and actions, removing the need to rebuild complex agent codebases within training frameworks.
Extremely Lightweight Design
The entire framework consists of about 3,500 lines of code, making the agent RL control plane easy to audit, modify, and extend.
Collocated Async RL
Allows rollout and model updates to share the same GPUs asynchronously, achieving a 2x speedup and maximizing hardware utilization.
Native Kubernetes Support
Runs agent rollouts directly as standard Kubernetes jobs, bypassing the need for paid commercial sandbox services.
How it works
Why it matters
It bridges the gap between agent training and deployment. Instead of rewriting agent logic inside monolithic RL frameworks, developers can plug real-world agent harnesses (like OpenHands or Claude Code) directly into RL pipelines. Native Kubernetes support and Collocated Async RL democratize agent tuning by eliminating sandbox subscription fees and significantly lowering GPU overhead.
Who it affects
- AI Developer
- AI Researcher
- Enterprise Leader
How to use it
- 1Training software engineering agents to solve GitHub issues on benchmarks like SWE-bench.
- 2Creating a seamless pipeline that synchronizes production agent configurations directly with RL training loops.
Limitations & caveats
- Splitting rollouts into discrete training samples during retokenization can shift token boundaries and complicate sample merging.
- Scheduling highly variable workloads (sample count and length) onto fixed GPU and parallel configurations remains challenging.
Related
RECAST: Active Evidence Construction via Adaptive Routing and Computation
RECAST:透過自適應證據路由,為 LLM 計算出正確的上下文
RECAST is an agentic framework that treats evidence construction as a sequential decision process, utilizing a Router-Compiler loop to actively compute and derive the optimal context for RAG rather than relying on passive retrieval.
Agent in a Bottle: Can LLM Agents Package Their Capabilities into Cheap, Scalable Artifacts?
打造低成本 AI 工件:LLM Agent 是否具備「能力封裝」的本領?
The study introduces the BOTTLED benchmark to evaluate if LLM agents can autonomously package their capabilities into low-cost, task-specific artifacts, revealing that strong zero-shot performance does not guarantee successful bottling.
AdvSim2Real: Co-Evolving Web Agents and Adversaries inside a Web World Model
利用 Web 世界模型進行對抗訓練:AdvSim2Real 提升 Web Agent 抵禦適應性提示詞注入之能力
AdvSim2Real co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen world model, boosting the robustness of a 4B agent against adaptive prompt injections by 33.6%.