{"id":1187920,"date":"2026-10-07T09:00:00","date_gmt":"2026-10-07T16:00:00","guid":{"rendered":"https:\/\/www.microsoft.com\/en-us\/research\/?p=1187920"},"modified":"2026-10-07T06:20:47","modified_gmt":"2026-10-07T13:20:47","slug":"agent-lightning-v1-0-a-3500-line-lightweight-agentic-rl-framework-for-training-agents-with-real-harnesses","status":"publish","type":"post","link":"https:\/\/www.microsoft.com\/en-us\/research\/blog\/agent-lightning-v1-0-a-3500-line-lightweight-agentic-rl-framework-for-training-agents-with-real-harnesses\/","title":{"rendered":"Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses"},"content":{"rendered":"\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1400\" height=\"788\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1.jpg\" alt=\"System architecture diagram. On the left, agents with harnesses \u2014 mini-SWE-agent, OpenHands, and OpenClaw \u2014 run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.\" class=\"wp-image-1188789\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1.jpg 1400w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1065x599-1.jpg 1065w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1279x720.jpg 1279w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-653x368.jpg 653w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-675x380.jpg 675w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w\" sizes=\"auto, (max-width: 1400px) 100vw, 1400px\" \/><\/figure>\n\n\n\n<div style=\"padding-bottom:0;padding-top:0\" class=\"wp-block-msr-immersive-section alignfull row\">\n\t\n\t<div class=\"container\">\n\t\t<div class=\"wp-block-msr-immersive-section__inner wp-block-msr-immersive-section__inner--narrow\">\n\t\t\t<div class=\"wp-block-columns mb-10 pb-1 pr-1 is-layout-flex wp-container-core-columns-is-layout-8f761849 wp-block-columns-is-layout-flex\" style=\"box-shadow:var(--wp--preset--shadow--outlined)\">\n<div class=\"wp-block-column is-layout-flow wp-block-column-is-layout-flow\">\n<h2 id=\"at-a-glance\" class=\"wp-block-heading h3\">At a glance<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Harnessed Agentic RL: Microsoft Research Asia introduces a training paradigm in which the same agent harness used in deployment participates directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.<\/li>\n\n\n\n<li>Lightweight by design: Agent Lightning v1.0 delivers a complete agent RL control plane in roughly 3,500 lines of code.<\/li>\n\n\n\n<li>Native Kubernetes support: agents run as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes, or local infrastructure, with no dependency on paid commercial sandbox services.<\/li>\n\n\n\n<li>Data-efficient training recipe: an end-to-end coding agent pipeline raised Qwen3.5-9B from 41.8% to 56.4% Pass@1 on SWE-bench Verified, a 14.6 percentage point gain, using only about 6,000 training samples based on open sourced dataset.<\/li>\n<\/ul>\n<\/div>\n<\/div>\t\t<\/div>\n\t<\/div>\n\n\t<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">AI agents have evolved from single models to complex full-stack systems built from models, tools, and execution environments. Their capabilities increasingly depend on the agent harness that coordinates them from outside the model. Reinforcement learning (RL) is an approach where AI systems learn through trial and error, guided by rewards and penalties for their actions. RL can make those agents better, but most agent RL systems require developers to reimplement the agent inside the training framework. That is costly, and it means the agent being trained is not quite the agent that gets deployed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To address this, researchers at Microsoft Research Asia have introduced the Harnessed Agentic RL training paradigm and open-sourced a fully rebuilt <a class=\"msr-external-link glyph-append glyph-append-open-in-new-tab glyph-append-xsmall\" href=\"https:\/\/github.com\/microsoft\/agent-lightning\" target=\"_blank\" rel=\"noopener noreferrer\">Agent Lightning v1.0<span class=\"sr-only\"> (opens in new tab)<\/span><\/a>. Compared with the original, Agent Lightning, v1.0 puts more emphasis on staying lightweight, on integrating with real harnesses, and on a complete, reproducible agent RL training pipeline.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Agent Lightning v1.0 was rebuilt around Harnessed Agentic RL, with key improvements:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Lightweight<\/strong>: the entire framework is about 3,500 lines of code. Agent Lightning v1.0 implements a complete Harnessed Agentic RL system in a codebase that is small and clear enough to understand, modify, and extend.<\/li>\n\n\n\n<li><strong>Training on a real agent harness<\/strong>: agents reach the model through the large language model (LLM) proxy in Agent Lightning v1.0, leaving existing harness code unchanged.<\/li>\n\n\n\n<li><strong>Native Kubernetes support<\/strong>: agents run directly as Kubernetes jobs, without external commercial sandbox services. Self-managed clusters and local infrastructure alike can support rollouts at scale.<\/li>\n\n\n\n<li><strong>A complete coding agent training example<\/strong>: an end-to-end pipeline built on Qwen3.5-9B raised Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points, using only about 6,000 training samples.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"the-limits-of-traditional-agentic-rl\" class=\"wp-block-heading\">The limits of traditional agentic RL<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Traditional agentic RL assumes the training framework owns the interaction loop with the environment. In a ReAct-style loop, the model generates an action, the environment returns an observation, the observation is appended to the context, and the model generates the next action, so the whole rollout maps onto one continuous token trajectory. Early RL systems such as verl, AReaL, and slime were built this way, which meant training an agent required rebuilding its loop inside the RL framework.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Real harnesses have outgrown that assumption. Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code, and Codex each bring their own context management, tool protocols, execution logic, and dependencies, as do general-purpose agent systems. Rebuilding one for training is expensive, and the rebuilt agent may no longer behave in the same way as the deployed agent.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Agent Lightning takes a different route. It places an LLM proxy between the agent and the model. The agent continues to run as before: simply point the endpoint that previously called the model API at Agent Lightning, and the training framework can observe and record its model calls. In v1.0, the researchers go further and formally define this paradigm as Harnessed Agentic RL: whichever agent harness is used in deployment is the harness that takes part directly in reinforcement learning during training (Figure 1).<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1083\" height=\"472\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1.jpg\" alt=\"Figure 1: Side-by-side comparison of two training loops. In Agentic RL, the environment exchanges actions and observations with a tokenizer, which passes action and observation tokens to the policy model. In Harnessed Agentic RL, an agent harness handling context and orchestration sits between the environment and an OpenAI-like API, which exchanges input and output tokens with the policy model.\" class=\"wp-image-1187926\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1.jpg 1083w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1-300x131.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1-1024x446.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1-768x335.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.1-240x105.jpg 240w\" sizes=\"auto, (max-width: 1083px) 100vw, 1083px\" \/><figcaption class=\"wp-element-caption\">Figure 1. Traditional agentic RL compared with Harnessed Agentic RL. In traditional agentic RL, the training framework manages the environment and the agent loop. In Harnessed Agentic RL, the harness manages both.<\/figcaption><\/figure>\n\n\n\n<h2 id=\"four-challenges-in-training-with-real-harnesses\" class=\"wp-block-heading\">Four challenges in training with real harnesses<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A core difference between Harnessed Agentic RL and traditional agentic RL is that the environment interaction loop is handled by the agent harness rather than the training framework. The training system can only observe a series of LLM request and response pairs, so a single rollout may be split into a variable number of training samples. This brings four key challenges:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Retokenization and sample merging<\/strong>: Harnesses keep context as text, but RL training needs the token IDs sampled during the rollout. Passing text through the chat template and tokenizer again can shift token boundaries, so adjacent calls cannot always be merged into one sample.<\/li>\n\n\n\n<li><strong>Advantage calculation<\/strong>: Retokenization, subagents and context summarization can split one rollout into several samples. Computing baselines and advantages directly at the sample level causes rollouts that produce more samples to be counted repeatedly, which alters the original statistical relationships at the rollout level.<\/li>\n\n\n\n<li><strong>Loss normalization<\/strong>: Averaging loss by sample count gives more weight to rollouts that produce more samples. Since sample count is often just a product of harness behavior, loss normalization must also avoid being distorted by it.<\/li>\n\n\n\n<li><strong>Training backend scheduling<\/strong>: Sample count and length are known only after the harness finishes, while GPU counts and data\/tensor parallel configurations are usually fixed. The backend has to map a variable workload into fixed resources.<\/li>\n<\/ul>\n\n\n\n\t<div class=\"border-bottom border-top border-gray-300 mt-5 mb-5 msr-promo text-center text-md-left alignwide\" data-bi-aN=\"promo\" data-bi-id=\"670821\">\n\t\t\n\n\t\t<p class=\"msr-promo__label text-gray-800 text-center text-uppercase\">\n\t\t<span class=\"px-4 bg-white display-inline-block font-weight-semibold small\">Spotlight: Microsoft research newsletter<\/span>\n\t<\/p>\n\t\n\t<div class=\"row pt-3 pb-4 align-items-center\">\n\t\t\t\t\t\t<div class=\"msr-promo__media col-12 col-md-5\">\n\t\t\t\t<a class=\"bg-gray-300 display-block\" href=\"https:\/\/info.microsoft.com\/ww-landing-microsoft-research-newsletter.html\" aria-label=\"Microsoft Research Newsletter\" data-bi-cn=\"Microsoft Research Newsletter\" target=\"_blank\">\n\t\t\t\t\t<img decoding=\"async\" class=\"w-100 display-block\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2019\/09\/Newsletter_Banner_08_2019_v1_1920x1080.png\" alt=\"\" \/>\n\t\t\t\t<\/a>\n\t\t\t<\/div>\n\t\t\t\n\t\t\t<div class=\"msr-promo__content p-3 px-5 col-12 col-md\">\n\n\t\t\t\t\t\t\t\t\t<h2 class=\"h4\">Microsoft Research Newsletter<\/h2>\n\t\t\t\t\n\t\t\t\t\t\t\t\t<p id=\"microsoft-research-newsletter\" class=\"large\">Stay connected to the research community at Microsoft.<\/p>\n\t\t\t\t\n\t\t\t\t\t\t\t\t<div class=\"wp-block-buttons justify-content-center justify-content-md-start\">\n\t\t\t\t\t<div class=\"wp-block-button is-style-fill-chevron\">\n\t\t\t\t\t\t<a href=\"https:\/\/info.microsoft.com\/ww-landing-microsoft-research-newsletter.html\" aria-describedby=\"microsoft-research-newsletter\" class=\"btn btn-brand glyph-append glyph-append-chevron-right\" data-bi-cn=\"Microsoft Research Newsletter\" target=\"_blank\">\n\t\t\t\t\t\t\tSubscribe today\t\t\t\t\t\t<\/a>\n\t\t\t\t\t<\/div>\n\t\t\t\t<\/div>\n\t\t\t\t\t\t\t<\/div><!--\/.msr-promo__content-->\n\t<\/div><!--\/.msr-promo__inner-wrap-->\n\t<\/div><!--\/.msr-promo-->\n\t\n\n\n<h2 id=\"building-a-complete-agent-rl-control-plane-in-3-500-lines-of-code\" class=\"wp-block-heading\">Building a complete agent RL control plane in 3,500 lines of code<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In system design, Agent Lightning v1.0 treats simplicity as its first principle. The entire framework is about 3,500 lines of code, with three core components: the API Gateway, the Rollout Controller, and the Customized Trainer (Figure 2).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The API gateway stores rollouts, models, and events, and serves as an OpenAI-compatible LLM proxy. It links every model call from the harness to its rollout and records the prompts, responses, and log probabilities that training needs. The rollout controller starts and manages agent execution, either as local processes or as standard Kubernetes jobs, keeping agent execution separate from the trainer. The customized trainer, built on verl, creates rollouts, waits for them to finish, collects samples, and assembles the final training samples through a sample adapter. As a result, for an existing agent harness, simply pointing the model endpoint at the Agent Lightning proxy is usually enough to connect quickly to RL training.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"865\" height=\"293\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.2.png\" alt=\"Figure 2: System architecture diagram. On the left, agents with harnesses \u2014 mini-SWE-agent, OpenHands, and OpenClaw \u2014 run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.\" class=\"wp-image-1187930\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.2.png 865w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.2-300x102.png 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.2-768x260.png 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.2-240x81.png 240w\" sizes=\"auto, (max-width: 865px) 100vw, 865px\" \/><figcaption class=\"wp-element-caption\">Figure 2. The Agent Lightning v1.0 system architecture showing the API gateway, rollout controller, and customized trainer.<\/figcaption><\/figure>\n\n\n\n<h2 id=\"collocated-async-rl\" class=\"wp-block-heading\">Collocated async RL<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Rollout times vary widely across agents. Synchronous RL waits for the slowest agent in a batch and leaves GPUs idle, while fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training. In response, Agent Lightning v1.0 introduces Collocated Async RL, which lets rollout and model updates share the same set of GPUs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once the system has collected enough rollouts, the update begins: the API Gateway pauses accepting new requests and waits for requests already in progress to finish, and rollout resumes after the update completes. The entire state transition is transparent to the external agent harness. In experiments, this approach achieved about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL (Figure 3).<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1196\" height=\"380\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3.jpg\" alt=\"Figure 3: Three GPU scheduling timelines. Synchronous RL uses four GPUs at low efficiency, with long idle gaps before a single update block. Collocated Async RL uses the same four GPUs at high efficiency, interleaving full and partial rollouts with update blocks. Asynchronous RL reaches high efficiency but requires eight GPUs. Bars are colored for full rollout, partial rollout, and update.\" class=\"wp-image-1187927\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3.jpg 1196w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3-300x95.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3-1024x325.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3-768x244.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.3-240x76.jpg 240w\" sizes=\"auto, (max-width: 1196px) 100vw, 1196px\" \/><figcaption class=\"wp-element-caption\">Figure 3. Synchronous RL, asynchronous RL, and Collocated Async RL compared. Collocated Async RL raises utilization while occupying fewer GPUs.<\/figcaption><\/figure>\n\n\n\n<h2 id=\"running-agents-on-kubernetes\" class=\"wp-block-heading\">Running agents on Kubernetes<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Collecting enough rollouts means running many agents at once, which consumes substantial CPU, memory, and compute resources. Other Harnessed Agentic RL frameworks often host those agents on commercial sandbox services such as Modal Sandbox or E2B, where cost climbs quickly with scale. Instead, Agent Lightning v1.0 runs them as standard Kubernetes jobs, reusing existing self-managed clusters, cloud Kubernetes, or local infrastructure (Figure 4). Existing compute resources are used more efficiently, large rollouts cost less, and the whole pipeline stays open source and reproducible.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1217\" height=\"332\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4.jpg\" alt=\"Figure 4: Flow diagram. An API Gateway holds three rollouts, two queueing and one running. The Rollout Controller polls the gateway and uses a Kubernetes reconciler to create jobs on a Kubernetes cluster, and a local reconciler to watch and list local processes. Status updates flow back to the gateway.\" class=\"wp-image-1187928\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4.jpg 1217w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4-300x82.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4-1024x279.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4-768x210.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.4-240x65.jpg 240w\" sizes=\"auto, (max-width: 1217px) 100vw, 1217px\" \/><figcaption class=\"wp-element-caption\">Figure 4. The Rollout Controller in Agent Lightning v1.0 provides native Kubernetes support, running agents directly as standard Kubernetes jobs.<\/figcaption><\/figure>\n\n\n\n<h2 id=\"6-000-training-samples-a-14-6-point-performance-gain\" class=\"wp-block-heading\">6,000 training samples, a 14.6-point performance gain<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To test the approach, researchers built a full pipeline on SWE-smith, mini-SWE-agent, and Qwen3.5-9B, covering data cleaning, environment construction, reward-hacking safeguards, and RL training. The training set holds about 6,000 samples and needs no large-scale compute. RL training alone raised Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified, a gain of 14.6 percentage points.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The coding agent experiments further confirm the earlier analysis of two challenges: advantage calculation and loss normalization. Compared with sample-level handling, rollout-level advantage combined with rollout-level normalization achieves a higher validation reward and keeps policy entropy more stable during training (Figure 5).<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1268\" height=\"463\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5.jpg\" alt=\"Figure 5: Two line charts plotting 200 training steps. On the left, validation reward: rollout-level advantage combined with rollout-level normalization reaches the highest reward at about 0.37, above rollout-level advantage alone and sample-level advantage. On the right, policy entropy: rollout-level advantage alone climbs steeply to about 0.65, while the combined method stays lower and steadier.\" class=\"wp-image-1187929\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5.jpg 1268w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5-300x110.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5-1024x374.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5-768x280.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-blog_EXT_FA_figure.5-240x88.jpg 240w\" sizes=\"auto, (max-width: 1268px) 100vw, 1268px\" \/><figcaption class=\"wp-element-caption\">Figure 5. Pass rate and policy entropy for Qwen3.5-9B on the SWE-smith validation set.<\/figcaption><\/figure>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-fe48e5de wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button is-style-fill\"><a data-bi-type=\"button\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/www.microsoft.com\/en-us\/research\/publication\/agent-lightning-v1-0-towards-harnessed-agentic-rl\/\">Technical report<\/a><\/div>\n\n\n\n<div class=\"wp-block-button is-style-fill-github\"><a data-bi-type=\"button\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/github.com\/microsoft\/agent-lightning\" target=\"_blank\" rel=\"noopener\">GitHub project<\/a><\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.<\/p>\n","protected":false},"author":43868,"featured_media":1188789,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr-author-ordering":[{"type":"user_nicename","value":"Zhiyuan He","user_id":"41743"},{"type":"user_nicename","value":"Yuqing Yang","user_id":"40654"}],"msr_hide_image_in_river":0,"footnotes":""},"categories":[1],"tags":[],"research-area":[13556,13547],"msr-region":[],"msr-event-type":[],"msr-locale":[268875],"msr-post-option":[243984],"msr-impact-theme":[],"msr-promo-type":[],"msr-podcast-series":[],"class_list":["post-1187920","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-research-blog","msr-research-area-artificial-intelligence","msr-research-area-systems-and-networking","msr-locale-en_us","msr-post-option-blog-homepage-featured"],"msr_event_details":{"start":"","end":"","location":""},"podcast_url":"","podcast_episode":"","msr_research_lab":[199560],"msr_impact_theme":[],"related-publications":[],"related-downloads":[],"related-videos":[],"related-academic-programs":[],"related-groups":[815140,881388],"related-projects":[],"related-events":[],"related-researchers":[{"type":"user_nicename","value":"Zhiyuan He","user_id":41743,"display_name":"Zhiyuan He","author_link":"<a href=\"https:\/\/www.microsoft.com\/en-us\/research\/people\/zhiyuhe\/\" aria-label=\"Visit the profile page for Zhiyuan He\">Zhiyuan He<\/a>","is_active":false,"last_first":"He, Zhiyuan","people_section":0,"alias":"zhiyuhe"},{"type":"user_nicename","value":"Yuqing Yang","user_id":40654,"display_name":"Yuqing Yang","author_link":"<a href=\"https:\/\/www.microsoft.com\/en-us\/research\/people\/yuqyang\/\" aria-label=\"Visit the profile page for Yuqing Yang\">Yuqing Yang<\/a>","is_active":false,"last_first":"Yang, Yuqing","people_section":0,"alias":"yuqyang"}],"msr_type":"Post","featured_image_thumbnail":"<img width=\"960\" height=\"540\" src=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-960x540.jpg\" class=\"img-object-cover\" alt=\"System architecture diagram. On the left, agents with harnesses \u2014 mini-SWE-agent, OpenHands, and OpenClaw \u2014 run on a Kubernetes cluster. They connect to three components: an API Gateway containing a Rollout API and an LLM API Proxy, a Rollout Controller containing a local reconciler and a Kubernetes reconciler, and a Customized Trainer containing a sample adapter and monitoring. These connect in turn to an inference engine and a training engine holding the model.\" decoding=\"async\" loading=\"lazy\" srcset=\"https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-960x540.jpg 960w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-300x169.jpg 300w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1024x576.jpg 1024w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-768x432.jpg 768w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1065x599-1.jpg 1065w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1279x720.jpg 1279w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-653x368.jpg 653w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-675x380.jpg 675w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-240x135.jpg 240w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-640x360.jpg 640w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1-1280x720.jpg 1280w, https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2026\/09\/AgentLightning-BlogHeroFeature-1400x788-1.jpg 1400w\" sizes=\"auto, (max-width: 960px) 100vw, 960px\" \/>","byline":"<a href=\"https:\/\/www.microsoft.com\/en-us\/research\/people\/zhiyuhe\/\" title=\"Go to researcher profile for Zhiyuan He\" aria-label=\"Go to researcher profile for Zhiyuan He\" data-bi-type=\"byline author\" data-bi-cN=\"Zhiyuan He\">Zhiyuan He<\/a> and <a href=\"https:\/\/www.microsoft.com\/en-us\/research\/people\/yuqyang\/\" title=\"Go to researcher profile for Yuqing Yang\" aria-label=\"Go to researcher profile for Yuqing Yang\" data-bi-type=\"byline author\" data-bi-cN=\"Yuqing Yang\">Yuqing Yang<\/a>","formattedDate":"October 7, 2026","formattedExcerpt":"Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.","locale":{"slug":"en_us","name":"English","native":"","english":"English"},"_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/posts\/1187920","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/users\/43868"}],"replies":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/comments?post=1187920"}],"version-history":[{"count":14,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/posts\/1187920\/revisions"}],"predecessor-version":[{"id":1188803,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/posts\/1187920\/revisions\/1188803"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media\/1188789"}],"wp:attachment":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1187920"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/categories?post=1187920"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/tags?post=1187920"},{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1187920"},{"taxonomy":"msr-region","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-region?post=1187920"},{"taxonomy":"msr-event-type","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-event-type?post=1187920"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1187920"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=1187920"},{"taxonomy":"msr-impact-theme","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-impact-theme?post=1187920"},{"taxonomy":"msr-promo-type","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-promo-type?post=1187920"},{"taxonomy":"msr-podcast-series","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-podcast-series?post=1187920"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}