Skip to content
The AI-First Web

Microsoft researchers train AI agents inside the real harness they will run in

Agent Lightning v1.0 trains agents in the same harness they will be deployed in, not a rebuilt copy

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Microsoft researchers train AI agents inside the real harness they will run in
In brief
  • Microsoft Research Asia released Agent Lightning v1.0, which trains an AI agent inside the same harness used in deployment instead of a rebuilt copy.
  • In Microsoft's test, a 9B-parameter model rose from 41.8% to 56.4% on SWE-bench Verified. The result comes from one model and one benchmark.
  • Leaders should ask whether tested and deployed agents match, and who controls the store of recorded prompts and responses.

Teams that improve an AI agent often train a stand-in. They rebuild the agent inside the training software, tune the copy, and ship the original. Microsoft Research says that means the agent being trained is "not quite the agent that gets deployed." The lesson is that the gap between practice and production is a real engineering cost, one Microsoft now designs around. The report does not measure the size of that gap.

What Microsoft released

Researchers at Microsoft Research Asia open-sourced Agent Lightning v1.0 and published it on October 7. They call the approach Harnessed Agentic RL. The harness is the software around a model that manages its tools, memory and environment. Tools like OpenHands, Claude Code and Codex ship with their own version of that setup.

RL, or reinforcement learning, is training by trial and error. The system rewards good actions and penalizes bad ones. The researchers' rule is simple: whatever harness runs in deployment is the one that takes part in training.

How it works

The design puts a relay in the middle. The agent sends its requests to Agent Lightning instead of straight to the model, and carries on as before. A team points the address that used to reach the model at the relay. The researchers say this is usually enough to connect an existing harness to training.

Three parts make up the framework. The API gateway is the relay. It ties each model call to its run and keeps a log of what was asked, what came back, and log probabilities, which are numbers showing how likely the model judged each of its own word choices. A rollout controller starts agents as local processes or standard Kubernetes jobs. A trainer, built on the verl framework, collects the results and assembles training samples.

About 3,500 lines of code
Size of the framework
Source: Microsoft Research, Agent Lightning v1.0 (October 7, 2026)

Why this is harder than it sounds

In older systems, the trainer ran the agent's loop and saw everything as one continuous stream. Now the harness runs the loop. The trainer sees only pairs of requests and responses. One run can split into a varying number of training samples.

That creates bookkeeping traps. Harnesses keep context as text, and converting it back to tokens can shift boundaries. Scoring each sample separately also counts runs that produce more samples more than once. Think of a class where the student who hands in more drafts gets more votes.

The researchers score and normalize at the level of the whole run instead. They report higher validation reward and steadier policy entropy than sample-level handling.

What the results show, and do not show

Microsoft built a full pipeline using SWE-smith, mini-SWE-agent and the Qwen3.5-9B model. It covered data cleaning, environment setup, safeguards against reward hacking (an agent gaming the score) and training. RL training alone lifted the model on SWE-bench Verified, a coding benchmark.

41.8% to 56.4% (+14.6 points)
Gain on SWE-bench Verified
Source: Microsoft Research, Agent Lightning v1.0 (October 7, 2026)

Treat this as one data point. It covers one model, one benchmark and one harness, and the figures are Microsoft's own. The report does not describe independent replication. It also does not say how the gain carries over to a company's private codebase.

Where the money goes

Running many agents at once takes heavy compute. The researchers say some competing frameworks rent commercial sandboxes, such as Modal Sandbox or E2B, to host agents, and that the bill rises fast as runs multiply. Agent Lightning instead uses ordinary Kubernetes jobs on capacity a team may already have, whether self-managed, in the cloud or local.

A second saving is on GPUs. Its Collocated Async RL lets rollouts and model updates share the same GPUs.

About 2x end to end
Speedup over synchronous RL
Source: Microsoft Research, Agent Lightning v1.0 (October 7, 2026)

Questions for your team

First, ask whether the agent you evaluate is the agent you deploy. If a rebuilt copy sits in between, test results describe the copy.

Second, ask where the recorded traffic goes. The gateway stores prompts and responses for every model call. This is WebPulse's view, not a finding of the report. If agents touch source code or customer data, that store needs the same access controls as the data itself.

Third, ask who defines the reward and who checks for gaming. Microsoft lists reward-hacking safeguards as part of its pipeline. An agent that learns to please the score can look better on paper than in use.

Fourth, ask what sandbox services cost today compared with the Kubernetes capacity you already own.

An agent is only as dependable as the match between its practice and its job. Training in the deployed harness is meant to remove one source of mismatch, and it makes the record of that practice something worth guarding.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Microsoft Research.

Share this insight