Training agents in real harnesses makes observability part of model improvement
Microsoft Research has rebuilt Agent Lightning around reinforcement learning inside existing agent harnesses. The practical skill shift is not just reward design: teams must make tool actions, context and failures inspectable enough to train on.

What happened
Microsoft Research published a 7 October technical overview of Agent Lightning v1.0, a roughly 3,500-line open-source framework designed to train agents while retaining their real tools, context and control flow.
Why it matters
Training inside an operational harness can narrow the gap between a benchmark and deployed work, but only if traces distinguish model choices from orchestration, tool and environment failures.
Microsoft Research described Agent Lightning v1.0 on 7 October. The framework puts a proxy between an existing agent harness and an OpenAI-like model API so reinforcement learning can keep tools, context and control flow in the loop. The technical report describes an implementation of roughly 3,500 lines and a Harnessed Agentic RL approach.
This is an architecture and research release, not evidence that reinforcement learning will improve every production agent. Results depend on reward validity, task coverage, environment stability and the quality of traces.
Treat observability as training data
A team cannot assign useful credit if it cannot tell whether failure came from the model, prompt, memory, tool, permission, orchestrator or environment. Instrument each step with stable action identifiers, inputs, outputs, latency, permission checks and failure categories. Preserve unsuccessful paths; deleting them biases the learning signal.
Reward design should combine task completion with constraints such as provenance, reversibility, cost and safe abstention. Before any policy update, replay held-out tasks and compare new failures, not just average reward. Keep a rollbackable policy version and a human-readable incident slice.
The counterargument is that richer traces increase storage and expose sensitive content. Use selective capture, redaction and short retention, but keep enough structure to reproduce the failure. The capability decision is whether the harness produces trustworthy learning evidence before the team spends on reinforcement learning infrastructure.
Start with an offline shadow run. Let the candidate policy observe the same tasks without controlling production tools, then compare its proposed actions with the existing agent and human outcomes. Track reward disagreement separately from execution failure: a policy can optimise the recorded score while violating the real objective. Promotion should require stable gains across held-out tasks, no new critical failure class and an auditable link from reward to trace. That is the minimum operational gate.