Self-Improving Agents
Self-improving agents generate changes to components that govern their own future behavior, such as prompts, code or tool-selection policies. A separate evaluation and acceptance process decides whether a proposed change is useful, making improvement an experimentally tested system modification rather than an agent's claim about itself.
What it is
The system exposes some part of its implementation as a candidate for revision. An agent proposes a change, evaluates the candidate on specified tasks and selects or rejects it under a comparison procedure. The modified component may be an inference-time program, not the underlying model weights. This distinguishes self-improvement from ordinary self-refinement, which revises an answer within a task, and from model retraining, which updates learned parameters. In a coding-agent setting, the search process can edit the agent's scaffold while keeping the evaluation harness fixed. The scope of editable components determines what the process can actually improve.
What the work involves
The practitioner identifies editable components, freezes an evaluation protocol and keeps candidate versions recoverable. Evaluation should include held-out tasks, resource costs and behavior that the proposed change might break. Acceptance needs an external decision rule rather than allowing the candidate to redefine success. Useful artifacts include candidate diffs, comparison results and a rollback path. Restricting access to evaluator code and deployment credentials separates generating an improvement proposal from granting it authority over the system that judges or releases that proposal.
Illustrative example
A coding agent proposes replacing a broad file-reading step with targeted search in its own task scaffold. The candidate is evaluated on unfamiliar repositories alongside the existing scaffold. The comparison checks task completion, missed context and resource use. A change that works on the development tasks but fails on an unseen project is rejected. An accepted candidate remains versioned so later regressions can be traced to the modified search policy.
Limits and common mistakes
Optimizing a fixed evaluator can overfit its tasks or exploit weaknesses in its measurements. More successful attempts under a larger budget do not necessarily establish a better agent at equal cost. An agent should not be allowed to alter the acceptance criterion unnoticed. Quality requires preserved evaluation boundaries, reproducible comparisons and monitoring after adoption. The term does not imply unlimited recursive improvement or general progress; it describes a bounded search over changes that a particular system can generate and test.
Prerequisites
- hardAI Agent Design
Self-improvement extends the agent architecture.
Self-rewriting agents optimize their own prompts.
Sources and further reading
- A Self-Improving Coding Agent
Studies an agent that modifies its own coding scaffold and evaluates candidate changes in a controlled task setting.
Last updated: 2026-10-10