Atlas · skill

Agent Evaluation

Agent evaluation measures whether an agent completes tasks correctly while using tools, state and permissions appropriately across an execution trajectory. It goes beyond scoring a final sentence: the evaluator may need to inspect actions, intermediate evidence, environment changes and whether the result actually satisfies the user's goal.

conceptEvaluation Design

What it is

An agent interacts with an environment through several decisions and actions. Its final answer can sound successful even when no requested change occurred or an unauthorized action was taken. Evaluation therefore combines outcome checks with trajectory or policy checks. Benchmarks such as WebArena and τ-bench provide specific environments and task protocols, rather than one universal agent metric. A trajectory can be assessed for tool argument correctness, recovery behavior or unnecessary actions, while outcome evaluation inspects the resulting state. Equivalent valid paths should be allowed where the task does not prescribe one exact sequence.

What the work involves

The practitioner defines success in observable environment terms and separates it from permission and efficiency constraints. Test environments reset state and avoid unintended external effects. Repeated trials capture stochastic reliability, and traces retain tool calls and results for review. Useful artifacts include task fixtures, outcome validators, policy checks and trajectory annotations. A failure taxonomy distinguishes planning errors, wrong tool use, missing evidence and execution failures. Evaluation also checks whether the agent stops and reports uncertainty when the allowed environment cannot support completion.

Illustrative example

A support agent must update a customer's delivery preference after confirming the relevant account. The test verifies the stored preference, the account used and the absence of unrelated changes. Another case denies the update permission and expects an appropriate escalation. A final message saying the update is complete earns no success credit if the database state is unchanged. Conversely, a shorter valid action sequence is accepted when it achieves the same authorized outcome without following the evaluator's preferred wording.

Limits and common mistakes

Outcome checks can be incomplete, simulated environments may miss real constraints and model-based trajectory judges can misclassify actions. Repeated tasks can also leak into agent tuning. Quality reports should state the environment, success validator and allowed budget. Reliable behavior on one benchmark does not establish broad autonomy. Evaluation needs realistic permissions, interruptions and failure paths as well as successful tasks, with state inspection independent of the agent's own account of its work.

Prerequisites

Sources and further reading

Last updated: 2026-10-10