Atlas · skill

Multi-armed Bandits

Multi-armed bandits learn which actions to choose while receiving feedback only for the selected action. They balance exploring uncertain alternatives with exploiting those that currently appear best. The competence includes reward design, exploration policy and evaluation of partially observed outcomes, especially when historical data were collected by a different selection policy.

conceptSequential Decision-Making

What it is

A bandit repeatedly selects an arm and observes its reward. In the basic setting, each choice is evaluated by its immediate reward without modeling how it changes a future state. Contextual bandits condition the action choice on information available for the current decision. Exploration strategies such as uncertainty-based selection or randomization allow the system to learn about less-used alternatives. Logged feedback is selective: the reward for an unchosen action is usually unknown. This distinguishes bandit learning from ordinary supervised learning, where each example can provide its target directly, and from broader reinforcement learning with delayed consequences through state transitions.

What the work involves

Define eligible actions, context and a reward that arrives on a usable timescale. Select an exploration policy compatible with operational constraints and record action probabilities, chosen actions and outcomes. Compare against a fixed policy and assess whether offline evaluation methods have sufficient overlap with the candidate policy. Monitor reward drift and performance across contexts. The practical result is an auditable selection policy with a deliberate exploration budget, rather than a greedy ranking system that never gathers evidence about alternatives it initially undervalued.

Illustrative example

Imagine an illustrative help center choosing among several article suggestions. A contextual bandit uses the current question category and learns from whether the user resolves the issue. Some eligible articles are occasionally explored, with their selection probability recorded. An analyst checks whether resolution feedback is missing more often for particular users and compares the policy with a fixed suggestion rule. The system cannot infer the value of an article never shown solely from the outcomes of other articles.

Limits and common mistakes

Rewards can be delayed, biased or misaligned with user benefit. Insufficient exploration makes alternatives hard to assess, while excessive exploration can reduce short-term quality. Offline estimates require assumptions about logging probabilities and action coverage and may have high variance. A basic bandit is inappropriate when choices materially change future states or long-term opportunities. Distinguish genuine reward changes from changes in exposure or measurement, and evaluate the policy rather than only the accuracy of its internal reward predictor.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10