Atlas · skill

Federated Learning

Federated learning trains a shared model from data held by separate participants without routinely collecting all raw examples in one place. The competence is designing local updates, aggregation and evaluation under heterogeneous data and unreliable participation. Keeping data local is an architectural property, not a complete privacy guarantee.

conceptTraining Infrastructure

What it is

A coordinator distributes model parameters to selected participants, each trains locally, and an aggregation rule combines returned updates. Federated averaging weights local model contributions according to an agreed procedure, commonly related to local data volume. Participants may be devices or organizations and can have different data distributions, compute capacity and availability. Multiple local steps reduce communication but can increase divergence between participants. The system must account for who contributes and how updates are validated. Secure aggregation and differential privacy are additional mechanisms with separate assumptions; the ordinary federated training protocol does not automatically prevent sensitive information from appearing in transmitted updates.

What the work involves

Define participants, participation rules and the trust model before choosing a protocol. Measure data heterogeneity and compare local-only, centralized reference where permitted and federated baselines. Configure local steps, weighting and communication rounds, and test dropped or delayed participants. Keep evaluation users or organizations separate from training participants where the generalization question requires it. Evaluate each participant group as well as the aggregate. The result includes an aggregation procedure, privacy and robustness controls appropriate to the threat model, and evidence that the shared model serves participants beyond those dominating the update volume.

Illustrative example

In an illustrative collaboration, several laboratories hold differently distributed sensor observations. Each trains a shared classifier locally and returns an update. Evaluation shows that an aggregation favoring large laboratories performs poorly on a smaller site's operating conditions. The team inspects per-site results, changes the participation and weighting policy and tests a previously unseen laboratory. They also assess update exposure separately, because no raw-data transfer alone does not establish confidentiality.

Limits and common mistakes

Nonidentical data, uneven participation and stale updates can slow convergence or disadvantage small groups. Malicious or faulty updates can corrupt aggregation, and ordinary updates may leak information. Privacy protections introduce their own utility and system costs. A centralized test set can conceal participant-specific failures, while local tests may not be comparable. Federated learning differs from distributed training over centrally managed data and from secure multiparty computation. State the actual trust and privacy mechanisms and evaluate both shared and participant-level utility.

Prerequisites

No prerequisites.

Related skills

Sources and further reading

Last updated: 2026-10-10