Federated Learning
Federated learning trains a shared model from data held by separate participants without routinely collecting all raw examples in one place. The competence is designing local updates, aggregation and evaluation under heterogeneous data and unreliable participation. Keeping data local is an architectural property, not a complete privacy guarantee.
What it is
A coordinator distributes model parameters to selected participants, each trains locally, and an aggregation rule combines returned updates. Federated averaging weights local model contributions according to an agreed procedure, commonly related to local data volume. Participants may be devices or organizations and can have different data distributions, compute capacity and availability. Multiple local steps reduce communication but can increase divergence between participants. The system must account for who contributes and how updates are validated. Secure aggregation and differential privacy are additional mechanisms with separate assumptions; the ordinary federated training protocol does not automatically prevent sensitive information from appearing in transmitted updates.
What the work involves
Define participants, participation rules and the trust model before choosing a protocol. Measure data heterogeneity and compare local-only, centralized reference where permitted and federated baselines. Configure local steps, weighting and communication rounds, and test dropped or delayed participants. Keep evaluation users or organizations separate from training participants where the generalization question requires it. Evaluate each participant group as well as the aggregate. The result includes an aggregation procedure, privacy and robustness controls appropriate to the threat model, and evidence that the shared model serves participants beyond those dominating the update volume.
Illustrative example
In an illustrative collaboration, several laboratories hold differently distributed sensor observations. Each trains a shared classifier locally and returns an update. Evaluation shows that an aggregation favoring large laboratories performs poorly on a smaller site's operating conditions. The team inspects per-site results, changes the participation and weighting policy and tests a previously unseen laboratory. They also assess update exposure separately, because no raw-data transfer alone does not establish confidentiality.
Limits and common mistakes
Nonidentical data, uneven participation and stale updates can slow convergence or disadvantage small groups. Malicious or faulty updates can corrupt aggregation, and ordinary updates may leak information. Privacy protections introduce their own utility and system costs. A centralized test set can conceal participant-specific failures, while local tests may not be comparable. Federated learning differs from distributed training over centrally managed data and from secure multiparty computation. State the actual trust and privacy mechanisms and evaluate both shared and participant-level utility.
Prerequisites
Related skills
- → is subcategory of: Distributed Training
Sources and further reading
- Communication-Efficient Learning of Deep Networks from Decentralized Data
Federated averaging, local updates and decentralized heterogeneous data.
- Advances and Open Problems in Federated Learning
Research account of heterogeneity, privacy, robustness and deployment constraints.
Last updated: 2026-10-10