What is federated learning?
Federated learning trains a shared model on data that stays on the devices that hold it: only the model’s updates travel. What it is, and when it is useful.

Federated learning trains one model on data held in several places without moving the data: each site trains on what it holds and sends only the model's updates to a coordinator, which averages them. It answers one question, how to learn from data that cannot leave its site, and leaves the others open: the labels still have to be made on each site, and the updates themselves have to be protected.
Where it comes from
The method was named and formalised by McMahan et al. (2017) at Google, for training keyboard models on phones without uploading what people type. Their algorithm, federated averaging, is still the reference: a coordinator sends the current model to a set of participants, each trains it for a few passes on its own data, each sends back its new weights, and the coordinator replaces the model by the weighted average of what it received. One such exchange is a round; a training run is a few hundred to a few thousand rounds. A survey by Kairouz et al. (2021), with 58 authors from 25 institutions, lists the open problems the first four years had raised.
Two families
| Cross-device | Cross-silo | |
|---|---|---|
| Participants | Millions of phones or sensors, few available at a time | A handful of organisations or sites, all available |
| Data per participant | Small | Large, often years of recordings |
| Reliability | Participants drop out mid-round | Participants stay for the run |
| Example | Next-word prediction on a keyboard | Hospitals, banks, or sites of one programme that each keep their recordings |
The distinction is Kairouz et al.'s. The cross-silo case is the one that concerns signal work: a few sites, each with a large archive that cannot be copied, and a network between them that may be slow or intermittent.
What it does not solve
- Labels. A supervised model needs labeled examples on every site, made with the same taxonomy and the same rules. Federated learning moves the training, not the annotation.
- Leakage through the updates. Model weights are computed from the data and can reveal some of it. Secure aggregation (Bonawitz et al., 2017) lets the coordinator see only the sum of the updates, never one participant's; differential privacy bounds what any single record can change. Both cost accuracy or rounds.
- Heterogeneous data. Sites record with different sensors, in different environments, at different sampling rates. Federated averaging assumes the sites are alike; when they are not, the model drifts towards the largest site or converges slowly. Adapting a network to each carrier is a problem of its own, the one Ezako and XPERT took up in 2025 for underwater platforms.
- Isolated networks. A site with no connection at all cannot take part in rounds. The plain alternative works: train on the site, validate there, and carry the model, not the data.
When it is the right tool
Federated learning is worth its complexity when three conditions hold at once: several sites hold data of the same kind, none of them may release it, and a model trained on one site alone is measurably worse than a model trained on all. When only the first two hold, a model trained and validated on each site, then compared, is simpler and easier to audit. When the third holds but the sites can share a dataset under agreement, a central dataset with immutable versions gives a traceable model and a test set everyone can read.
Ezako's software is built for the case where data stays where it is. Upalgo Labeling and UpalgoDB run on the customer's infrastructure, including air-gapped networks; models are trained on the customer's datasets and integrated into UpalgoDB there; no data is transmitted to Ezako. Security and the engineering principles describe how.
Sources
- H. B. McMahan, E. Moore, D. Ramage, S. Hampson, B. Agüera y Arcas, Communication-Efficient Learning of Deep Networks from Decentralized Data, AISTATS, 2017.
- P. Kairouz et al., Advances and Open Problems in Federated Learning, Foundations and Trends in Machine Learning 14(1-2), 2021.
- K. Bonawitz et al., Practical Secure Aggregation for Privacy-Preserving Machine Learning, ACM CCS, 2017.
- Ezako, Ezako and XPERT: deep learning for every underwater platform, 2025.
First published in January 2022; rewritten in September 2026 with the origin of the method, the two families, what it leaves unsolved and when it is worth using.