The current AI revolution originated from the impressive deep learning methodology. Deep neural networks and in particular recent foundation models rely heavily on vast amounts of data and computing power, which poses the question of how to efficiently train large AI models from big data distributed geographically in multiple organisations. Because of privacy concerns and commercial interest of the data providers, as well as mandatory regulation rules such as General Data Protection Regulation (GDPR) of the European Union, sensitive local data must not be leaked to others. This restriction makes traditional distributed learning paradigms an unfeasible solution in which local information can freely flow among nodes. Federated learning is an emerging distributed learning paradigm that features privacy protection by utilising sophisticated algorithms such as differential privacy and aggregation schemes to generate global results. However, while federated learning has made great progress and been deployed in many industrial systems, there remain obstacles in the canonical paradigm that prevent its advance in the era of large models. In this thesis, two important difficulties are identified: low training efficiency and lack of explainability. Through experimental results from the literature and collected operating stat- istics of deployed federated systems, the main training bottleneck boils down to the transfer time of large messages, which could add up to more than one day in a typical federated training across distant sites. On the other hand, explainability can help evaluate the quality of training samples, shorten training time, and facilitate the acceptance of complex AI models. These advantages are valuable in federated environments because there is no authoritative control over the data quantity of participants. Motivated by the aforementioned challenges of large-scale federated learning, the topic of this thesis is efficient and explainable large-scale federated learning. An extensive investigation of the literature and our works on the topic are presented. Specifically, our contributions are as follows: • We propose Hy pergradient for Data Relevance Analysis (HyDRA), which is an explainable approach that interprets the predictions made by deep neural networks as effects of training samples. Existing approaches like Influence Function generally estimate training sample contributions around the final model parameters and ignore how the training samples shape the optimisation trajectory. In contrast, HyDRA assesses the contribution of training samples to test samples throughout the optimisation trajectory. In order to accelerate computation, we provide an approximation formula and prove that, under moderate conditions, the approximation error is bounded. Consistent with the theoretical claim, empirical results indicate that the error is indeed small. In addition, we quantitatively demonstrate that HyDRA outperforms Influence Function in accurately estimating data contribution and detecting noisy data labels. • We propose the Guided T runcation G radient Shapley (GTG-Shapley) ap- proach to fairly evaluate the contributions of the participants to the perform- ance of the final model without exposing their private data. To sustain the long-term operation of a federated learning (FL) ecosystem, it is important to attract high-quality data owners with appropriate incentive schemes. Fair evaluation of participant contributions is an important building block of such incentive schemes. Although Shapley value-based techniques have been widely adopted to provide fair evaluation in such cases, existing approaches incur a significant computational cost, making them difficult to apply in practice. GTG-Shapley approximates sub-coalitions by aggregating sub- models, rather than repeatedly training from scratch. It also uses Monte Carlo sampling and other optimisations to further reduce the number of sub- coalition evaluations. The experimental results showed that GTG-Shapley can closely approximate the actual Shapley value, while significantly increasing the computational efficiency compared to other advanced methods, especially in non-i.i.d. settings. • We propose Fed erated O pportunistic B lock D ropout (FedOBD), which provides an efficient framework for federated learning. The key novelty is that it decomposes large neural networks into semantic blocks so that federated participants can opportunistically upload quantised blocks, which can greatly speed up training. Extensive experiments evaluating FedOBD against four competitive approaches based on common datasets show that it reduces the overall communication overhead by more than 88% compared to the best-performing baseline approach, while achieving the highest test accuracy. FedOBD was deployed in industry and helped the company reduce the training communication overhead by more than 70% compared to its previous AI engine, while maintaining model performance at a test F1 score of over 85%. Finally, we discuss further research directions, including: (1) integrating the pro- posed methods into a unified pipeline that provides hierarchical explanations at both sample and participant levels while maintaining communication efficiency; (2) extending the efficient communication techniques to federated graph learning, where convolutional operations over distributed sub-graphs incur high communication costs; and (3) balancing explainability and productivity in deployed industrial AI systems.
Read more