VeriFed: Cryptographic Accountability for Privacy-Preserving and Robust Federated LLM Fine-Tuning
Federated fine-tuning of large language models (LLMs) allows different participants to fine-tune a shared model, without any of these participants accessing the raw data. Rather than having sensitive information concentrated, parties train locally and provide only model updates, which enables the system to learn through diverse sources of data without breaching privacy and ownership of data; however, it is severely vulnerable to two potential attacks, namely (1) malicious clients may actively inject poisoned gradients in training and hence undermine integrity of the global model. and (2) accidental memorization and leakage of personally identifiable information (PII) during model updates. Traditional defense methods such as powerful aggregation schemes (e.g. Krum), and differential privacy can either protect against adaptive adversaries with limited protection or introduce substantial utility loss in large scale training of LLM. We can resolve this by indicating a privacy related federated framework called VeriFed with proposing a client to suitprove that it satisfies, under agreed upon conditions of privacy and safety, succinct cryptographic non-interactive arguments of knowledge (SNARKs). These are ℓ2 sensitivity bounded (to be DP-compliant), minimal embedding of PII tokens and a structured parameter efficient architecture (e.g., LoRA). Noteworthy, actual evidence is generated without disclosing the crude gradients or data, and secrecy is thus kept. We have a verification algorithm, using gradient drawing, hierarchical break down of restrictions, allowing scale reductions of up to billions of constraints in high-dimensional, high-order verification. VeriFed, evaluated on PubMedQA and Financial PhraseBank across five state-of-the-art LLMs, shows VeriFed achieves close utility, with a performance loss of less than 1% compared to FedAvg, and backdoor attacks success rates of 1.8%, corresponding to a suppression of over 98%. The leakage of privacy is reduced to a minimum, exposure to personally identifiable information (PII) = 0.08, which is 89% lower (LiRA = 0.08). Generation of proofs is still computationally viable, and takes less than 19 seconds to generate proofs with models of up to 33 million parameters. On the whole, VeriFed makes collaboration between LLM fine-tuning and jointly held accountability verifiable, because it is based on cryptographic assurances of reliability and not solely on statistical credibility.
Read more