- Research Article
477
- 10.1016/j.ins.2018.10.024
Secure Multi-Party Computation: Theory, practice and applications
- Oct 16, 2018
- Information Sciences
- Chuan Zhao + 6 more +6
Secure Multi-Party Computation: Theory, practice and applications
In a Secure Multiparty Computation (SMC), mutually distrusting parties use cryptographic techniques to cooperatively compute over their private data, in the process each party learns only explicitly revealed outputs. In this paper, we present Wysteria, a high-level programming language for writing SMCs. As with past languages, like Fairplay, Wysteria compiles secure computations to circuits that are executed by an underlying engine. Unlike past work, Wysteria provides support for mixed-mode programs, which combine local, private computations with synchronous SMCs. Wysteria complements a standard feature set with built-in support for secret shares and with wire bundles, a new abstraction that supports generic n-party computations. We have formalized Wysteria, its refinement type system, and its operational semantics. We show that Wysteria programs have an easy-to-understand single-threaded interpretation and prove that this view corresponds to the actual multi-threaded semantics. We also prove type soundness, a property we show has security ramifications, namely that information about one party's data can only be revealed to another via (agreed upon) secure computations. We have implemented Wysteria, and used it to program a variety of interesting SMC protocols from the literature, as well as several new ones. We find that Wysteria's performance is competitive with prior approaches while making programming far easier, and more trustworthy.
Secure Multi-Party Computation: Theory, practice and applications
Secure Multi-Party Computation: Theory, practice and applications
Metamorphic Testing of Secure Multi-party Computation (MPC) Compilers
The demanding need to perform privacy-preserving computations among multiple data owners has led to the prosperous development of secure multi-party computation (MPC) protocols. MPC offers protocols for parties to jointly compute a function over their inputs while keeping those inputs private. To date, MPC has been widely adopted in various real-world, privacy-sensitive sectors, such as healthcare and finance. Moreover, to ease the adoption of MPC, industrial and academic MPC compilers have been developed to automatically translate programs describing arbitrary MPC procedures into low-level MPC executables. Compiling high-level descriptions into high-efficiency MPC executables is challenging: the compilation often involves converting high-level languages into several intermediate representations (IR), e.g., arithmetic or boolean circuits, optimizing the computation/communication cost, and picking proper MPC protocols (and underlying virtual machines) for a particular task and threat model. Various optimizations and heuristics are employed during the compilation procedure to improve the efficiency of the generated MPC executables. Despite the prosperous adoption of MPC compilers by industrial vendors and academia, a principled and systematic understanding of the correctness of MPC compilers does not yet exist. To fill this critical gap, this paper introduces MT-MPC, a metamorphic testing (MT) framework specifically designed for MPC compilers to effectively uncover erroneous compilations. Our approach proposes three metamorphic relations (MRs) that are tailored for MPC programs to mutate high-level MPC programs (compiler inputs). We then examine if MPC compilers yield semantics-equivalent MPC executables regarding the original and mutated MPC programs by comparing their execution results. Real-world MPC compilers exhibit a high level of engineering quality. Nevertheless, we detected 4,772 inputs that can result in erroneous compilations in three popular MPC compilers available on the market. While the discovered error-triggering inputs do not cause the MPC compilers to crash directly, they can lead to the generation of incorrect MPC executables, jeopardizing the underlying dependability of the computation. With substantial manual effort and help from the MPC compiler developers, we uncovered thirteen bugs in these MPC compilers by debugging them using the error-triggering inputs. Our proposed testing frameworks and findings can be used to guide developers in their efforts to improve MPC compilers.
Read moreConsiderations for using Privacy Preserving Machine Learning Techniques for Safeguards
In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.
Read moreAchieving Privacy-preserving Computation on Data Grids
This paper proposes a generic Grid privacy-preserving computation(G2PC) model which supports privacy-preserving data analysis and computation on multiple distributed datasets without compromising both the raw data privacy of Grid nodes and data statistics (intermediate result) privacy. The center of the design is our novel Data Privacy-Preserving Broker (D2PB) that combines the GSI (Grid Security infrastructure) with a number of cryptographic primitives. G2PC model requires neither one-to-all interactions among participating entities, nor reassignment of security parameters when membership or data changes. Therefore, it is efficient, scalable, and suited to large-scale Data Grid systems that are expected to host thousands of dynamic nodes. The privacy-preserving variance computation and privacy-preserving k-means clustering algorithm have been used as examples to demonstrate the efficacy and efficiency of our proposed framework.
Read moreSPRINT: Scalable Secure & Differentially Private Inference for Transformers
Machine learning as a service (MLaaS) enables scalable model deployment and inference on cloud servers. However, MLaaS exposes user queries and model parameters to servers. To guarantee confidentiality of queries and model parameters, multi-party computation (MPC) enables secure inference by distributing data and computations across multiple service providers. MPC eliminates single points of failure, mitigates provider breaches and ensures confidentiality beyond legal agreements. Beyond confidentiality of queries and parameters, the model itself can memorize and leak training data during inference. To mitigate privacy concerns, differential privacy (DP) provides a formal privacy guarantee for training data, which can be satisfied by injecting carefully calibrated noise into gradients during training. However, naive combinations of DP and MPC amplify accuracy loss due to DP noise and MPC approximations, and incur high computational and communication overhead due to cryptographic operations. We present SPRINT, the first scalable solution for efficient MPC inference on DP fine-tuned models with high accuracy. SPRINT fine-tunes public pre-trained models on private data using DP. It integrates DP-specific optimizations, e.g., parameter-efficient fine-tuning and noise-aware optimizers, with MPC optimizations, e.g., cleartext public parameters and efficient approximations of non-linear functions. We evaluate SPRINT on GLUE benchmark with RoBERTa, achieving up to 1.6x faster MPC inference than the state-of-the-art non-DP solution SHAFT, reducing communication by 1.6x. Notably, SPRINT maintains high accuracy during MPC inference, with <1 percentage point gap compared to cleartext accuracy.
Read morePrivacy-preserving computation scheme for the maximum and minimum values of the sums of keyword-corresponding values in cross-chain data exchange
In the specific value computation scenario of blockchain cross-chain data exchange, participants can organize raw data in the form of data pairs, making the computed specific values become key information for cross-chain collaboration. The main challenges faced in this scenario include: the lack of appropriate privacy protection mechanisms leading to easy leakage of sensitive data; the vulnerability of unencrypted or unvalidated data to tampering or malicious attacks; a crisis of trust among users. To address the above problems, this paper introduces secure multi-party computation (MPC) into the cross-chain data exchange process to ensure the fairness of the computation process and protect data privacy. To address the problem of computing the maximum and minimum values of the sums of keyword-corresponding values in cross-chain interactions, it is transformed into the secure computation scheme of computing the maximum and minimum values of the sums of corresponding elements in the intersection of sets (MMSI). A secure computation scheme for maximum and minimum values based on the fully homomorphic NTRU (FH-NTRU) encryption algorithm is proposed. First, a secure computation protocol for MMSI without a universal set under the semi-honest model is proposed. Then, to address potential malicious behaviors in the protocol, a secure computation protocol for MMSI under the malicious model is proposed, using the cut-and-choose method. The protocol under the malicious model is analyzed for correctness and the security of the protocol is proved using the real/ideal model paradigm. Finally, the efficiency analysis and experimental simulations show that the protocol is efficient and reliable, it can resist malicious adversary attacks and ensure the correctness of the computation results while effectively improving the security during cross-chain interactions.Supplementary InformationThe online version contains supplementary material available at 10.1038/s41598-025-14629-1.
Read moreSecure multiparty computation protocol based on homomorphic encryption and its application in blockchain
Secure multiparty computation protocol based on homomorphic encryption and its application in blockchain
Fair and Secure Multi-Party Computation with Cheater Detection
Secure multi-party computation (SMC) is a cryptographic protocol that allows participants to compute the desired output without revealing their inputs. A variety of results related to increasing the efficiency of SMC protocol have been reported, and thus, SMC can be used in various applications. With the SMC protocol in smart grids, it becomes possible to obtain information for load balancing and various statistics, without revealing sensitive user information. To prevent malicious users from tampering with input values, SMC requires cheater detection. Several studies have been conducted on SMC with cheater detection, but none of these has been able to guarantee the fairness of the protocol. In such cases, only a malicious user can obtain a correct output prior to detection. This can be a critical problem if the result of the computation is real-time information of considerable economic value. In this paper, we propose a fair and secure multi-party computation protocol, which detects malicious parties participating in the protocol before computing the final output and prevents them from obtaining it. The security of our protocol is proven in the universal composability framework. Furthermore, we develop an enhanced version of the protocol that is more efficient when computing an average after detecting cheaters. We apply the proposed protocols to a smart grid as an application and analyze their efficiency in terms of computational cost.
Read moreA systemic comparison comparison of concurrent multiparty secret sharing with SGD regression and classification
In recent years Secure Multiparty Computation (MPC) via Secret Sharing has emerged as a key intersection between the fields of cryptography and machine learning. MPC's ability to allow N-number of participants to hide their inputs while able to share their outputs has presented itself as a solution for creating machine learning models for datasets with sensitive information such as those used in healthcare. Prior to or without MPC creating these models required that the information in these datasets would have to be shared with all participants in cleartext which then could be stolen or misused by dishonest participants or even in transit between the parties. As the demand for answers from machine learning models grows so does the need for privacy and confidentiality for data in those models. But MPC does have some pitfalls with both an increase in the cost of computations of secret shares and the increase in the amount of data transferred between the participating parties. In the context of my research, I Introduce A Systemic Comparison of Concurrent Multiparty Secret Sharing with SGD Regression and Classification in which I create and run experiments that compares different methods of MPC and different machine learning algorithms. The results of the experiments provide valuable insights into the practical viability of secure multiparty computation for machine learning applications. By comparing the performance of different secure computation methods, this research contributes to the understanding of the trade-offs between security and efficiency.
Read moreBuilding AI Pipelines That Comply with GDPR, HIPAA, and Industry Standards
The high take-up rate of artificial intelligence (AI) by different sectors has triggered critical interest in data security, privacy, and ethics compliance. This work discusses the design and process blueprint for the establishment of AI pipelines that comply with the General Data Protection Regulation (GDPR), the Health Insurance Portability and Accountability Act (HIPAA), and different sectoral standards. The aim is to offer a sound methodology to incorporate compliance checks at every stage of the AI lifecycle, from data collection to deployment. Our method focuses on incorporating privacy-preserving methods, secure data engineering principles, and auditability. Through a comparative analysis of AI development processes and regulatory requirements, we identify friction points and suggest technical as well as organizational solutions. The paper also analyzes case studies to test the proposed framework and addresses its implication for AI governance and ethical AI deployment. The research highlights the need for cross-disciplinary cooperation between legal professionals, data scientists, and organizational stakeholders to provide end-to-end compliance. We propose certain tools and technologies, including federated learning and secure multiparty computation, that enable the enforcement of regulatory requirements without undermining model effectiveness. This paper makes a contribution to current discussion by offering a guidebook to action for businesses to enact ethical AI while preserving competitive edge in a highly regulated digital environment. Keywords- AI pipeline, GDPR, HIPAA, data privacy, compliance, ethical AI, data governance, security standards, machine learning, privacy-preserving computation
Read moreSynCirc: Efficient Syn thesis of Depth-Optimized Circ uits from High-Level Languages
Secure Multi-Party Computation (MPC) enables secure computation on private data. Many of today’s efficient MPC protocols need a representation of the evaluated function as circuit composed of Boolean or Lookup Tables (LUTs). To improve the practicality of MPC, we present SynCirc, a hardware synthesis framework optimized for MPC applications. Built on Verilog and the open-source tool Yosys-ABC, SynCirc introduces custom libraries and constraints for multi-input AND gates, achieving up to 3× reduction in multiplicative depth and online rounds compared to TinyGMW (Demmler et al., CCS’15). <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SynCirc also offers an expanded library of efficient building blocks like comparison, multiplexers and equality checks, and incorporates Boolean and LUT circuits. For these building blocks, we achieve improvements in multiplicative depth/online rounds between 22.3% and 66.7% over ShallowCC (Büscher et al., ESORICS’16). Our evaluation using the FLUTE framework (Brüggemann et al., IEEE S&P’23) shows that SynCirc has 116× less online communication than the multi-input AND gate protocol of Trifecta (Faraji and Kerschbaum, PETS’23). <p xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">SynCirc introduces new capabilities, including enhanced support for High-Level Synthesis (HLS) with the XLS tool, enabling developers to create secure functions in C/C++ without the need for expertise in hardware definition languages like Verilog. SynCirc is an open-source toolchain that democratizes secure computation, simplifies circuit synthesis, and makes advanced privacy-preserving technologies more accessible.
Read moreA probabilistic look ahead of anonymization
Data anonymization is an expensive process, and sometimes the utility of the anonymized data may not justify the cost of anonymization. For example in a distributed setting where the data reside at different sites and needs to be anonymized without a trusted server, Secure Multiparty Computation (SMC) protocols need to be employed. However, the cost of SMC protocols could be prohibitive, and therefore the parties may want to look ahead of anonymization to decide if it is worth running the expensive SMC protocols. In this work, we describe a probabilistic fast look ahead of k-anonymization of horizontally partitioned data. The look ahead returns an upper bound on the probability that k-anonymity will be achieved at a certain utility where the utility is quantified by commonly used metrics from the anonymization literature. The look ahead process exploits prior information such as total data size, attribute distributions, or attribute correlations, all of which require simple SMC operations to compute. More specifically, given only statistics on the private dataset, we show how to calculate the probability that a mapping of values to generalizations will make a private dataset k-anonymous.
Read moreSecure error correction using multiparty computation
In the last couple of decades, error correction techniques play a prominent role in various scientific and engineering fields such as information theory, communication, networking, to name a few. These techniques are mainly utilized to locate and fix corrupted data over noisy channels. The data might be corrupted due to various reasons, for instance, communication failures, noise, or adversarial activities. On the other hand, data-privacy has been in the center of attention by many researchers in recent years. As such, it's important to be able to use error correction techniques over private data. This paper therefore proposes a secure error correction method by using secure multiparty computation (MPC). To the best of our knowledge, our proposed approach is the first solution in the literature. In secure MPC protocols, parties first share their private inputs by cryptographic primitives in order to jointly compute a function without revealing those private inputs. At the end of the protocol, only the function value will be revealed to all parties. Our secure MPC protocol efficiently implements the error-locator function of Berlekamp-Welch algorithm. After locating errors, we utilize another cryptographic technique, named enrollment protocol, to fix the errors.
Read moreImplementing Support for Pointers to Private Data in a General-Purpose Secure Multi-Party Compiler
Recent compilers allow a general-purpose program (written in a conventional programming language) that handles private data to be translated into a secure distributed implementation of the corresponding functionality. The resulting program is then guaranteed to provably protect private data using secure multi-party computation techniques. The goals of such compilers are generality, usability, and efficiency, but the complete set of features of a modern programming language has not been supported to date by the existing compilers. In particular, recent compilers PICCO and the two-party ANSI C compiler strive to translate any C program into its secure multi-party implementation, but they currently lack support for pointers and dynamic memory allocation, which are important components of many C programs. In this work, we mitigate the limitation and add support for pointers to private data and consequently dynamic memory allocation to the PICCO compiler, enabling it to handle a more diverse set of programs over private data. Because doing so opens up a new design space, we investigate the use of pointers to private data (with known as well as private locations stored in them) in programs and report our findings. Aside from dynamic memory allocation, we examine other important topics associated with common pointer use such as reference by pointer/address, casting, and building various data structures in the context of secure multi-party computation. This results in enabling the compiler to automatically translate a user program that uses pointers to private data into its distributed implementation that provably protects private data throughout the computation. We empirically evaluate the constructions and report on the performance of representative programs.
Read moreCompiling Low Depth Circuits for Practical Secure Computation
With the rise of practical Secure Multi-party Computation (MPC) protocols, compilers have been developed that create Boolean or Arithmetic circuits for MPC from functionality descriptions in a high-level language. Previous compilers focused on the creation of size-minimal circuits. However, many MPC protocols, such as GMW and SPDZ, have a round complexity that is dependent on the circuit’s depth. When deploying these protocols in real world network settings, with network latencies in the range of tens or hundreds of milliseconds, the round complexity quickly becomes a significant performance bottleneck.
Read more