Modeling and Verification of Distributed Algorithms in Theorem Proving Environments
The field of application for distributed algorithms is growing with the ongoing demand for new network applications. Wherever a global goal must be achieved by a number of peers, a distributed algorithm consisting of local computations and communication protocols must be developed. In asynchronous environments the design of a distributed algorithm must respect the fact that processes may vary in speed and that messages might be delayed for an arbitrarily long time. Moreover, it is often required that the algorithm still works in the presence of crash failures, i.e., in scenarios, where some processes stop working at some time. The complex nature of concurrent executions in distributed systems often hinder the application of common methods for modeling and verification. Although many different formal approaches for modeling fault tolerant distributed algorithms, like e.g. I/O automata and process calculi have been proposed, many works in the field of distributed systems use informal pseudo-code representations to introduce new algorithms. On most cases this leads to informal and crude arguments for correctness, which usually become incomprehensible if all race conditions must be taken into account. This work proposes new formal strategies to model and verify distributed algorithms, whereas we focus on two main goals. Firstly, we introduce a means to satisfy a designer’s demands by using methods that are easy to understand and to work with and which are very similar to common applied methods as e.g. TLA and Abstract State Machines. Secondly, our strategies are suitable for the formal use in a theorem proving environment. This enables mechanical verification of our algorithms and therefore, we are able to produce machinechecked proofs for correctness. Our model adopts abstractions from the work of Fuzzati, which examines two similar distributed algorithms in a logical framework. We lifted these abstractions to a more general level to make them suitable and reusable for as many distributed algorithms as possible. Furthermore, we present a new method to place local computation steps in the respective global context. The advantage of our method is that, on the on hand, like with pseudo code, computations can be described as local steps, but, on the other hand, it is possible to formally reason over global system states. Our model for communication mechanisms comprises modules for different kinds of message passing, broadcasts, and shared memory access. We are convinced that our model can easily be extended to support almost all kinds of communication infrastructures that can formally be described. We show how properties that have to be verified can be expressed in terms of our model to enable formal reasoning. Furthermore, we present appropriate strategies to verify different classes of properties. For the theorem proving environment Isabelle/HOL we provide a library which incorporates all introduced abstractions and can, therefore, be reused and extended for further applications, new algorithms, and further verification projects. Finally, we present case studies for the verification of different distributed algorithms to demonstrate the applicability of our strategies. Four known algorithms are verified in our model using our framework for Isabelle/HOL. The algorithms vary in complexity, the used communication infrastructure, the failure model, and other constraints, so that many different modeling and verification aspects are exemplified.
Read more