- Research Article
- 10.34121/1028-9763-2026-1-100-108
Reliability assessment of cluster computing systems.
- Jan 01, 2026
- Mathematical machines and systems
- N.v Cespedes Garcia
The article is devoted to the study of international approaches to assessing the reliability of cluster computing systems and their combination with modern domestic methodologies. Cluster systems are widely used in fundamental and applied research, numerical modeling of complex physical processes, climate forecasting, analysis of large volumes of data, the aerospace and automotive industries, energy, cloud computing, and other areas. Therefore, assessing the relia-bility of cluster computing systems is critically important for ensuring their continuous opera-tion and resistance to failures under conditions of intensive load. A comprehensive assessment of the reliability of cluster systems consists of the following stages: analysis of system structure components, calculation of reliability metrics, modeling of failure scenarios, analysis of com-bined failure probabilities, and assessment of redundancy. It has been determined that, in gen-eral, cluster systems have a so-called «k» with «n» structure, therefore, reliability metrics are calculated according to the corresponding domestic methodologies. In the general case, failures of cluster computing systems are non-monotonic in nature. Consequently, to calculate the prob-ability of a failure-free operation, modern domestic methodologies based on the DN-distribution of failures can be used. Assessment of the reliability of cluster computing systems is a multi-level approach that combines analytical, experimental, and model research methods. Such a comprehensive approach allows the identification of potential bottlenecks in the system architecture, identify critical components, and assess the risks of failures. This enables timely implementation of effective mechanisms for fault tolerance, redundancy, and automatic recov-ery.
Read more