- Research Article
1
- 10.1007/s11227-025-07972-7
Distributed fast and accurate simulation platform for advanced ARM- and RISC-V-based HPC systems
- Oct 21, 2025
- The Journal of Supercomputing
- Nikolaos Tampouratzis + 5 more +5
Publications from 2021 to 2026
Showing 4 of 4 papers
Distributed fast and accurate simulation platform for advanced ARM- and RISC-V-based HPC systems
Numerical Uncertainty in Analytical Pipelines Lead to Impactful Variability in Brain Networks
Abstract The analysis of brain-imaging data requires complex processing pipelines to support findings on brain function or pathologies. Recent work has shown that variability in analytical decisions, small amounts of noise, or computational environments can lead to substantial differences in the results, endangering the trust in conclusions 1-7 . We explored the instability of results by instrumenting a connectome estimation pipeline with Monte Carlo Arithmetic 8,9 to introduce random noise throughout. We evaluated the reliability of the connectomes, their features 10,11 , and the impact on analysis 12,13 . The stability of results was found to range from perfectly stable to highly unstable. This paper highlights the potential of leveraging induced variance in estimates of brain connectivity to reduce the bias in networks alongside increasing the robustness of their applications in the classification of individual differences. We demonstrate that stability evaluations are necessary for understanding error inherent to brain imaging experiments, and how numerical analysis can be applied to typical analytical workflows both in brain imaging and other domains of computational science. Overall, while the extreme variability in results due to analytical instabilities could severely hamper our understanding of brain organization, it also leads to an increase in the reliability of datasets.
Read moreCustom-Precision Mathematical Library Explorations for Code Profiling and Optimization
The typical processors used for scientific computing have fixed-width data-paths. This implies that mathematical libraries were specifically developed to target each of these fixed precisions (binary16, binary32, binary64). However, to address the increasing energy consumption and throughput requirements of scientific applications, library and hardware designers are moving beyond this one-size-fits-all approach. In this article we propose to study the effects and benefits of using user-defined floating-point formats and target accuracies in calculations involving mathematical functions. Our tool collects input-data profiles and iteratively explores lower precisions for each call-site of a mathematical function in user applications. This profiling data will be a valuable asset for specializing and fine-tuning mathematical function implementations for a given application. We demonstrate the tool's capabilities on SGP4, a satellite tracking application. The profile data shows the potential for specialization and provides insight into answering where it is useful to provide variable-precision designs for elementary function evaluation.
Read moreImplementing Breadth-First Search on a Compact Supercomputer Suiren
Cost and energy efficient supercomputers have received attention not only for scientific computation but for big data processing. In the fields of social networks and biology, the relationship between data is often represented by large target graphs that require huge computation costs to analyze. A new parallel BFS method called degree-chain traversal (DC) is proposed and implemented on the energy efficient compact supercomputer Suiren. In DC, by treating vertices that have the same parents as a form of ’chain’, both the communication amount and the number of memory accesses are reduced. Evaluation results show that the total amount of computation was reduced by 30%, and the execution time was shortened by 14%, when tasks are executed with four processes. We also tried to accelerate the execution with PEZY-SC, an MIMD accelerator attached to Suiren. However, the average execution time was not improved because of the large variation in the execution time depending on the root node. Through the analysis, an unbalanced task assignment and a bottleneck of the memory were pointed out. However, this bottleneck is eased by using new PEZY-SC2 which has wider memory bandwidth.
Read more