• Home
  • Search
  • SimHost: A Lightweight End-to-End Simulation Framework for HPC Network Systems
  • https://doi.org/10.1145/3767339Copy DOI Icon

SimHost: A Lightweight End-to-End Simulation Framework for HPC Network Systems

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

As high-performance computing (HPC) systems continue to grow in scale and complexity, simulation has emerged as a crucial tool for evaluating network performance, especially when real hardware is unavailable. However, achieving highly accurate results quickly and with low overhead in large-scale network evaluations remains a significant challenge. To tackle this issue, we propose SimHost, a lightweight end-to-end offline simulation framework featuring flexible communication and host models tailored for large-scale network simulations. SimHost enables the replay of application communication events without requiring an actual or virtualized execution environment or hardware. It accurately restores the message passing and system processing overheads of real application execution, delivering evaluation results consistent with physical deployments. With its cross-platform and discrete event design, SimHost integrates seamlessly with existing simulators, needing only minimal modifications for synchronous operation. Evaluations and use cases confirm that SimHost’s workloads accurately replicate real-world scenarios. Experimental results indicate that SimHost achieves less than 5.12% error in application completion time compared to the testbed, while maintaining only 7% simulation runtime overhead when scaled to 2048 nodes relative to synthetic workloads.

Similar Papers
  • Book Chapter

Code Modernization Tools for Assisting Users in Migrating to Future Generations of Supercomputers

  • Jan 01, 2017
  • Ritu Arora +1
  • Conference Article
  • Citations4

Simulation Framework for Studying Optical Cable Failures in Dragonfly Topologies

  • May 01, 2019
  • Tiffany A Connors +3
  • Conference Article
  • Citations1

Node level Power Profiling and Thermal Management in HPC system

  • Feb 01, 2016
  • Sherin M A +4
  • Research Article
  • Citations82

Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning

  • Apr 01, 2019
  • IEEE Transactions on Parallel and Distributed Systems
  • Ozan Tuncer +7
  • Research Article
  • Citations13

Cost-oriented proactive fault tolerance approach to high performance computing (HPC) in the cloud

  • Jan 22, 2014
  • International Journal of Parallel, Emergent and Distributed Systems
  • Ifeanyi P Egwutuoha +4
  • Conference Article
  • Citations4

Analysis and Prediction of Data Transfer Throughput for Data-Intensive Workloads

  • Dec 01, 2019
  • Devarshi Ghoshal +3
  • Conference Article

Evaluation of process level redundant checkpointing/restart for HPC systems

  • Nov 01, 2011
  • Ifeanyi P Egwutuoha +2
  • Conference Article
  • Citations5

Using Monitoring Data to Improve HPC Performance via Network-Data-Driven Allocation

  • Sep 20, 2021
  • Yijia Zhang +7
  • Research Article
  • Citations1

Python-based social science applications’ profiling and optimization on HPC systems using task and data parallelism

  • Sep 26, 2023
  • The Scientific Temper
  • S Prabagar +5
  • Conference Article
  • Citations21

First Experiences in Performance Benchmarking with the New SPEChpc 2021 Suites

  • May 01, 2022
  • Holger Brunst +9
  • Conference Article
  • Citations4

Energy-Efficient Workload Allocation in Distributed HPC System

  • Jul 01, 2019
  • Piotr Arabas +1
  • Research Article

ScaleQsim: Highly Scalable Quantum Circuit Simulation Framework for Exascale HPC Systems

  • Dec 01, 2025
  • Proceedings of the ACM on Measurement and Analysis of Computing Systems
  • Changjong Kim +7
  • Research Article
  • Citations29

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale

  • Sep 01, 2017
  • Supercomputing Frontiers and Innovations
  • Saurabh Hukerikar +1
  • Research Article

A Scalable Runtime Fault Localization Framework for High-Performance Computing Systems

  • Sep 30, 2017
  • International Journal of Parallel Programming
  • Jian Gao +3
  • Conference Article
  • Citations7

Adaptive memory power management techniques for HPC workloads

  • Dec 01, 2011
  • Karthik Elangovan +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.