• Home
  • Search
  • Node level Power Profiling and Thermal Management in HPC system
  • Cite Icon1
  • https://doi.org/10.1109/icghpc.2016.7508064Copy DOI Icon

Node level Power Profiling and Thermal Management in HPC system

  • Feb 1, 2016
  • Sherin M A +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In addition to the performance, power consumption has become a major concern in High Performance Computing (HPC) systems. Typically the cooling system and the IT loads are the major contributors to the power bills. Understanding the power consumption at different granular levels in the HPC system is a first step to quantify the problems in HPC system. By proper monitoring and effective utilization of the cooling system, the power requirements of the HPC facility can be effectively met. In this paper we present a system where node level power measurement and WSN based rack level temperature measurement are used to provide localized control of cold air supply. A Smart Power Monitoring and Distribution Unit (PMDU) is designed and developed, to replace the existing Power Distribution Unit (PDU) in HPC. This can measure and report the power consumption, to support power profiling of large scale HPC system. This measured data are communicated to a base station via Ethernet. This base station collects all such measurements which can be used for power profiling of IT load of the HPC system. This helps to provide better insight into the power utilization pattern. Wireless Sensor Network (WSN) is used to collect exhaust and inlet air temperature of the server nodes and this information is used for directing the cold air effectively. A vent control system is designed and fabricated for intelligently directing the air flow to the server node inlet. It takes the node power and the temperature data as its inputs. This enables supplying/ redirecting more cold air towards under-cooled nodes without creating an extra load on the cooling system, thereby bringing in effective cooling at reduced power consumption. As an added advantage this could help in hot-spots mitigation.

Similar Papers
  • Book Chapter

Code Modernization Tools for Assisting Users in Migrating to Future Generations of Supercomputers

  • Jan 01, 2017
  • Ritu Arora +1
  • Conference Article
  • Citations4

Simulation Framework for Studying Optical Cable Failures in Dragonfly Topologies

  • May 01, 2019
  • Tiffany A Connors +3
  • Research Article
  • Citations82

Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning

  • Apr 01, 2019
  • IEEE Transactions on Parallel and Distributed Systems
  • Ozan Tuncer +7
  • Research Article
  • Citations13

Cost-oriented proactive fault tolerance approach to high performance computing (HPC) in the cloud

  • Jan 22, 2014
  • International Journal of Parallel, Emergent and Distributed Systems
  • Ifeanyi P Egwutuoha +4
  • Conference Article
  • Citations5

Using Monitoring Data to Improve HPC Performance via Network-Data-Driven Allocation

  • Sep 20, 2021
  • Yijia Zhang +7
  • Conference Article
  • Citations4

Analysis and Prediction of Data Transfer Throughput for Data-Intensive Workloads

  • Dec 01, 2019
  • Devarshi Ghoshal +3
  • Conference Article

Evaluation of process level redundant checkpointing/restart for HPC systems

  • Nov 01, 2011
  • Ifeanyi P Egwutuoha +2
  • Supplementary Content

Power-Aware Job Dispatching in High Performance Computing Systems

  • May 15, 2017
  • AMS Dottorato Institutional Doctoral Theses Repository (University of Bologna)
  • Andrea Borghesi
  • Research Article
  • Citations1

Python-based social science applications’ profiling and optimization on HPC systems using task and data parallelism

  • Sep 26, 2023
  • The Scientific Temper
  • S Prabagar +5
  • Conference Article
  • Citations21

First Experiences in Performance Benchmarking with the New SPEChpc 2021 Suites

  • May 01, 2022
  • Holger Brunst +9
  • Conference Article
  • Citations4

Energy-Efficient Workload Allocation in Distributed HPC System

  • Jul 01, 2019
  • Piotr Arabas +1
  • Research Article
  • Citations29

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale

  • Sep 01, 2017
  • Supercomputing Frontiers and Innovations
  • Saurabh Hukerikar +1
  • Research Article

A Scalable Runtime Fault Localization Framework for High-Performance Computing Systems

  • Sep 30, 2017
  • International Journal of Parallel Programming
  • Jian Gao +3
  • Conference Article
  • Citations10

Spark Meets MPI: Towards High-Performance Communication Framework for Spark using MPI

  • Sep 01, 2022
  • Kinan Al-Attar +4
  • Conference Article
  • Citations68

Entropy-Aware I/O Pipelining for Large-Scale Deep Learning on HPC Systems

  • Sep 01, 2018
  • Yue Zhu +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.