• Home
  • Search
  • FgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms
  • Cite Icon14
  • https://doi.org/10.1145/3512770Copy DOI Icon

FgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Sparse matrix-sparse vector (SpMSpV) multiplication is one of the fundamental and important operations in many high-performance scientific and engineering applications. The inherent irregularity and poor data locality lead to two main challenges to scaling SpMSpV over high-performance computing (HPC) systems: (i) a large amount of redundant data limits the utilization of bandwidth and parallel resources; (ii) the irregular access pattern limits the exploitation of computing resources. This paper proposes a fine-grained parallel SpMSpV (fgSpMSpV) framework on Sunway TaihuLight supercomputer to alleviate the challenges for large-scale real-world applications. First,fgSpMSpVadopts an MPI\( + \)OpenMP\( +X \)parallelization model to exploit the multi-stage and hybrid parallelism of heterogeneous HPC architectures and accelerate both pre-/post-processing and main SpMSpV computation. Second,fgSpMSpVutilizes an adaptive parallel execution to reduce the pre-processing, adapt to the parallelism and memory hierarchy of the Sunway system, while still tame redundant and random memory accesses in SpMSpV, including a set of techniques like the fine-grained partitioner, re-collection method, and Compressed Sparse Column Vector (CSCV) matrix format. Third,fgSpMSpVuses several optimization techniques to further utilize the computing resources.fgSpMSpVon the Sunway TaihuLight gains a noticeable performance improvement from the key optimization techniques with various sparsity of the input. Additionally,fgSpMSpVis implemented on an NVIDIA Tesal P100 GPU and applied to the breath-first-search (BFS) application.fgSpMSpVon a P100 GPU obtains the speedup of up to\( 134.38\times \)over the state-of-the-art SpMSpV algorithms, and the BFS application usingfgSpMSpVachieves the speedup of up to\( 21.68\times \)over the state-of-the-arts.

Similar Papers
  • Book Chapter

Code Modernization Tools for Assisting Users in Migrating to Future Generations of Supercomputers

  • Jan 01, 2017
  • Ritu Arora +1
  • Conference Article
  • Citations22

Software-only based Diverse Redundancy for ASIL-D Automotive Applications on Embedded HPC Platforms

  • Jan 01, 2020
  • Sergi Alcaide +3
  • Conference Article
  • Citations4

Parallel beam dynamics calculations on high performance computers

  • Jan 01, 1997
  • AIP conference proceedings
  • Robert Ryne +1
  • Book Chapter

Cooperative Preprocessing at Petabytes on High Performance Computing System

  • Jan 01, 2018
  • Rujun Sun +2
  • Conference Article
  • Citations13

Experimental Verification and Analysis of Dynamic Loop Scheduling in Scientific Applications

  • Jun 01, 2018
  • Ali Mohammed +4
  • Conference Article
  • Citations4

Simulation Framework for Studying Optical Cable Failures in Dragonfly Topologies

  • May 01, 2019
  • Tiffany A Connors +3
  • Research Article
  • Citations11

Heterogeneous parallel algorithm design and performance optimization for WENO on the Sunway Taihulight supercomputer

  • Aug 13, 2019
  • Tsinghua Science and Technology
  • Jianqiang Huang +3
  • Conference Article
  • Citations1

Node level Power Profiling and Thermal Management in HPC system

  • Feb 01, 2016
  • Sherin M A +4
  • Conference Article
  • Citations4

Analysis and Prediction of Data Transfer Throughput for Data-Intensive Workloads

  • Dec 01, 2019
  • Devarshi Ghoshal +3
  • Conference Article
  • Citations23

SciChain: Blockchain-enabled Lightweight and Efficient Data Provenance for Reproducible Scientific Computing

  • Apr 01, 2021
  • Abdullah Al-Mamun +2
  • Research Article
  • Citations13

Cost-oriented proactive fault tolerance approach to high performance computing (HPC) in the cloud

  • Jan 22, 2014
  • International Journal of Parallel, Emergent and Distributed Systems
  • Ifeanyi P Egwutuoha +4
  • Conference Article
  • Citations7

ICE: A General and Validated Energy Complexity Model for Multithreaded Algorithms

  • Dec 01, 2016
  • Vi Ngoc-Nha Tran +1
  • Research Article
  • Citations82

Online Diagnosis of Performance Variation in HPC Systems Using Machine Learning

  • Apr 01, 2019
  • IEEE Transactions on Parallel and Distributed Systems
  • Ozan Tuncer +7
  • Conference Article

Evaluation of process level redundant checkpointing/restart for HPC systems

  • Nov 01, 2011
  • Ifeanyi P Egwutuoha +2
  • Conference Article
  • Citations2

Parallel Multi-hop Routing Protocol for 5G Backhauling Network Using HPC Platform

  • Dec 15, 2020
  • Somaya A Aboulrous +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.