• Home
  • Search
  • A Hierarchical Tridiagonal System Solver for Heterogenous Supercomputers
  • Cite Icon6
  • https://doi.org/10.1109/scala.2014.12Copy DOI Icon

A Hierarchical Tridiagonal System Solver for Heterogenous Supercomputers

  • Nov 1, 2014
  • Xinliang Wang +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Tridiagonal system solver is an important kernel in many scientific and engineering applications. Even though quite a few parallel algorithms and implementations have been addressed in recent years, challenges still remain when solving large-scale tridiagonal system on heterogenous supercomputers. In this paper, a hierarchical algorithm framework SPIKE (pronounced 'SPIKE squared') is proposed to minimize the parallel overhead and to achieve the best utilization of CPU-GPU hybrid systems. In these systems, a layered and adaptive partitioning is presented based on the SPIKE algorithm to effectively control the sequential parts while efficiently exploiting the computation and communication overlapping in heterogeneous computing node. Moreover, the SPIKE algorithm is reformulated to reduce the matrix computations to only 1/3 in our hierarchical algorithm framework. Meanwhile, an improved implementation of the tiled-PCR-pThomas algorithm is employed for the GPU architecture, and the shared memory usage on the GPU can be reduced by 1/3 using careful dependence analysis on solving unit vector tridiagonal systems. Our experiments on Tianhe-1A show ideal weak scalability on up to 128 nodes when solving a tridiagonal system with a size of 1920M in the largest run and good strong scalability (70%) from 32 nodes to 256 nodes when solving a tridiagonal system with a size of 480M. Furthermore, the adaptive task partition across the CPU and GPU can get over 10% performance improvement in the strong scaling test with 256 nodes. In one computing node of Tianhe-1A, our GPU-only code can outperform the CUSPARSE version (non-pivoting tridiagonal solver) by 30%, and our hybrid code is about 6.7 times faster than the Intel SPIKE multi-process version for tridiagonal systems having a size of 3M, 5M, and 15M.

Similar Papers
  • Research Article
  • Citations5

Parallel prefix operations on GPU: tridiagonal system solvers and scan operators

  • Nov 02, 2018
  • The Journal of Supercomputing
  • Adrián P Diéguez +2
  • Research Article
  • Citations4

Performance comparison of a set of periodic and non-periodic tridiagonal solvers on SP2 and Paragon parallel computers

  • Aug 01, 1997
  • Concurrency: Practice and Experience
  • Xian-He Sun +1
  • Conference Article
  • Citations1

A parallel algorithm to solve symmetric tridiagonal linear systems

  • Jan 01, 2010
  • Yan Zhong +2
  • Conference Article
  • Citations18

Parallel algorithms for solution of tridiagonal systems on multicomputers

  • Jan 01, 1989
  • Xian-He Sun +2
  • Conference Article
  • Citations4

A Stable Parallel Algorithm for Diagonally Dominant Tridiagonal Linear Systems

  • Dec 01, 2015
  • S Chandra Sekhara Rao +1
  • Research Article
  • Citations2

Graph Transformation and Designing Parallel Sparse Matrix Algorithms beyond Data Dependence Analysis

  • Jan 01, 2004
  • Scientific Programming
  • H.X Lin
  • Research Article
  • Citations9

A class of novel parallel algorithms for the solution of tridiagonal systems

  • Apr 29, 2005
  • Parallel Computing
  • J Verkaik +1
  • PDF
  • Research Article
  • Citations5

Parallel Thomas approach development for solving tridiagonal systems in GPU programming − steady and unsteady flow simulation

  • Jan 01, 2020
  • Mechanics & Industry
  • Milad Souri +2
  • Book Chapter

A new Parallel algorithm for solving general linear systems of equations

  • Jan 01, 1986
  • Liao Qui-Wei
  • Conference Article
  • Citations25

A divide-and-conquer method of solving tridiagonal systems on hypercube massively parallel computers

  • Dec 02, 1991
  • Xiaojing Wang +1
  • Conference Article
  • Citations3

A Study on Network Sharing and Radio Resource Management in 3G and Beyond Mobiles Wireless Networks Supporting Heterogeneous Traffic

  • Oct 16, 2006
  • S.A Alqahtani +3
  • Research Article
  • Citations13

The parallel recursive decoupling algorithm for solving tridiagonal linear systems

  • May 01, 1993
  • Parallel Computing
  • G Spaletta +1
  • Conference Article
  • Citations219

Fast tridiagonal solvers on the GPU

  • Jan 09, 2010
  • Yao Zhang +2
  • Book Chapter
  • Citations8

A recursive doubling algorithm for solution of tridiagonal systems on hypercube multiprocessors

  • Jan 01, 1990
  • Advances in Parallel Computing
  • Ömer Eğecioǧlu +2
  • Conference Article
  • Citations60

Adaptive Partitioning for Large-Scale Dynamic Graphs

  • Jun 01, 2014
  • Luis M Vaquero +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.