• Home
  • Search
  • Multi-GPU System Design with Memory Networks
  • Cite Icon56
  • https://doi.org/10.1109/micro.2014.55Copy DOI Icon

Multi-GPU System Design with Memory Networks

  • Dec 1, 2014
  • Gwangsun Kim +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

GPUs are being widely used to accelerate different workloads and multi-GPU systems can provide higher performance with multiple discrete GPUs interconnected together. However, there are two main communication bottlenecks in multi-GPU systems -- accessing remote GPU memory and the communication between GPU and the host CPU. Recent advances in multi-GPU programming, including unified virtual addressing and unified memory from NVIDIA, has made programming simpler but the costly remote memory access still makes multi-GPU programming difficult. In order to overcome the communication limitations, we propose to leverage the memory network based on hybrid memory cubes (HMCs) to simplify multi-GPU memory management and improve programmability. In particular, we propose scalable kernel execution (SKE) where multiple GPUs are viewed as a single virtual GPU as a single kernel can be executed across multiple GPUs without modifying the source code. To fully enable the benefits of SKE, we explore alternative memory network designs in a multi-GPU system. We propose a GPU memory network (GMN) to simplify data sharing between the discrete GPUs while a CPU memory network (CMN) is used to simplify data communication between the host CPU and the discrete GPUs. These two types of networks can be combined to create a unified memory network (UMN) where the communication bottleneck in multi-GPU can be significantly minimized as both the CPU and GPU share the memory network. We evaluate alternative network designs and propose a sliced flattened butterfly topology for the memory network that scales better than previously proposed alternative topologies by removing local HMC channels. In addition, we propose an overlay network organization for unified memory network to minimize the latency for CPU access while providing high bandwidth for the GPUs. We evaluate trade-offs between the different memory network organization and show how UMN significantly reduces the communication bottleneck in multi-GPU systems.

Similar Papers
  • Book Chapter
  • Citations52

A Multi-GPU Programming Library for Real-Time Applications

  • Jan 01, 2012
  • Sebastian Schaetz +1
  • Research Article

Optimization of machine learning model training procedure on multi-gpu systems to enhance cyber security in telecommunication networks

  • Sep 05, 2024
  • Scientific Bulletin of UNFU
  • O P Kuziv +1
  • Research Article
  • Citations27

Efficient breadth first search on multi-GPU systems

  • Jun 06, 2013
  • Journal of Parallel and Distributed Computing
  • Enrico Mastrostefano +1
  • Research Article
  • Citations14

Scalable framework for mapping streaming applications onto multi-GPU systems

  • Feb 25, 2012
  • ACM SIGPLAN Notices
  • Huynh Phung Huynh +3
  • Research Article
  • Citations5

Fast Acquisition of Spread Spectrum Signals Using Multiple GPUs

  • Dec 01, 2019
  • IEEE Transactions on Aerospace and Electronic Systems
  • Ying Liu +2
  • Dissertation
  • Citations1

Enabling collaborative heterogeneous computing

  • Jan 01, 2020
  • Yifan Sun
  • Conference Article
  • Citations5

Converting data-parallelism to task-parallelism by rewrites: purely functional programs across multiple GPUs

  • Aug 30, 2015
  • Bo Joel Svensson +4
  • Research Article
  • Citations6

Effect of Alzheimer's Pathology on Task-Related Brain Network Reconfiguration in Aging.

  • Aug 21, 2023
  • The Journal of Neuroscience
  • Kaitlin E Cassady +7
  • Conference Article
  • Citations24

Coordinated Page Prefetch and Eviction for Memory Oversubscription Management in GPUs

  • May 01, 2020
  • Qi Yu +5
  • Research Article
  • Citations11

PET Metabolic Imaging of Time-Dependent Reorganization of Olfactory Cued Fear Memory Networks in Rats.

  • Oct 20, 2021
  • Cerebral Cortex
  • Anne-Marie Mouly +5
  • Book Chapter
  • Citations12

Hands on with OpenMP4.5 and Unified Memory: Developing Applications for IBM’s Hybrid CPU + GPU Systems (Part I)

  • Jan 01, 2017
  • Leopold Grinberg +2
  • Research Article

Scalable QR Factorisation of Ill-Conditioned Tall-and-Skinny Matrices on Distributed GPU Systems

  • Nov 11, 2025
  • Mathematics
  • Nenad Mijić +3
  • Conference Article
  • Citations5

Simulation of Information Propagation over Complex Networks: Performance Studies on Multi-GPU

  • Oct 01, 2013
  • Jiangming Jin +4
  • Conference Article
  • Citations2

Fine-Grained Air Quality Prediction using Attention Based Neural Network

  • Jul 01, 2018
  • Tianyu Liu +4
  • Book Chapter
  • Citations6

Gating Sensory Noise in a Spiking Subtractive LSTM

  • Jan 01, 2018
  • Isabella Pozzi +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.