• Home
  • Search
  • MoDNN: Memory Optimal Deep Neural Network Training on Graphics Processing Units
  • Cite Icon21
  • https://doi.org/10.1109/tpds.2018.2866582Copy DOI Icon

MoDNN: Memory Optimal Deep Neural Network Training on Graphics Processing Units

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Graphics processing units (GPUs) have been widely adopted to accelerate the training of deep neural networks (DNNs). Although the computational performance of GPUs has been improving steadily, the memory size of modern GPUs is still quite limited, which restricts the sizes of DNNs that can be trained on GPUs, and hence raises serious challenges. This paper introduces a framework, referred to as moDNN (memory optimal DNN training on GPUs), to optimize the memory usage in DNN training. moDNN supports automatic tuning of DNN training code to match any given memory budget (not smaller than the theoretical lower bound). By taking full advantage of overlapping computations and data transfers, we develop new heuristics to judiciously schedule data offloading and prefetching transfers, together with convolution algorithm selection, to optimize memory usage. We further devise a new sub-batch size selection method which also greatly reduces memory usage. moDNN can save memory usage up to 59×, compared with an ideal case which assumes that the GPU memory is sufficient to hold all data. When executing moDNN on a GPU with 12 GB memory, the training time is increased by only 3 percent, which is much shorter than that incurred by the best known approach, vDNN. Furthermore, we propose an optimization strategy for moDNN on multiple GPUs again by utilizing the idea of overlapping data transfers and GPU computations. The results show that 3.7× speedup is attained on four GPUs.

Similar Papers
  • Conference Article
  • Citations3

AccDP: Accelerated Data-Parallel Distributed DNN Training for Modern GPU-Based HPC Clusters

  • Dec 01, 2022
  • Nawras Alnaasan +4
  • Conference Article
  • Citations3

Victream

  • Dec 05, 2017
  • Jun Suzuki +6
  • Research Article
  • Citations2

Gpubased, microsecond latency, hectochannel mimo feedback control of magnetically confined plasmas

  • Jan 01, 2013
  • Columbia Academic Commons (Columbia University)
  • M E Mauel +1
  • Research Article
  • Citations6

A Guessing Entropy-Based Framework for Deep Learning-Assisted Side-Channel Analysis

  • Jan 01, 2023
  • IEEE Transactions on Information Forensics and Security
  • Ziyue Zhang +2
  • Book Chapter
  • Citations3

Utilizing GPU Virtualization to Protect the Private Keys of GPU Cryptographic Computation

  • Jan 01, 2018
  • Ziyang Wang +3
  • Conference Article
  • Citations3

Efficient GPU-Based Query Processing with Pruned List Caching in Search Engines

  • Dec 01, 2017
  • Dongdong Wang +6
  • Book Chapter
  • Citations1

Distributed Training of Deep Neural Network for Segmentation-Free Telugu Word Recognition

  • Jan 01, 2021
  • Koteswara Rao Devarapalli +1
  • Conference Article
  • Citations3

FlexGPU: A Flexible and Efficient Scheduler for GPU Sharing Systems

  • May 01, 2020
  • Qichen Chen +3
  • Conference Article
  • Citations1

Automatic Parallelization of GPU Applications Using OpenCL

  • Jul 01, 2015
  • Lizandro D Solano-Quinde +2
  • Research Article
  • Citations6

Analysis of heterogeneous computing approaches to simulating heat transfer in heterogeneous material

  • Jun 17, 2019
  • Journal of Parallel and Distributed Computing
  • Andrew Loeb +1
  • Research Article
  • Citations3

A Novel Approach for Efficient Training of Deep Neural Networks

  • Sep 01, 2018
  • Indonesian Journal of Electrical Engineering and Computer Science
  • D.T.V Dharmajee Rao +1
  • Conference Article
  • Citations13

GLoP

  • Sep 09, 2014
  • Xavier J A Bellekens +4
  • Research Article
  • Citations248

Survey on Deep Neural Networks in Speech and Vision Systems

  • Jul 26, 2020
  • Neurocomputing
  • M Alam +4
  • Conference Article
  • Citations53

Improving the Performance of CA-GMRES on Multicores with Multiple GPUs

  • May 01, 2014
  • Ichitaro Yamazaki +4
  • Conference Article
  • Citations19

Deep Neural Network Training Accelerator Designs in ASIC and FPGA

  • Oct 21, 2020
  • Shreyas K Venkataramanaiah +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.