• Home
  • Search
  • SoftMemoryBox II: A Scalable, Shared Memory Buffer Framework for Accelerating Distributed Training of Large-Scale Deep Neural Networks
  • Cite Icon4
  • https://doi.org/10.1109/access.2020.3038112Copy DOI Icon

SoftMemoryBox II: A Scalable, Shared Memory Buffer Framework for Accelerating Distributed Training of Large-Scale Deep Neural Networks

Show More
  • Abstract
  • Highlights & Summary
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Distributed processing using high-performance computing resources is essential for developers to train large-scale deep neural networks (DNNs). The major impediment to distributed DNN training is the communication bottleneck during the parameter exchange among the distributed DNN training workers. The communication bottleneck increases training time and decreases the utilization of the computational resources. Our previous study, SoftMemoryBox (SMB1) presented considerably superior performance compared to message passing interface (MPI) in the parameter communication of distributed DNN training. However, SMB1 had disadvantages such as the limited scalability of the distributed DNN training due to the restricted communication bandwidth from a single memory server, inability to provide a synchronization function for the shared memory buffer, and low portability/usability as a consequence of the kernel-level implementation. This paper proposes a scalable, shared memory buffer framework, called SoftMemoryBox II (SMB2), which overcomes the shortcomings of SMB1. With SMB2, distributed training processes can easily share virtually unified shared memory buffers composed of memory segments provided from remote memory servers and can exchange DNN parameters at high speed through the shared memory buffer. The scalable communication bandwidth of the SMB2 framework facilitates the reduction of DNN distributed training times compared to SMB1. According to intensive evaluation results, the communication bandwidth of the proposed SMB2 is 6.3 times greater than that of SMB1 when the SMB2 framework is scaled out to use eight memory servers. Moreover, the training time of SMB2-based asynchronous distributed training of five DNN models is up to 2.4 times faster than SMB1-based training.

Similar Papers
  • Research Article
  • Citations11

Effective Scheduler for Distributed DNN Training Based on MapReduce and GPU Cluster

  • Feb 22, 2021
  • Journal of Grid Computing
  • Jie Xu +5
  • Conference Article
  • Citations3

AccDP: Accelerated Data-Parallel Distributed DNN Training for Modern GPU-Based HPC Clusters

  • Dec 01, 2022
  • Nawras Alnaasan +4
  • Conference Article
  • Citations13

Distance-aware DNNs for robust speech recognition

  • Sep 06, 2015
  • Yajie Miao +1
  • Dissertation

Software-hardware co-design: towards ultimate efficiency in deep learning acceleration

  • Jan 01, 2024
  • Peiyan Dong
  • Conference Article
  • Citations36

On decomposing a deep neural network into modules

  • Nov 07, 2020
  • Rangeet Pan +1
  • Book Chapter
  • Citations1

Handwritten Digit Recognition Using Very Deep Convolutional Neural Network

  • Jan 01, 2022
  • M Dhilsath Fathima +2
  • PDF
  • Research Article
  • Citations13

Yielding Multi-Fold Training Strategy for Image Classification of Imbalanced Weeds

  • Apr 07, 2021
  • Applied Sciences
  • Vo Hoang Trong +3
  • Research Article
  • Citations3

Developing a Deep Neural Network model for COVID-19 diagnosis based on CT scan images.

  • Jan 01, 2023
  • Mathematical biosciences and engineering : MBE
  • Javad Hassannataj Joloudari +7
  • PDF
  • Research Article
  • Citations10

An Adaptive Task Migration Scheduling Approach for Edge‐Cloud Collaborative Inference

  • Jan 01, 2022
  • Wireless Communications and Mobile Computing
  • Boyin Zhang +4
  • Research Article
  • Citations6

A Guessing Entropy-Based Framework for Deep Learning-Assisted Side-Channel Analysis

  • Jan 01, 2023
  • IEEE Transactions on Information Forensics and Security
  • Ziyue Zhang +2
  • Research Article
  • Citations185

An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks

  • Apr 09, 2022
  • ACM Transactions on Software Engineering and Methodology
  • Lizhi Liao +3
  • Research Article
  • Citations52

Deep neural network modeling of unknown partial differential equations in nodal space

  • Oct 15, 2021
  • Journal of Computational Physics
  • Zhen Chen +3
  • Research Article
  • Citations5

SSAT: Active Authorization Control and User’s Fingerprint Tracking Framework for DNN IP Protection

  • Oct 29, 2024
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Mingfu Xue +5
  • Research Article

Data-Driven Modeling of a Direct Detection System With an Electro-Absorption Modulated Laser

  • Jun 15, 2025
  • Journal of Lightwave Technology
  • Alan Yi-Lun Yuan +1
  • Research Article
  • Citations6

Accelerating Distributed SGD With Group Hybrid Parallelism

  • Jan 01, 2021
  • IEEE Access
  • Kyung-No Joo +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.