• Home
  • Search
  • Accelerating Tensor Swapping in GPUs With Self-Tuning Compression
  • Cite Icon5
  • https://doi.org/10.1109/tpds.2022.3193867Copy DOI Icon

Accelerating Tensor Swapping in GPUs With Self-Tuning Compression

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Data swapping between CPUs and GPUs is widely used to address the GPU memory shortage issue when training deep neural networks (DNNs) requiring a larger amount of memory than that a GPU may have. Data swapping may become a bottleneck when its latency is longer than the latency of DNN computations. Tensor compression in GPUs can reduce the data swapping time. However, existing works on compressing tensors in the virtual memory of GPUs have three major issues: lack of portability because its implementation requires additional (de)compression units in memory controllers, sub-optimal compression performance for varying tensor compression ratios and sizes, and poor adaptation to dense tensors because they only focus on sparse tensors. We propose a self-tuning tensor compression framework, named <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CSwap+</small> , for improving the virtual memory management of GPUs. It uses GPUs for (de)compression directly and thus has high portability and is minimally dependent on GPU architecture features. Furthermore, it only applies compression on tensors that are deemed to be cost-effective considering their compression ratio, size, and the characteristics of compression algorithms at runtime. Finally, to adapt to DNN models with dense tensors, it also supports cost-effective lossy compression for dense tensors with nearly no model training accuracy degradation. We conduct the experiments through six representative memory-intensive DNN models. Compared to vDNN, <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CSwap+</small> reduces tensor swapping latency by up to 50.9% and 46.1% with NVIDIA V100 GPU, for DNN models with sparse and dense tensors, respectively.

Similar Papers
  • Research Article
  • Citations185

An Empirical Study of the Impact of Hyperparameter Tuning and Model Optimization on the Performance Properties of Deep Neural Networks

  • Apr 09, 2022
  • ACM Transactions on Software Engineering and Methodology
  • Lizhi Liao +3
  • Conference Article
  • Citations13

Distance-aware DNNs for robust speech recognition

  • Sep 06, 2015
  • Yajie Miao +1
  • Research Article
  • Citations20

Optimizing makespan and resource utilization for multi-DNN training in GPU cluster

  • Jun 24, 2021
  • Future Generation Computer Systems
  • Zhongjin Li +5
  • Research Article
  • Citations3

Developing a Deep Neural Network model for COVID-19 diagnosis based on CT scan images.

  • Jan 01, 2023
  • Mathematical biosciences and engineering : MBE
  • Javad Hassannataj Joloudari +7
  • Conference Article
  • Citations36

On decomposing a deep neural network into modules

  • Nov 07, 2020
  • Rangeet Pan +1
  • Research Article
  • Citations52

Deep neural network modeling of unknown partial differential equations in nodal space

  • Oct 15, 2021
  • Journal of Computational Physics
  • Zhen Chen +3
  • Research Article
  • Citations5

SSAT: Active Authorization Control and User’s Fingerprint Tracking Framework for DNN IP Protection

  • Oct 29, 2024
  • ACM Transactions on Multimedia Computing, Communications, and Applications
  • Mingfu Xue +5
  • PDF
  • Research Article
  • Citations7

Assessment of Therapeutic Responses Using a Deep Neural Network Based on 18F-FDG PET and Blood Inflammatory Markers in Pyogenic Vertebral Osteomyelitis

  • Nov 21, 2022
  • Medicina
  • Hyunkwang Shin +4
  • Research Article
  • Citations6

A Guessing Entropy-Based Framework for Deep Learning-Assisted Side-Channel Analysis

  • Jan 01, 2023
  • IEEE Transactions on Information Forensics and Security
  • Ziyue Zhang +2
  • PDF
  • Research Article
  • Citations10

An Adaptive Task Migration Scheduling Approach for Edge‐Cloud Collaborative Inference

  • Jan 01, 2022
  • Wireless Communications and Mobile Computing
  • Boyin Zhang +4
  • PDF
  • Research Article

Optimal deep neural network architecture design with improved generalization for data-driven cooling load estimation problem

  • May 02, 2025
  • Neural Computing and Applications
  • Baris Baykant Alagoz +4
  • Research Article
  • Citations33

Channel Characteristic-Based Deep Neural Network Models for Accurate Eye Diagram Estimation in High Bandwidth Memory (HBM) Silicon Interposer

  • Feb 01, 2022
  • IEEE Transactions on Electromagnetic Compatibility
  • Daehwan Lho +11
  • Research Article
  • Citations4

SoftMemoryBox II: A Scalable, Shared Memory Buffer Framework for Accelerating Distributed Training of Large-Scale Deep Neural Networks

  • Jan 01, 2020
  • IEEE Access
  • Shinyoung Ahn +1
  • Research Article

Industrial software technology: R Mitchell (ed) Peter Peregrinus Ltd (on behalf of the Institution of Electrical Engineers), London, UK (1987) 289pp £37.00

  • May 01, 1989
  • Information and Software Technology
  • R Veryard
  • Research Article

DeepClassifier: A Data Sampling-Based Hybrid BiLSTM-BiGRU Neural Network for Enhanced Type 2 Diabetes Prediction

  • Jan 01, 2026
  • Computer Modeling in Engineering & Sciences
  • Abdullahi Abubakar Imam +11
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.