• Open Access IconOpen Access
  • Cite Icon215
  • https://doi.org/10.1145/3458817.3476205Copy DOI Icon

ZeRO-infinity

  • Nov 13, 2021
  • Samyam Rajbhandari +4 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In the last three years, the largest dense deep learning models have grown over 1000x to reach hundreds of billions of parameters, while the GPU memory has only grown by 5x (16 GB to 80 GB). Therefore, the growth in model scale has been supported primarily though system innovations that allow large models to fit in the aggregate GPU memory of multiple GPUs. However, we are getting close to the GPU memory wall. It requires 800 NVIDIA V100 GPUs just to fit a trillion parameter model for training, and such clusters are simply out of reach for most data scientists. In addition, training models at that scale requires complex combinations of parallelism techniques that puts a big burden on the data scientists to refactor their model. In this paper we present ZeRO-Infinity, a novel heterogeneous system technology that leverages GPU, CPU, and NVMe memory to allow for unprecedented model scale on limited resources without requiring model code refactoring. At the same time it achieves excellent training throughput and scalability, unencumbered by the limited CPU or NVMe bandwidth. ZeRO-Infinity can fit models with tens and even hundreds of trillions of parameters for training on current generation GPU clusters. It can be used to fine-tune trillion parameter models on a single NVIDIA DGX-2 node, making large models more accessible. In terms of training throughput and scalability, it sustains over 25 petaflops on 512 NVIDIA V100 GPUs (40% of peak), while also demonstrating super linear scalability. An open source implementation of ZeRO-Infinity is available through DeepSpeed 1.

Similar Papers
  • Conference Article
  • Citations42

GPUswap

  • Mar 14, 2015
  • Jens Kehne +2
  • PDF
  • Research Article

Computing Quantiles of Functions of the Agent Distribution Using t-Digests

  • Sep 26, 2023
  • Computational Economics
  • Robert Kirkby
  • PDF
  • Research Article
  • Citations22

Efficient Use of GPU Memory for Large-Scale Deep Learning Model Training

  • Nov 04, 2021
  • Applied Sciences
  • Hyeonseong Choi +1
  • Conference Article
  • Citations13

Enabling energy-efficient DNN training on hybrid GPU-FPGA accelerators

  • Jun 03, 2021
  • Xin He +6
  • Conference Article
  • Citations5

Profiling Heterogeneous Computing Performance with VTune Profiler

  • Apr 27, 2021
  • Vladimir Tsymbal +1
  • PDF
  • Research Article
  • Citations3

Improving Global Performance on GPU for Algorithms with Main Loop Containing a Reduction Operation: Case of Dijkstra’s Algorithm

  • Jan 01, 2015
  • Journal of Computer and Communications
  • Amadou Chaibou +1
  • Conference Article
  • Citations210

VDNN: virtualized deep neural networks for scalable, memory-efficient neural network design

  • Oct 15, 2016
  • Minsoo Rhu +4
  • Research Article
  • Citations7

Portable performance on heterogeneous architectures

  • Mar 16, 2013
  • ACM SIGPLAN Notices
  • Phitchaya Mangpo Phothilimthana +3
  • Research Article
  • Citations14

Multiresolution Volume Filtering in the Tensor Compressed Domain.

  • Nov 09, 2017
  • IEEE Transactions on Visualization and Computer Graphics
  • Rafael Ballester-Ripoll +2
  • Conference Article
  • Citations33

Blocked All-Pairs Shortest Paths Algorithm for Hybrid CPU-GPU System

  • Sep 01, 2011
  • Kazuya Matsumoto +2
  • Research Article
  • Citations100

Efficient 3D Point Cloud Feature Learning for Large-Scale Place Recognition.

  • Jan 01, 2022
  • IEEE Transactions on Image Processing
  • Le Hui +4
  • Conference Article
  • Citations4

FreeLunch: Compression-based GPU Memory Management for Convolutional Neural Networks

  • Nov 01, 2021
  • Shaurya Patel +2
  • Conference Article
  • Citations29

Automatic Graph Partitioning for Very Large-scale Deep Learning

  • May 01, 2021
  • Masahiro Tanaka +3
  • Research Article
  • Citations5

Accelerating Tensor Swapping in GPUs With Self-Tuning Compression

  • Dec 01, 2022
  • IEEE Transactions on Parallel and Distributed Systems
  • Ping Chen +6
  • Research Article

SU‐E‐J‐60: Efficient Monte Carlo Dose Calculation On CPU‐GPU Heterogeneous Systems

  • May 29, 2014
  • Medical Physics
  • K Xiao +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.