• Home
  • Search
  • GPU Domain Specialization via Composable On-Package Architecture
  • Cite Icon10
  • https://doi.org/10.1145/3484505Copy DOI Icon

GPU Domain Specialization via Composable On-Package Architecture

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

As GPUs scale their low-precision matrix math throughput to boost deep learning (DL) performance, they upset the balance between math throughput and memory system capabilities. We demonstrate that a converged GPU design trying to address diverging architectural requirements between FP32 (or larger)-based HPC and FP16 (or smaller)-based DL workloads results in sub-optimal configurations for either of the application domains. We argue that a C omposable O n- PA ckage GPU (COPA-GPU) architecture to provide domain-specialized GPU products is the most practical solution to these diverging requirements. A COPA-GPU leverages multi-chip-module disaggregation to support maximal design reuse, along with memory system specialization per application domain. We show how a COPA-GPU enables DL-specialized products by modular augmentation of the baseline GPU architecture with up to 4× higher off-die bandwidth, 32× larger on-package cache, and 2.3× higher DRAM bandwidth and capacity, while conveniently supporting scaled-down HPC-oriented designs. This work explores the microarchitectural design necessary to enable composable GPUs and evaluates the benefits composability can provide to HPC, DL training, and DL inference. We show that when compared to a converged GPU design, a DL-optimized COPA-GPU featuring a combination of 16× larger cache capacity and 1.6× higher DRAM bandwidth scales per-GPU training and inference performance by 31% and 35%, respectively, and reduces the number of GPU instances by 50% in scale-out training scenarios.

Similar Papers
  • Conference Article
  • Citations6

Sectum: Accurate Latency Prediction for TEE-hosted Deep Learning Inference

  • Jul 01, 2022
  • Yan Li +3
  • Conference Article
  • Citations1

Simulated Annealing for Timeliness and Energy aware Deep Learning Job Assignment

  • Oct 01, 2019
  • Dong-Ki Kang +1
  • Conference Article

Safe Deep Reinforcement Learning Based on Sample Value Evaluation

  • Dec 02, 2022
  • Rongjun Ye +2
  • Research Article

Evaluasi Pelaksanaan Program Pelatihan Pembelajaran Mendalam (PM) bagi Guru SMPN 2 Praya Tengah Menggunakan Model Countenance Evaluation Stake

  • Feb 21, 2026
  • Jurnal Pendidikan, Sains, Geologi, dan Geofisika (GeoScienceEd Journal)
  • Nurul Hidayati +1
  • Conference Article
  • Citations29

SecDeep

  • May 18, 2021
  • Renju Liu +4
  • Research Article
  • Citations8

Development of training environment for deep learning with medical images on supercomputer system based on asynchronous parallel Bayesian optimization

  • Jan 20, 2020
  • The Journal of Supercomputing
  • Yukihiro Nomura +11
  • Research Article
  • Citations5

D3: Differential Testing of Distributed Deep Learning With Model Generation

  • Jan 01, 2025
  • IEEE Transactions on Software Engineering
  • Jiannan Wang +6
  • PDF
  • Research Article
  • Citations34

GaNDLF: the generally nuanced deep learning framework for scalable end-to-end clinical workflows

  • May 16, 2023
  • Communications Engineering
  • Sarthak Pati +41
  • Conference Article
  • Citations32

Offloaded Execution of Deep Learning Inference at Edge: Challenges and Insights

  • Mar 01, 2019
  • Swarnava Dey +2
  • Research Article
  • Citations43

Medical Image Classification Algorithm Based on Visual Attention Mechanism-MCNN.

  • Jan 01, 2021
  • Oxidative Medicine and Cellular Longevity
  • Fengping An +2
  • Conference Article
  • Citations1

Breast cancer classification using deep learning and FPGA inferencing

  • Jan 01, 2023
  • AIP conference proceedings
  • E-Hong Wong +4
  • Conference Article
  • Citations1

Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane

  • Aug 27, 2025
  • Yinchao Zhang +11
  • Research Article
  • Citations4

Exploring Bitslicing Architectures for Enabling FHE-Assisted Machine Learning

  • Nov 01, 2022
  • IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
  • Soumik Sinha +7
  • Conference Article
  • Citations4

Deep learning Achievements and Opportunities in Domain of Electronic Warfare Applications

  • Dec 05, 2021
  • Mohamed A Ammar +3
  • Research Article
  • Citations10

∇-Prox: Differentiable Proximal Algorithm Modeling for Large-Scale Optimization

  • Jul 26, 2023
  • ACM Transactions on Graphics
  • Zeqiang Lai +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.