• Home
  • Search
  • Highly Concurrent Latency-tolerant Register Files for GPUs
  • Cite Icon9
  • https://doi.org/10.1145/3419973Copy DOI Icon

Highly Concurrent Latency-tolerant Register Files for GPUs

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Graphics Processing Units (GPUs) employ large register files to accommodate all active threads and accelerate context switching. Unfortunately, register files are a scalability bottleneck for future GPUs due to long access latency, high power consumption, and large silicon area provisioning. Prior work proposes hierarchical register file to reduce the register file power consumption by caching registers in a smaller register file cache. Unfortunately, this approach does not improve register access latency due to the low hit rate in the register file cache. In this article, we propose the Latency-Tolerant Register File (LTRF) architecture to achieve low latency in a two-level hierarchical structure while keeping power consumption low. We observe that compile-time interval analysis enables us to divide GPU program execution into intervals with an accurate estimate of a warp’s aggregate register working-set within each interval. The key idea of LTRF is to prefetch the estimated register working-set from the main register file to the register file cache under software control, at the beginning of each interval, and overlap the prefetch latency with the execution of other warps. We observe that register bank conflicts while prefetching the registers could greatly reduce the effectiveness of LTRF. Therefore, we devise a compile-time register renumbering technique to reduce the likelihood of register bank conflicts. Our experimental results show that LTRF enables high-capacity yet long-latency main GPU register files, paving the way for various optimizations. As an example optimization, we implement the main register file with emerging high-density high-latency memory technologies, enabling 8× larger capacity and improving overall GPU performance by 34%.

Similar Papers
  • Research Article
  • Citations15

ContextPreRF: Enhancing the Performance and Energy of GPUs With Nonuniform Register Access

  • Jan 01, 2016
  • IEEE Transactions on Very Large Scale Integration (VLSI) Systems
  • Michael Moeng +3
  • Research Article

Power Estimation of Partitioned Register Files in a Clustered Architecture with Performance Evaluation

  • Mar 01, 2007
  • IEICE Transactions on Information and Systems
  • Y Sato +2
  • Research Article
  • Citations7

Exploring Soft-Error Robust and Energy-Efficient Register File in GPGPUs using Resistive Memory

  • Jan 28, 2016
  • ACM Transactions on Design Automation of Electronic Systems
  • Jingweijia Tan +3
  • Conference Article
  • Citations1

Duplicated Register File Design for Embedded Simultaneous Multithreading Microprocessor

  • Dec 01, 2005
  • Chengjie Zang +2
  • Conference Article
  • Citations4

Exploiting Zero Data to Reduce Register File and Execution Unit Dynamic Power Consumption in GPGPUs

  • Jul 01, 2020
  • Ahmad M Radaideh +1
  • Conference Article
  • Citations11

GPU Register Packing: Dynamically Exploiting Narrow-Width Operands to Improve Performance

  • Aug 01, 2017
  • Xin Wang +1
  • Research Article
  • Citations13

An Adaptive Thread Scheduling Mechanism With Low-Power Register File for Mobile GPUs

  • Jan 01, 2014
  • IEEE Transactions on Multimedia
  • Chih-Chieh Hsiao +2
  • Conference Article
  • Citations25

Architecting energy-efficient STT-RAM based register file on GPGPUs via delta compression

  • Jun 05, 2016
  • Hang Zhang +3
  • Conference Article
  • Citations28

Reducing power consumption of embedded processors through register file partitioning and compiler support

  • Jul 01, 2008
  • Xuan Guan +1
  • Research Article
  • Citations19

A low-level software-based fault tolerance approach to detect SEUs in GPUs' register files

  • Jul 15, 2017
  • Microelectronics Reliability
  • Marcio Gonçalves +3
  • Research Article

Pulsed-latch-based register file architecture for multiport

  • Mar 30, 2025
  • International Journal of Science and Research Archive
  • Fazal Shah +4
  • Research Article
  • Citations115

Multiple-banked register file architectures

  • May 01, 2000
  • ACM SIGARCH Computer Architecture News
  • José-Lorenzo Cruz +3
  • Conference Article
  • Citations2

Support of Paged Register Files for Improving Context Switching on Embedded Processors

  • Jan 01, 2009
  • Chung-Wen Huang +3
  • Dissertation

Improving multithreading performance for clustered VLIW architectures.

  • Jun 14, 2013
  • Manoj Gupta
  • Conference Article
  • Citations13

Bank stealing for conflict mitigation in GPGPU Register File

  • Jul 01, 2015
  • Naifeng Jing +5
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.