• Home
  • Search
  • An Efficient GPU Cache Architecture for Applications with Irregular Memory Access Patterns
  • Cite Icon17
  • https://doi.org/10.1145/3322127Copy DOI Icon

An Efficient GPU Cache Architecture for Applications with Irregular Memory Access Patterns

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

GPUs provide high-bandwidth/low-latency on-chip shared memory and L1 cache to efficiently service a large number of concurrent memory requests. Specifically, concurrent memory requests accessing contiguous memory space are coalesced into warp-wide accesses. To support such large accesses to L1 cache with low latency, the size of L1 cache line is no smaller than that of warp-wide accesses. However, such L1 cache architecture cannot always be efficiently utilized when applications generate many memory requests with irregular access patterns especially due to branch and memory divergences that make requests uncoalesced and small. Furthermore, unlike L1 cache, the shared memory of GPUs is not often used in many applications, which essentially depends on programmers. In this article, we propose Elastic-Cache, which can efficiently support both fine- and coarse-grained L1 cache line management for applications with both regular and irregular memory access patterns to improve the L1 cache efficiency. Specifically, it can store 32- or 64-byte words in non-contiguous memory space to a single 128-byte cache line. Furthermore, it neither requires an extra memory structure nor reduces the capacity of L1 cache for tag storage, since it stores auxiliary tags for fine-grained L1 cache line managements in the shared memory space that is not fully used in many applications. To improve the bandwidth utilization of L1 cache with Elastic-Cache for fine-grained accesses, we further propose Elastic-Plus to issue 32-byte memory requests in parallel, which can reduce the processing latency of memory instructions and improve the throughput of GPUs. Our experiment result shows that Elastic-Cache improves the geometric-mean performance of applications with irregular memory access patterns by 104% without degrading the performance of applications with regular memory access patterns. Elastic-Plus outperforms Elastic-Cache and improves the performance of applications with irregular memory access patterns by 131%.

Similar Papers
  • Research Article
  • Citations16

MEMORY HIERARCHY PERFORMANCE PREDICTION FOR BLOCKED SPARSE ALGORITHMS

  • Sep 01, 1999
  • Parallel Processing Letters
  • Basilio B Fraguela +2
  • Research Article
  • Citations8

A Compile/Run-time Environment for the Automatic Transformation of Linked List Data Structures

  • Sep 24, 2008
  • International Journal of Parallel Programming
  • H L A Van Der Spek +3
  • Conference Article
  • Citations27

Adaptive Cache Bypass and Insertion for Many-core Accelerators

  • Jun 15, 2014
  • Xuhao Chen +6
  • Conference Article
  • Citations3

Boosting Memory Performance of Many-Core FPGA Device through Dynamic Precedence Graph

  • Apr 01, 2013
  • Yu Bai +4
  • Conference Article
  • Citations6

Managing shared memory spaces in an object-oriented real-time simulation

  • Aug 10, 1998
  • David Geyer +5
  • Conference Article
  • Citations10

A private level-1 cache architecture to exploit the latency and capacity tradeoffs in multicores operating at near-threshold voltages

  • Oct 01, 2013
  • Farrukh Hijaz +2
  • Conference Article
  • Citations7

HitME: low power Hit MEmory buffer for embedded systems

  • Jan 19, 2009
  • Andhi Janapsatya +2
  • Conference Article
  • Citations7

Software Prefetching for Unstructured Mesh Applications

  • Nov 01, 2018
  • Ioan Hadade +3
  • Conference Article
  • Citations17

Increasing GPU Translation Reach by Leveraging Under-Utilized On-Chip Resources

  • Oct 17, 2021
  • Jagadish B Kotra +3
  • Conference Article
  • Citations48

Accelerating Graph Analytics by Co-Optimizing Storage and Access on an FPGA-HMC Platform

  • Feb 15, 2018
  • Soroosh Khoram +3
  • Conference Article

Locality-Aware Dynamic Mapping for Multithreaded Applications

  • Feb 01, 2012
  • Betul Demiroz +3
  • Conference Article
  • Citations9

Demystifying memory access patterns of FPGA-based graph processing accelerators

  • Jun 20, 2021
  • Jonas Dann +2
  • Conference Article
  • Citations13

A Traversal Cache Framework for FPGA Acceleration of Pointer Data Structures: A Case Study on Barnes-Hut N-body Simulation

  • Dec 01, 2009
  • James Coole +2
  • Research Article
  • Citations13

Intel Xeon Phi acceleration of Hybrid Total FETI solver

  • May 10, 2017
  • Advances in Engineering Software
  • Michal Merta +6
  • Conference Article
  • Citations1

Planaria: Pattern Directed Cross-page Composite Prefetcher

  • Jun 23, 2024
  • Yuhang Liu +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.