• Home
  • Search
  • Locality aware memory assignment and tiling
  • Cite Icon5
  • https://doi.org/10.1145/3195970.3196070Copy DOI Icon

Locality aware memory assignment and tiling

  • Jun 24, 2018
  • Samuel Rogers +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

With the trend toward specialization, an efficient memory-path design is vital to capitalize customization in data-path. A monolithic memory hierarchy is often highly inefficient for irregular applications, traditionally targeted for CPUs. New approaches and tools are required to offer application-specific memory customization combining the benefits of cache and scratchpad memory simultaneously. This paper introduces a novel approach for automated application-specific on-chip memory assignment and tiling. The approach offers two major tools: (1) static memory access analysis and (2) variable-level memory assignment. Static memory analysis performs at the LLVM abstraction. It extracts target-independent pointer behaviors, measures the access strides and analyze the prefetchability of variables. (2) variable-level memory assignment creates a memory allocation graph for memory assignment (cache vs. scratchpad) based on the variables size and their estimated locality. It also explores the opportunity for tiling memory access. For the exploration and results, this paper uses Machsuite benchmarks (with both regular & irregular memory access behaviors), and gem5-Aladdin tool for performance & power evaluation. The proposed approach optimizes the memory hierarchy by automatically combining the benefits of cache, (tiled-) scratchpad at variable level granularity per individual applications. The results demonstrate more than 45% improvement in our power-stall product, on average, over the monolithic cache or scratchpad design.

Similar Papers
  • Conference Article
  • Citations2

Hardware-based fast exploration of cache hierarchies in application specific MPSoCs

  • Jan 01, 2014
  • Isuru Nawinne +3
  • Research Article
  • Citations2

DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor

  • Jun 30, 2025
  • ACM Transactions on Architecture and Code Optimization
  • Gia Bao Thieu +2
  • Research Article
  • Citations8

Compiler-Assisted Data Streaming for Regular Code Structures

  • Apr 28, 2020
  • IEEE Transactions on Computers
  • Nuno Neves +2
  • Conference Article

DVFS method of memory hierarchy based on CPU microarchitectural information

  • Oct 24, 2022
  • Bumgyu Park +6
  • Conference Article
  • Citations72

7.2 4Mb STT-MRAM-based cache with memory-access-aware power optimization and write-verify-write / read-modify-write scheme

  • Jan 01, 2016
  • Hiroki Noguchi +12
  • Conference Article
  • Citations38

Compiler support for selective page migration in NUMA architectures

  • Aug 24, 2014
  • Guilherme Piccoli +5
  • Book Chapter
  • Citations4

Improving Memory Access Performance of In-Memory Key-Value Store Using Data Prefetching Techniques

  • Jan 01, 2015
  • Pengfei Zhu +3
  • Conference Article
  • Citations39

Universal mechanisms for data-parallel architectures

  • Dec 03, 2003
  • Karthikeyan Sankaralingam +3
  • Research Article
  • Citations32

Adaptive Scratch Pad Memory Management for Dynamic Behavior of Multimedia Applications

  • Apr 01, 2009
  • IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
  • Doosan Cho +5
  • Research Article
  • Citations3

Specific read only data management for memory hierarchy optimization

  • Jan 22, 2015
  • ACM SIGBED Review
  • Gregory Vaumourin +3
  • Research Article
  • Citations43

Affinity-Based Thread and Data Mapping in Shared Memory Systems

  • Dec 05, 2016
  • ACM Computing Surveys
  • Matthias Diener +4
  • Conference Article
  • Citations16

Hybrid source-level simulation of data caches using abstract cache models

  • Mar 01, 2012
  • S Stattelmann +4
  • Conference Article
  • Citations7

HitME: low power Hit MEmory buffer for embedded systems

  • Jan 19, 2009
  • Andhi Janapsatya +2
  • Research Article
  • Citations31

Design and implementation of a pipelined 8 bit-serial single-flux-quantum microprocessorwith cache memories

  • Oct 18, 2007
  • Superconductor Science and Technology
  • M Tanaka +10
  • Research Article

Dhcm: a dynamic hierarchy coordination mechanism for memory optimization

  • Apr 10, 2025
  • The Journal of Supercomputing
  • Juan Fang +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.