• Home
  • Search
  • Coherent Profiles: Enabling Efficient Reuse Distance Analysis of Multicore Scaling for Loop-based Parallel Programs
  • Open Access IconOpen Access
  • Cite Icon39
  • https://doi.org/10.1109/pact.2011.58Copy DOI Icon

Coherent Profiles: Enabling Efficient Reuse Distance Analysis of Multicore Scaling for Loop-based Parallel Programs

  • Oct 1, 2011
  • Meng-Ju Wu +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Reuse distance (RD) analysis is a powerful memory analysis tool that can potentially help architects study multicore processor scaling. One key obstacle though is multicore RD analysis requires measuring concurrent reuse distance (CRD) profiles across thread-interleaved memory reference streams. Sensitivity to memory interleaving makes CRD profiles architecture dependent, preventing them from analyzing different processor configurations. For loop-based parallel programs, CRD profiles shift coherently to larger CRD values with core count scaling because interleaving threads are symmetric. Simple techniques can predict such shifting, making the analysis of numerous multicore configurations from a small set of CRD profiles feasible. Given the ubiquity and scalability of loop-level parallelism, such techniques will be extremely valuable for studying future large multicore designs. This paper investigates using RD analysis to efficiently analyze multicore cache performance for loop-based parallel programs, making several contributions. First, we provide in depth analysis on how CRD profiles change with core count scaling. Second, we develop techniques to predict CRD profile scaling, in particular employing reference groups to predict coherent shift, and evaluate prediction accuracy. Third, we show core count scaling only degrades performance for last level caches (LLCs) below 16MB for our benchmarks and problem sizes, increasing to 64 -- 128MB if problem size scales by 64x. Finally, we apply CRD profiles to analyze multicore cache performance. When combined with existing problem scaling prediction, our techniques can predict LLC MPKI to within 11.1% of simulation across 1,728 configurations using only 36 measured CRD profiles.

Similar Papers
  • Research Article
  • Citations5

Efficient Cache Resizing policy for DRAM-based LLCs in ChipMultiprocessors

  • Sep 17, 2020
  • Journal of Systems Architecture
  • Bindu Agarwalla +2
  • Conference Article
  • Citations1

LLC Buffer for Arbitrary Data Sharing in Heterogeneous Systems

  • Dec 01, 2016
  • Yu Licheng +5
  • Conference Article
  • Citations8

GPUs Cache Performance Estimation using Reuse Distance Analysis

  • Oct 01, 2019
  • Yehia Arafa +5
  • Conference Article
  • Citations18

Parity++: Lightweight Error Correction for Last Level Caches

  • Jun 01, 2018
  • Irina Alam +3
  • Book Chapter
  • Citations4

A Machine Learning Approach for a Scalable, Energy-Efficient Utility-Based Cache Partitioning

  • Jan 01, 2015
  • Isa Ahmet Guney +5
  • Research Article
  • Citations9

RESTRAIN: A dynamic and cost-efficient resource management scheme for addressing performance interference in NFV-based systems

  • Jan 07, 2022
  • Journal of Network and Computer Applications
  • Venkatarami Reddy Chintapalli +3
  • Research Article
  • Citations2

Energy-Efficient Cache Partitioning Using Machine Learning for Embedded Systems

  • Jan 01, 2023
  • Jordan Journal of Electrical Engineering
  • Samar Nour +2
  • Conference Article
  • Citations4

Energy Efficient Last Level Caches via Last Read/Write Prediction

  • Oct 01, 2013
  • Marco A.Z Alves +3
  • Book Chapter
  • Citations3

A Survey of Low Power Design Techniques for Last Level Caches

  • Jan 01, 2018
  • Emmanuel Ofori-Attah +2
  • PDF
  • Preprint Article

Performance analysis of coalesce private or spill shared (CPOSS) replacement strategy overLRU to handle directory entry evictions from the slice last level cache (LLC) using a scalable coherent sparse directory in Multicore processor

  • May 06, 2024
  • Research Square
  • Narottam Sahu +3
  • Conference Article
  • Citations2

PV-aware Replacement Policy for Two-level Shared Cache

  • Dec 01, 2022
  • Bindu Agarwalla +2
  • Conference Article

A Novel Prefetch Technique for High Performance Embedded System

  • Oct 01, 2014
  • Hong Jun Choi +3
  • Conference Article
  • Citations9

Improving energy efficiency of embedded DRAM caches for high-end computing systems

  • Jun 23, 2014
  • Sparsh Mittal +2
  • Book Chapter

Towards Efficient Dynamic LLC Home Bank Mapping with NoC-Level Support

  • Jan 01, 2013
  • Mario Lodde +2
  • Conference Article
  • Citations6

Exploiting Secrets by Leveraging Dynamic Cache Partitioning of Last Level Cache

  • Feb 01, 2021
  • Anurag Agarwal +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.