• Home
  • Search
  • Epoch Profiles: Microarchitecture-Based Application Analysis and Optimization
  • Cite Icon2
  • https://doi.org/10.1109/lca.2014.2329873Copy DOI Icon

Epoch Profiles: Microarchitecture-Based Application Analysis and Optimization

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The performance of data-intensive applications, when running on modern multi- and many-core processors, is largely determined by their memory access behavior. Its most important contributors are the frequency and latency of off-chip accesses and the extent to which long-latency memory accesses can be overlapped with useful computation or with each other. In this paper we present two methods to better understand application and microarchitectural interactions. An <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">epoch profile</i> is an intuitive way to understand the relationships between three important characteristics: theon-chip cache size, the size of the reorder window of an out-of-order processor, and the frequency of processor stalls caused by long-latency,off-chip requests (epochs). By relating these three quantities one can more easily understand an application’s memory reference behavior and thus significantly reduce the design space. While epoch profiles help to provide insight into the behavior of a single application, developing an understanding of a number of applications in the presence of area and core count constraints presents additional challenges. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> Epoch-based microarchitectural analysis</i> is presented as a better way to understand the trade-offs for memory-bound applications in the presence of these physical constraints. Through epoch profiling and optimization, one can significantly reduce the multidimensional design space for hardware/software optimization through the use of high-level model-driven techniques.

Similar Papers
  • Conference Article
  • Citations39

Performance Analysis and Optimization of Sparse Matrix-Vector Multiplication on Modern Multi- and Many-Core Processors

  • Aug 01, 2017
  • Athena Elafrou +2
  • Research Article
  • Citations28

Can traditional programming bridge the ninja performance gap for parallel computing applications?

  • Apr 23, 2015
  • Communications of the ACM
  • Nadathur Satish +7
  • Research Article
  • Citations13

Enhancing Performance Optimization of Multicore/Multichip Nodes with Data Structure Metrics

  • May 01, 2014
  • ACM Transactions on Parallel Computing
  • Ashay Rane +1
  • Single Report

The High Performance and Wide Area Analysis and Mining of Scientific & Engineering Data

  • Dec 01, 2002
  • R Grossman
  • Supplementary Content

Many-Core Architectures: Hardware-Software Optimization and Modeling Techniques

  • May 19, 2015
  • AMS Dottorato Institutional Doctoral Theses Repository (University of Bologna)
  • Christian Pinto
  • Research Article
  • Citations3

Genetic algorithm based estimation of non–functional properties for GPGPU programs

  • Dec 07, 2019
  • Journal of Systems Architecture
  • Adrian Horga +3
  • Research Article
  • Citations4

Optimizing memory access traffic via runtime thread migration for on-chip distributed memory systems

  • Jun 24, 2014
  • The Journal of Supercomputing
  • Weiwei Fu +3
  • Conference Article
  • Citations31

STM: Cloning the spatial and temporal memory access behavior

  • Feb 01, 2014
  • Amro Awad +1
  • Conference Article
  • Citations9

Locality-aware bank partitioning for shared DRAM MPSoCs

  • Jan 01, 2017
  • Yangguo Liu +3
  • Research Article
  • Citations21

Locality-Aware Task Scheduling and Data Distribution for OpenMP Programs on NUMA Systems and Manycore Processors

  • Jan 01, 2015
  • Scientific Programming
  • Ananya Muddukrishna +2
  • Conference Article
  • Citations26

Many-Core Graph Workload Analysis

  • Nov 01, 2018
  • Stijn Eyerman +4
  • Conference Article
  • Citations1

A Memory Access Performance Detection and Optimization Driven by Address Mapping

  • May 20, 2022
  • Weijia Hua +1
  • Conference Article

Performance of Graph Analytics Applications on Many-Core Processors

  • Sep 01, 2018
  • Jenna Wise +3
  • Research Article
  • Citations43

Affinity-Based Thread and Data Mapping in Shared Memory Systems

  • Dec 05, 2016
  • ACM Computing Surveys
  • Matthias Diener +4
  • Dissertation

Efficient openMP over sequentially consistent distributed shared memory systems

  • Jul 20, 2011
  • Juan José Costa Prats
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.