• Home
  • Search
  • Enhancing Performance Optimization of Multicore/Multichip Nodes with Data Structure Metrics
  • Cite Icon13
  • https://doi.org/10.1145/2588788Copy DOI Icon

Enhancing Performance Optimization of Multicore/Multichip Nodes with Data Structure Metrics

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Program performance optimization is usually based solely on measurements of execution behavior of code segments using hardware performance counters. However, memory access patterns are critical performance limiting factors for today's multicore chips where performance is highly memory bound. Therefore diagnoses and selection of optimizations based only on measurements of the execution behavior of code segments are incomplete because they do not incorporate knowledge of memory access patterns and behaviors. This article presents a low-overhead tool (MACPO) that captures memory traces and computes metrics for the memory access behavior of source-level (C, C++, Fortran) data structures. MACPO explicitly targets the measurement and metrics important to performance optimization for multicore chips. The article also presents a complete process for integrating measurement and analyses of code execution with measurements and analyses of memory access patterns and behaviors for performance optimization, specifically targeting multicore chips and multichip nodes of clusters. MACPO uses more realistic cache models for computation of latency metrics than those used by previous tools. Evaluation of the effectiveness of adding memory access behavior characteristics of data structures to performance optimization was done on subsets of the ASCI, NAS and Rodinia parallel benchmarks and two versions of one application program from a domain not represented in these benchmarks. Adding characteristics of the behavior of data structures enabled easier diagnoses of bottlenecks and more accurate selection of appropriate optimizations than with only code centric behavior measurements. The performance gains ranged from a few percent to 38 percent.

Similar Papers
  • Conference Article
  • Citations1

Trace-based Automatic Padding for Locality Improvement with Correlative Data Visualization Interface

  • Sep 15, 2007
  • M Hobbel +2
  • Research Article
  • Citations48

The NAS Parallel Benchmarks for evaluating C++ parallel programming frameworks on shared-memory architectures

  • Jul 19, 2021
  • Future Generation Computer Systems
  • Júnior Löff +6
  • Research Article

Dhcm: a dynamic hierarchy coordination mechanism for memory optimization

  • Apr 10, 2025
  • The Journal of Supercomputing
  • Juan Fang +2
  • Conference Article
  • Citations2

An FPGA implementation of a DWT with 5/3 filter using semi-programmable hardware

  • Nov 01, 2008
  • Akira Yamawaki +2
  • Conference Article
  • Citations17

Unstructured grid applications on GPU

  • Mar 05, 2011
  • Lizandro Solano-Quinde +3
  • Conference Article
  • Citations3

Split'n Trace NVM: Leveraging Library OSes for Semantic Memory Tracing

  • Aug 01, 2020
  • Christian Hakert +7
  • Research Article
  • Citations171

Optimizing database architecture for the new bottleneck: memory access

  • Dec 01, 2000
  • The VLDB Journal
  • Stefan Manegold +2
  • Research Article
  • Citations1

PatternS: An intelligent hybrid memory scheduler driven by page pattern recognition

  • May 16, 2024
  • Journal of Systems Architecture
  • Yanjie Zhen +5
  • Conference Article
  • Citations2

Experience report: Memory accesses for avionic applications and operating systems on a multi-core platform

  • Nov 01, 2015
  • Andreas Lofwenmark +1
  • Research Article

Methodology to define a static allocation mapping based on memory access patterns and the Signature of MPI applications in HPC systems

  • Oct 18, 2024
  • Journal of Computer Science and Technology
  • Gerard Enrique +5
  • Conference Article
  • Citations95

Optimizing Memory Efficiency for Deep Convolutional Neural Networks on GPUs

  • Nov 01, 2016
  • Chao Li +4
  • Conference Article
  • Citations5

Performance analysis of parallel hash join algorithms on a distributed shared memory machine implementation and evaluation on HP exemplar SPP 1600

  • Feb 23, 1998
  • M Nakano +2
  • Research Article
  • Citations3

Genetic algorithm based estimation of non–functional properties for GPGPU programs

  • Dec 07, 2019
  • Journal of Systems Architecture
  • Adrian Horga +3
  • Research Article
  • Citations2

Epoch Profiles: Microarchitecture-Based Application Analysis and Optimization

  • Jan 01, 2015
  • IEEE Computer Architecture Letters
  • Trevor E Carlson +3
  • Conference Article
  • Citations9

COMP: Compiler Optimizations for Manycore Processors

  • Dec 01, 2014
  • Linhai Song +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.