• Home
  • Search
  • Combined partitioning and data padding for scheduling multiple loop nests
  • Cite Icon10
  • https://doi.org/10.1145/502217.502228Copy DOI Icon

Combined partitioning and data padding for scheduling multiple loop nests

  • Jan 1, 2001
  • Zhong Wang +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

With the widening performance gap between processors and main memory, efficient memory accessing behavior is necessary for good program performance. Loop partition is an effective way to exploit the data locality. Traditional loop partition techniques, however, consider only a singleton nested loop. This paper presents multiple loop partition scheduling technique, which combines the loop partition and data padding to generate the detailed partition schedule. The computation and data prefetching are balanced in the partition schedule, such that the long memory latency can be hidden efficiently. Multiple loop partition scheduling explores parallelism among computations, and exploit the data locality between different loop nests as well in each loop nest. Data padding is applied in our technique to eliminate the cache interference, which overcomes the problem of cache conflict misses arisen from loop partition. Therefore, our technique can be applied in architectures with low associativity cache. The experiments show that multiple loop partition scheduling can achieve the significant improvement over the existing methods

Similar Papers
  • Research Article
  • Citations15

Partition Scheduling on Heterogeneous Multicore Processors for Multi-dimensional Loops Applications

  • Jul 15, 2016
  • International Journal of Parallel Programming
  • Yan Wang +2
  • Research Article

Minimizing write operation for multi-dimensional DSP applications via a two-level partition technique with complete memory latency hiding

  • Feb 01, 2015
  • Journal of Systems Architecture
  • Yan Wang +2
  • Book Chapter
  • Citations49

Iteration Space Slicing for Locality

  • Jan 01, 2000
  • William Pugh +1
  • Research Article
  • Citations1

Enhancing On-Device DNN Inference Performance With a Reduced Retention-Time MRAM-Based Memory Architecture

  • Jan 01, 2024
  • IEEE Access
  • Munhyung Lee +3
  • Conference Article
  • Citations2

Architecture-Aware Mapping and Scheduling of IMA partitions on Multicore platforms

  • Oct 10, 2018
  • Aishwarya Vasu +1
  • Conference Article
  • Citations59

Multi-dimensional incremental loop fusion for data locality

  • Jun 24, 2003
  • S Verdoolaege +3
  • Book Chapter
  • Citations8

MOSS-DB: A Hardware-Aware OLAP Database

  • Jan 01, 2010
  • Yansong Zhang +2
  • Conference Article
  • Citations6

Memory access micro-profiling for ASIP design

  • Jan 01, 2006
  • K Karuri +4
  • Conference Article

A multiple criteria-based spectral partitioning method for remotely sensed hyperspectral image classification

  • Oct 24, 2016
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Yi Liu +3
  • Conference Article
  • Citations236

Dynamic hot data stream prefetching for general-purpose programs

  • May 17, 2002
  • Trishul M Chilimbi +1
  • Conference Article
  • Citations7

A superscalar RISC processor with 160 FPRs for large scale scientific processing

  • Oct 10, 1999
  • K Shimada +4
  • Book Chapter
  • Citations3

Building Implicit Dictionaries Based on Extreme Random Clustering for Modality Recognition

  • Jan 01, 2012
  • Olivier Pauly +2
  • Conference Article

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

  • Sep 08, 2025
  • Elliott Binder +2
  • Research Article
  • Citations20

The virtual write queue

  • Jun 19, 2010
  • ACM SIGARCH Computer Architecture News
  • Jeffrey Stuecheli +4
  • Conference Article
  • Citations120

The virtual write queue

  • Jun 19, 2010
  • Jeffrey Stuecheli +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.