• Home
  • Search
  • An evaluation of memory consistency models for shared-memory systems with ILP processors
  • Cite Icon78
  • https://doi.org/10.1145/237090.237142Copy DOI Icon

An evaluation of memory consistency models for shared-memory systems with ILP processors

  • Sep 1, 1996
  • Vijay S Pai +3 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Relaxed consistency models have been shown to significantly outperform sequential consistency for single-issue, statically scheduled processors with blocking reads. However, current microprocessors aggressively exploit instruction-level parallelism (ILP) using methods such as multiple issue, dynamic scheduling, and non-blocking reads. Researchers have conjectured that two techniques, hardware-controlled non-binding prefetching and speculative loads, have the potential to equalize the hardware performance of memory consistency models on such processors.This paper performs the first detailed quantitative comparison of several implementations of sequential consistency and release consistency optimized for aggressive ILP processors. Our results indicate that hardware prefetching and speculative loads dramatically improve the performance of sequential consistency. However, the gap between sequential consistency and release consistency depends on the cache write policy and the complexity of the cache-coherence protocol implementation. In most cases, release consistency significantly outperforms sequential consistency, but for two applications, the use of a write-back primary cache and a more complex cache-coherence protocol nearly equalizes the performance of the two models.We also observe that the existing techniques, which require on-chip hardware modifications, enhance the performance of release consistency only to a small extent. We propose two new software techniques --- fuzzy acquires and selective acquires --- to achieve more overlap than allowed by the previous implementations of release consistency. To enhance methods for overlapping acquires, we also propose a technique to eliminate control dependences caused by an acquire loop, using a small amount of off-chip hardware called the synchronization buffer.

Similar Papers
  • Research Article
  • Citations6

The interaction of software prefetching with ILP processors in shared-memory systems

  • May 01, 1997
  • ACM SIGARCH Computer Architecture News
  • Parthasarathy Ranganathan +3
  • Dissertation

Efficient openMP over sequentially consistent distributed shared memory systems

  • Jul 20, 2011
  • Juan José Costa Prats
  • Research Article
  • Citations3

Efficient sequential consistency via conflict ordering

  • Mar 03, 2012
  • ACM SIGARCH Computer Architecture News
  • Changhui Lin +3
  • Research Article
  • Citations10

Efficient sequential consistency via conflict ordering

  • Mar 03, 2012
  • ACM SIGPLAN Notices
  • Changhui Lin +3
  • Research Article
  • Citations3

On the efficiency of image and video processing programs on instruction level parallel processors

  • Jul 01, 2002
  • Proceedings of the IEEE
  • N Zingirian +1
  • Book Chapter
  • Citations7

Implicit Transactional Memory in Kilo-Instruction Multiprocessors

  • Sep 28, 2017
  • Marco Galluzzi +7
  • Research Article
  • Citations3

A regulated transitive reduction (RTR) for longer memory race recording

  • Oct 20, 2006
  • ACM SIGOPS Operating Systems Review
  • Min Xu +2
  • Research Article
  • Citations1

A regulated transitive reduction (RTR) for longer memory race recording

  • Oct 20, 2006
  • ACM SIGPLAN Notices
  • Min Xu +2
  • Research Article
  • Citations6

A regulated transitive reduction (RTR) for longer memory race recording

  • Oct 20, 2006
  • ACM SIGARCH Computer Architecture News
  • Min Xu +2
  • Research Article
  • Citations198

Rsim: simulating shared-memory multiprocessors with ILP processors

  • Jan 01, 2002
  • Computer
  • C.J Hughes +3
  • Research Article
  • Citations1

Automated Robustness Verification of Concurrent Data Structure Libraries against Relaxed Memory Models

  • Oct 08, 2024
  • Proceedings of the ACM on Programming Languages
  • Kartik Nagar +3
  • Conference Article
  • Citations3

Impact of Instruction Re-Ordering on the Correctness of Shared-Memory Programs

  • Dec 07, 2005
  • L Higham +1
  • Research Article
  • Citations9

Scalability Analysis of Memory Consistency Models in NoC-Based Distributed Shared Memory SoCs

  • Jan 01, 2013
  • IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
  • A Naeem +2
  • Book Chapter
  • Citations3

Comprehensive Redundant Load Elimination for the IA-64 Architecture

  • Jan 01, 2000
  • Youngfeng Wu +1
  • Conference Article
  • Citations74

Lamport clocks

  • Jan 01, 1998
  • Manoj Plakal +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.