• Home
  • Search
  • Efficient openMP over sequentially consistent distributed shared memory systems
  • https://doi.org/10.5821/dissertation-2117-94576Copy DOI Icon

Efficient openMP over sequentially consistent distributed shared memory systems

  • Jul 20, 2011
  • Juan José Costa Prats
Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

Nowadays clusters are one of the most used platforms in High Performance Computing and most programmers use the Message Passing Interface (MPI) library to program their applications in these distributed platforms getting their maximum performance, although it is a complex task. On the other side, OpenMP has been established as the de facto standard to program applications on shared memory platforms because it is easy to use and obtains good performance without too much effort. So, could it be possible to join both worlds? Could programmers use the easiness of OpenMP in distributed platforms? A lot of researchers think so. And one of the developed ideas is the distributed shared memory (DSM), a software layer on top of a distributed platform giving an abstract shared memory view to the applications. Even though it seems a good solution it also has some inconveniences. The memory coherence between the nodes in the platform is difficult to maintain (complex management, scalability issues, high overhead and others) and the latency of the remote-memory accesses which can be orders of magnitude greater than on a shared bus due to the interconnection network. Therefore this research improves the performance of OpenMP applications being executed on distributed memory platforms using a DSM with sequential consistency evaluating thoroughly the results from the NAS parallel benchmarks. The vast majority of designed DSMs use a relaxed consistency model because it avoids some major problems in the area. In contrast, we use a sequential consistency model because we think that showing these potential problems that otherwise are hidden may allow the finding of some solutions and, therefore, apply them to both models. The main idea behind this work is that both runtimes, the OpenMP and the DSM layer, should cooperate to achieve good performance, otherwise they interfere one each other trashing the final performance of applications. We develop three different contributions to improve the performance of these applications: (a) a technique to avoid false sharing at runtime, (b) a technique to mimic the MPI behaviour, where produced data is forwarded to their consumers and, finally, (c) a mechanism to avoid the network congestion due to the DSM coherence messages. The NAS Parallel Benchmarks are used to test the contributions. The results of this work shows that the false-sharing problem is a relative problem depending on each application. Another result is the importance to move the data flow outside of the critical path and to use techniques that forwards data as early as possible, similar to MPI, benefits the final application performance. Additionally, this data movement is usually concentrated at single points and affects the application performance due to the limited bandwidth of the network. Therefore it is necessary to provide mechanisms that allows the distribution of this data through the computation time using an otherwise idle network. Finally, results shows that the proposed contributions improve the performance of OpenMP applications on this kind of environments.

Similar Papers
  • Research Article
  • Citations9

Scalability Analysis of Memory Consistency Models in NoC-Based Distributed Shared Memory SoCs

  • Jan 01, 2013
  • IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems
  • A Naeem +2
  • Research Article
  • Citations1

Performance visualization for distributed shared memory systems

  • Jan 01, 2001
  • Scalable Computing Practice and Experience
  • James E Lumpp +3
  • Research Article
  • Citations15

An interaction of coherence protocols and memory consistency models in DSM systems

  • Oct 01, 1997
  • ACM SIGOPS Operating Systems Review
  • Weisong Shi +2
  • Research Article

Towards implementation of a novel scheme for data prefetching on distributed shared memory systems

  • Feb 26, 2008
  • The Journal of Supercomputing
  • Hsiao-Hsi Wang +3
  • Research Article
  • Citations1

CBBR: enabling distributed shared memory-based coordination among mobile robots

  • Jul 18, 2016
  • Science China Information Sciences
  • Xue Jiang +1
  • Conference Article
  • Citations5

Performance analysis of parallel hash join algorithms on a distributed shared memory machine implementation and evaluation on HP exemplar SPP 1600

  • Feb 23, 1998
  • M Nakano +2
  • Research Article
  • Citations1

Random Priority-Based Thrashing Control for Distributed Shared Memory

  • Mar 01, 2020
  • IEEE Transactions on Parallel and Distributed Systems
  • Yi-Wei Ci +4
  • Research Article
  • Citations2

Specification-based Verification in a Distributed Shared Memory Simulation Model

  • Oct 22, 2009
  • SIMULATION
  • Worawan Marurngsith +1
  • Research Article
  • Citations32

Software distributed shared memory over virtual interface architecture: implementation and performance

  • Oct 10, 2000
  • View
  • Muralidharan Rangarajan +1
  • Conference Article
  • Citations59

High performance MPI design using unreliable datagram for ultra-scale InfiniBand clusters

  • Jun 17, 2007
  • Matthew J Koop +3
  • Conference Article
  • Citations74

Lamport clocks

  • Jan 01, 1998
  • Manoj Plakal +3
  • Research Article
  • Citations4

Using confidence interval to summarize the evaluating results of DSM systems

  • Jan 01, 2000
  • Journal of Computer Science and Technology
  • Weisong Shi +2
  • Conference Article
  • Citations12

Reliable distributed shared memory

  • Oct 11, 1990
  • B.D Fleisch
  • Research Article
  • Citations18

Distributed Parallel Computing Using Navigational Programming

  • Feb 01, 2004
  • International Journal of Parallel Programming
  • Lei Pan +5
  • Research Article

NONH: A new cache-based coherence protocol for linked list structure DSM system and its performance evaluation

  • Jul 01, 1996
  • Journal of Computer Science and Technology
  • Zhiyi Fang +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.