• Home
  • Search
  • Efficient hardware support for the Partitioned Global Address Space
  • Cite Icon25
  • https://doi.org/10.1109/ipdpsw.2010.5470851Copy DOI Icon

Efficient hardware support for the Partitioned Global Address Space

  • Apr 1, 2010
  • Holger Froning +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

We present a novel architecture of a communication engine for non-coherent distributed shared memory systems. The shared memory is composed by a set of nodes exporting their memory. Remote memory access is possible by forwarding local load or store transactions to remote nodes. No software layers are involved in a remote access, neither on origin or target side: a user level process can directly access remote locations without any kind of software involvement. We have implemented the architecture as an FPGA-based prototype in order to demonstrate the functionality of the complete system. This prototype also allows real world measurements in order to show the performance potential of this architecture, in particular for fine grain memory accesses like they are typically used for synchronization tasks.

Similar Papers
  • Conference Article
  • Citations1

Intra-Epoch Message Scheduling To Exploit Unused or Residual Overlapping Potential

  • Sep 09, 2014
  • Judicael A Zounmevo +1
  • Research Article
  • Citations4

Algorithmic optimizations of a conjugate gradient solver on shared memory architectures

  • Jan 01, 2006
  • International Journal of Parallel, Emergent and Distributed Systems
  • Henrik Löf +1
  • Conference Article
  • Citations6

Modeling and analysis of remote memory access programming

  • Oct 19, 2016
  • Andrei Marian Dan +3
  • Conference Article
  • Citations5

Optimizing NEURON brain simulator with Remote Memory Access on distributed memory systems

  • Dec 01, 2015
  • Danish Shehzad +1
  • Research Article
  • Citations4

Optimizing memory access traffic via runtime thread migration for on-chip distributed memory systems

  • Jun 24, 2014
  • The Journal of Supercomputing
  • Weiwei Fu +3
  • Dissertation

Contention resolution and memory load balancing algorithms on distributed shared memory multiprocessors

  • Aug 01, 2005
  • Mehmet Fatih Akay +1
  • Research Article
  • Citations78

NUMA-aware graph-structured analytics

  • Jan 24, 2015
  • ACM SIGPLAN Notices
  • Kaiyuan Zhang +2
  • Research Article
  • Citations34

Enabling Highly-Scalable Remote Memory Access Programming with MPI-3 One Sided

  • Jan 01, 2014
  • Scientific Programming
  • Robert Gerstenberger +2
  • Conference Article

EDM: An Ultra-Low Latency Ethernet Fabric for Memory Disaggregation

  • Mar 30, 2025
  • Weigao Su +1
  • Research Article
  • Citations15

NumaGiC

  • Mar 14, 2015
  • ACM SIGARCH Computer Architecture News
  • Lokesh Gidra +4
  • Book Chapter
  • Citations4

Compiler Optimizations for Non-contiguous Remote Data Movement

  • Jan 01, 2014
  • Timo Schneider +2
  • Conference Article
  • Citations5

Performance analysis of parallel hash join algorithms on a distributed shared memory machine implementation and evaluation on HP exemplar SPP 1600

  • Feb 23, 1998
  • M Nakano +2
  • Conference Article
  • Citations8

Extended task queuing: active messages for heterogeneous systems

  • Nov 13, 2016
  • Michael Lebeane +15
  • Conference Article
  • Citations11

RDMA control support for fine-grain parallel computations

  • Jan 01, 2004
  • A Smyk +1
  • Conference Article

Rmalloc() and rpipe()

  • Jun 12, 2018
  • Udayanga Wickramasinghe +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.