• Home
  • Search
  • Orchestrating data transfer for the cell/B.E. processor
  • Cite Icon26
  • https://doi.org/10.1145/1375527.1375570Copy DOI Icon

Orchestrating data transfer for the cell/B.E. processor

  • Jun 7, 2008
  • Tong Chen +2 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

In heterogeneous multi-core systems, such as the Cell/B.E. or certain embedded systems, the accelerator core has its own fast local memory without hardware supported coherence between the local and global memories. It is software's responsibility to dynamically transfer the working set into the local memory when the total data set is too large to fit in the local memory. The data can be transferred through either a software controlled cache or a direct buffer. Such a software cache can maintain correctness and exploit reuse among references, especially when complicated aliasing or data dependences exist. However, the software cache introduces the extra overhead of cache lookup. Direct buffering, on the other hand, is fast but is limited by the compiler's ability to disambiguate memory references. It is desirable to judiciously use both methods, for irregular and regular accesses respectively. However, when a datum resides in both the software cache and the direct buffer, coherence problems occur.In this paper, we propose a solution which provides compile time analysis and runtime maintenance to address this coherence issue. We use compiler analysis to guarantee that there is no access to software cache within the local live range of a direct buffer, and rely on runtime support to update values from or to software cache at the entry or exit of the direct buffer. Further, we present a global data flow analysis design to eliminate redundant coherence maintenance, and overlap computation and DMA accesses to reduce runtime overhead. We have implemented this method in our Single Source Compiler for Cell, and have conducted experiments with the NAS OpenMP benchmarks. The results show that our method maintains correctness while keeping most of the opportunities for direct buffering. The execution performance can increase more than 3x compared to approaches using only the software cache. Furthermore, compile time analysis can reduce 90% of the runtime updates, thereby improving performance by 20% further.

Similar Papers
  • Research Article
  • Citations9

Memory and genocide in graphic novels: the Holocaust as paradigm

  • Aug 22, 2017
  • Journal of Graphic Novels and Comics
  • Joanne Pettitt
  • PDF
  • Research Article
  • Citations1

The location of transcultural memory in Vikram Seth’s memoir Two Lives (2005)

  • Nov 22, 2019
  • Journal of Aesthetics & Culture
  • Nadia Butt
  • Book Chapter
  • Citations52

Offload – Automating Code Migration to Heterogeneous Multicore Systems

  • Jan 01, 2010
  • Pete Cooper +5
  • Conference Article
  • Citations4

Evaluation of compiler and runtime library approaches for supporting parallel regular applications

  • Mar 30, 1998
  • D.R Chakrabarti +2
  • Conference Article
  • Citations30

Optimizing memory system performance for communication in parallel computers

  • Jan 01, 1995
  • T Stricker +1
  • Conference Article
  • Citations3

Boosting Memory Performance of Many-Core FPGA Device through Dynamic Precedence Graph

  • Apr 01, 2013
  • Yu Bai +4
  • Research Article
  • Citations17

An Efficient GPU Cache Architecture for Applications with Irregular Memory Access Patterns

  • Jun 17, 2019
  • ACM Transactions on Architecture and Code Optimization
  • Bingchao Li +4
  • Conference Article
  • Citations5

Fine-Grained Task Migration for Graph Algorithms Using Processing in Memory

  • May 01, 2016
  • Paula Aguilera +3
  • PDF
  • Research Article
  • Citations14

Regular access and adherence to medications of the specialized component of pharmaceutical services

  • Nov 13, 2017
  • Revista de Saúde Pública
  • Janaína Soder Fritzen +2
  • Research Article
  • Citations12

JTAG-based remote configuration of FPGAs over optical fibers

  • Jan 01, 2015
  • Journal of Instrumentation
  • B Deng +14
  • Research Article
  • Citations1

Multiply-and-Fire: An Event-Driven Sparse Neural Network Accelerator

  • Dec 14, 2023
  • ACM Transactions on Architecture and Code Optimization
  • Miao Yu +3
  • Research Article

A study on the role of gram panchayats and awareness on Jal Shakti Abhiyan in South Andaman district

  • Nov 01, 2024
  • International Journal of Agriculture Extension and Social Development
  • Debendra Nath Dash +2
  • Research Article
  • Citations10

Implementing Sparse Matrix-Vector Multiplication with QCSR on GPU

  • Mar 01, 2013
  • Applied Mathematics & Information Sciences
  • Jilin Zhang +5
  • Conference Article
  • Citations2

Reducing Compiler-Inserted Instrumentation in Unified-Parallel-C Code Generation

  • Oct 01, 2014
  • Michail Alvanosl +4
  • Conference Article
  • Citations2

Analysis and VLSI architecture of update step in motion-compensated temporal filtering

  • May 21, 2006
  • Chih-Chi Cheng +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.