• Home
  • Search
  • An NoC Traffic Compiler for Efficient FPGA Implementation of Sparse Graph-Oriented Workloads
  • Cite Icon30
  • https://doi.org/10.1155/2011/745147Copy DOI Icon

An NoC Traffic Compiler for Efficient FPGA Implementation of Sparse Graph-Oriented Workloads

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Parallel graph-oriented applications expressed in the Bulk-Synchronous Parallel (BSP) and Token Dataflow compute models generate highly-structured communication workloads from messages propagating along graph edges. We can statially expose this structure to traffic compilers and optimization tools to reshape and reduce traffic for higher performance (or lower area, lower energy, lower cost). Such offline traffic optimization eliminates the need for complex, runtime NoC hardware and enables lightweight, scalable NoCs. We perform load balancing, placement, fanout routing, and fine-grained synchronization to optimize our workloads for large networks up to 2025 parallel elements for BSP model and 25 parallel elements for Token Dataflow. This allows us to demonstrate speedups between 1.2× and 22× (3.5× mean), area reductions (number of Processing Elements) between 3× and 15× (9× mean) and dynamic energy savings between 2× and 3.5× (2.7× mean) over a range of real-world graph applications in the BSP compute model. We deliver speedups of 0.5–13× (geomean 3.6×) for Sparse Direct Matrix Solve (Token Dataflow compute model) applied to a range of sparse matrices when using a high-quality placement algorithm. We expect such traffic optimization tools and techniques to become an essential part of the NoC application-mapping flow.

Similar Papers
  • Book Chapter
  • Citations2

Portability of Performance in the BSP Model

  • Jan 01, 1999
  • Jonathan Hill
  • Conference Article
  • Citations3

Agent based ServiceBSP Model with Superstep Service for Grid Computing

  • Aug 01, 2007
  • Weikai Miao +1
  • Conference Article
  • Citations3

Message passing over windows-based desktop grids

  • Nov 27, 2006
  • Carlos Queiroz +2
  • Research Article
  • Citations10

Mock BSPlib for Testing and Debugging Bulk Synchronous Parallel Software

  • Mar 01, 2017
  • Parallel Processing Letters
  • Wijnand Suijlen
  • Research Article
  • Citations36

FORMAL PROOFS OF FUNCTIONAL BSP PROGRAMS

  • Sep 01, 2003
  • Parallel Processing Letters
  • Frédéric Gava
  • Book Chapter
  • Citations18

Load Balancing Strategies in a Web Computing Environment

  • Jan 01, 2006
  • Olaf Bonorden +2
  • Conference Article
  • Citations92

RFH: A Resilient, Fault-Tolerant and High-Efficient Replication Algorithm for Distributed Cloud Storage

  • Sep 01, 2012
  • Yanzhen Qu +1
  • Research Article
  • Citations68

EASM: Efficiency-aware switch migration for balancing controller loads in software-defined networking

  • Feb 02, 2018
  • Peer-to-Peer Networking and Applications
  • Tao Hu +3
  • Research Article
  • Citations2

Equivalence between Graph Spectral Clustering and Column Subset Selection (Student Abstract)

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Guihong Wan +3
  • Research Article
  • Citations51

Enhancement of mechanical properties of low stacking fault energy brass processed by cryorolling followed by short-annealing

  • Nov 15, 2014
  • Materials & Design
  • Ravi Kumar +4
  • Research Article
  • Citations12

Logical Effort Framework for CNFET-Based VLSI Circuits for Delay and Area Optimization

  • Mar 01, 2019
  • IEEE Transactions on Very Large Scale Integration (VLSI) Systems
  • Muhammad Ali +2
  • Research Article
  • Citations48

Targeted implementation of cool roofs for equitable urban adaptation to extreme heat

  • Oct 29, 2021
  • Science of The Total Environment
  • Ashley M Broadbent +4
  • Conference Article
  • Citations1

An Indoor Planning Tool with Both Antennas Placement and Wiring Optimization

  • Oct 01, 1999
  • B Gloria +3
  • Conference Article
  • Citations29

Application of Parallel (MIMD) Computers to Reservoir Simulation

  • Feb 01, 1987
  • SPE Symposium on Reservoir Simulation
  • S L Scott +3
  • PDF
  • Research Article
  • Citations9

Area- and energy-efficient CORDIC accelerators in deep sub-micron CMOS technologies

  • Sep 18, 2012
  • Advances in Radio Science
  • U Vishnoi +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.