• Home
  • Search
  • Floating-point sparse matrix-vector multiply for FPGAs
  • Open Access IconOpen Access
  • Cite Icon165
  • https://doi.org/10.1145/1046192.1046203Copy DOI Icon

Floating-point sparse matrix-vector multiply for FPGAs

  • Feb 20, 2005
  • Michael Delorimier +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Large, high density FPGAs with high local distributed memory bandwidth surpass the peak floating-point performance of high-end, general-purpose processors. Microprocessors do not deliver near their peak floating-point performance on efficient algorithms that use the Sparse Matrix-Vector Multiply (SMVM) kernel. In fact, microprocessors rarely achieve 33% of their peak floating-point performance when computing SMVM. We develop and analyze a scalable SMVM implementation on modern FPGAs and show that it can sustain high throughput, near peak, floating-point performance. Our implementation consists of logic design as well as scheduling and data placement techniques. For benchmark matrices from the Matrix Market Suite we project 1.5 double precision Gflops/FPGA for a single VirtexII-6000-4 and 12 double precision Gflops for 16 Virtex IIs (750 Mflops/FPGA). We also analyze the asymptotic efficiency of our architecture as parallelism scales using a constant rent-parameter matrix model. This demonstrates that our data placement techniques provide an asymptotic scaling benefit. While FPGA performance is attractive, higher performance is possible if we re-balance the hardware resources in FPGAs with embedded memories. We show that sacrificing half the logic area for memory area rarely degrades performance and improves performance for large matrices, by up to 5 times. We also 0 the performance effect of adding custom floating-point using a simple area model to preserve total chip area. Sacrificing logic for memory and custom floating-point units increases single FPGA performance to 5 double precision Gflops.

Similar Papers
  • Conference Article
  • Citations4

On improving performance and energy profiles of sparse scientific applications

  • Apr 25, 2006
  • Konrad Malkowski +3
  • Conference Article
  • Citations19

Performing Floating-Point Accumulation on a Modern FPGA in Single and Double Precision

  • Jan 01, 2010
  • Tarek Ould Bachir +1
  • Conference Article
  • Citations4

Evaluation of a Directive-Based GPU Programming Approach for High-Order Unstructured Mesh Computational Fluid Dynamics

  • Jun 26, 2017
  • Kunal Puri +2
  • Research Article
  • Citations5

A sparse matrix–vector multiplication based algorithm for accurate density matrix computations on systems of millions of atoms

  • Feb 15, 2018
  • Computer Physics Communications
  • Purnima Ghale +1
  • Conference Article
  • Citations51

An FPGA framework for edge-centric graph processing

  • May 08, 2018
  • Shijie Zhou +3
  • Conference Article
  • Citations23

A reduced-precision streaming SpMV architecture for Personalized PageRank on FPGA

  • Jan 18, 2021
  • Virtual Community of Pathological Anatomy (University of Castilla La Mancha)
  • Alberto Parravicini +2
  • Conference Article
  • Citations1

Multi-Mode Transprecision Sparse Matrix-Vector Multiplication Engine for PageRank

  • Feb 06, 2022
  • Whijin Kim +3
  • Book Chapter
  • Citations4

General-Purpose DSP Processors

  • Jan 01, 2013
  • Jarmo Takala
  • Research Article

GPU Accelerated Reconstruction in Compton Scattering Tomography Using Matrix Compression

  • Feb 01, 2014
  • Applied Mechanics and Materials
  • Yu Fei Yu +5
  • Conference Article

Optimization and implementation of SCME channel model on GPP

  • Oct 01, 2013
  • Meng Xu +2
  • Conference Article
  • Citations8

Optimizing Sparse Matrix Vector Multiplication Using Diagonal Storage Matrix Format

  • Sep 01, 2010
  • Liang Yuan +3
  • Conference Article
  • Citations8

A memory accelerator with gather functions for bandwidth-bound irregular applications

  • Nov 13, 2011
  • Noboru Tanabe +6
  • Conference Article
  • Citations7

ICE: A General and Validated Energy Complexity Model for Multithreaded Algorithms

  • Dec 01, 2016
  • Vi Ngoc-Nha Tran +1
  • Conference Article
  • Citations3

Energy consumption optimization of the Total-FETI solver and BLAS routines by changing the CPU frequency

  • Jul 01, 2016
  • David Horak +4
  • Book Chapter
  • Citations3

Chapter 10 - Parallel Patterns: Sparse Matrix–Vector Multiplication: An Introduction to Compaction and Regularization in Parallel Algorithms

  • Jan 01, 2013
  • Programming Massively Parallel Processors
  • David B Kirk +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.