• Home
  • Search
  • Evaluating hybrid memory cube infrastructure to support high-performance sparse algorithms
  • Cite Icon1
  • https://doi.org/10.1145/3132402.3132435Copy DOI Icon

Evaluating hybrid memory cube infrastructure to support high-performance sparse algorithms

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This work is focused on analyzing potential performance improvements of HPC applications using stacked memories like the Hybrid Memory Cube, or HMC. We target a HPC sparse direct solver library, SuperLU [4], that performs LU decomposition and is a core piece of simulation codes like NIMROD [1]. To accelerate this library, we are interested in mapping both the computationally intense Spare Matrix-Vector (SpMV) kernels that can be implemented using matrix-matrix multiply (GEMM) calls and memory-intensive primitives like Scatter and Gather to a reconfigurable fabric tightly integrated with a 3D stacked memory. Here we provide initial results on mapping GEMM to OpenCL-based devices as well as a trace-driven evaluation of SuperLU's memory accesses with a combined FPGA and HMC platform.

Similar Papers
  • Conference Article
  • Citations35

HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations

  • May 01, 2016
  • John D Leidel +1
  • Conference Article
  • Citations3

Enabling energy efficient Hybrid Memory Cube systems with erasure codes

  • Jul 01, 2015
  • Shibo Wang +3
  • Conference Article
  • Citations2

An online parallel CRC32 realization for Hybrid Memory Cube protocol

  • Dec 01, 2013
  • Khaled Salah
  • Conference Article
  • Citations36

Large Vector Extensions Inside the HMC

  • Jan 01, 2016
  • Marco A Z Alves +3
  • Conference Article
  • Citations13

Exploring Cache Size and Core Count Tradeoffs in Systems with Reduced Memory Access Latency

  • Feb 01, 2016
  • Paulo C Santos +4
  • Book Chapter
  • Citations8

Effective FPGA Architecture for General CRC

  • Jan 01, 2019
  • Lukáš Kekely +2
  • Conference Article
  • Citations4

AWGR-based optical processor-to-memory communication for low-latency, low-energy vault accesses

  • Oct 01, 2018
  • Proceedings of the International Symposium on Memory Systems
  • Sebastian Werner +3
  • Research Article
  • Citations9

Two-level main memory co-design: Multi-threaded algorithmic primitives, analysis, and simulation

  • Jan 03, 2017
  • Journal of Parallel and Distributed Computing
  • Michael A Bender +9
  • Research Article
  • Citations20

Ultra-Short-Reach Interconnects for Die-to-Die Links: Global Bandwidth Demands in Microcosm

  • Jan 01, 2019
  • IEEE Solid-State Circuits Magazine
  • Behzad Dehlaghi +2
  • Conference Article
  • Citations3

Towards Application-Specific Address Mapping for Emerging Memory Devices

  • Sep 28, 2020
  • Shashank Adavally +1
  • Research Article

Power-Time Exploration Tools for NMP-Enabled Systems

  • Sep 28, 2019
  • Electronics
  • Chae Eun Rhee +4
  • Research Article
  • Citations16

Enabling fast and energy-efficient FM-index exact matching using processing-near-memory

  • Mar 02, 2021
  • The Journal of Supercomputing
  • Jose M Herruzo +3
  • Research Article
  • Citations1

3D Technology Applications Market Trends & Key Challenges

  • Jan 01, 2014
  • Additional Conferences (Device Packaging, HiTEC, HiTEN, and CICMT)
  • Thibault Buisson +3
  • Conference Article
  • Citations56

Multi-GPU System Design with Memory Networks

  • Dec 01, 2014
  • Gwangsun Kim +3
  • Conference Article
  • Citations146

GraphQ

  • Oct 12, 2019
  • Youwei Zhuo +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.