• Home
  • Search
  • HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations
  • Cite Icon35
  • https://doi.org/10.1109/ipdpsw.2016.43Copy DOI Icon

HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations

  • May 1, 2016
  • John D Leidel +1 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The recent advent of stacked memory devices has led to a resurgence of research associated with the fundamental memory hierarchy and associated memory pipeline. The bandwidth advantages provided by stacked logic and DRAM devices have inspired research associated with eliminating the bandwidth bottlenecks associated with many applications in high performance computing. Further, recent efforts have focused on utilizing stacked memory devices as last-level caches. In addition to the two aforementioned focus areas, a third area of research is emerging to explore augmenting the stacked memory logic layer with additional operations. This first generation of Hybrid Memory Cube (HMC) devices provided rudimentary atomic memory operations. The Gen2 Hybrid Memory Cube devices provide more expressive atomic memory operations that include primitive integer arithmetic operations. Despite the inclusion of more expressive arithmetic operations, many users have expressed interest in more complex and potentially orthogonal custom memory cube, or CMC, operations in future revisions of the Hybrid Memory Cube specification. This work presents recent development associated with the HMC-Sim Hybrid Memory Cube simulation framework that provides users a powerful infrastructure to experiment and research augmented custom memory cube, or CMC, operations within the current Gen2 Hybrid Memory Cube device infrastructure. We provide an overview of extending the original HMC-Sim simulation infrastructure to include support for CMC operations with requiring users to modify the core simulation code base. We also present three examples of building and utilizing custom, user-defined CMC operations in sample simulations to exhibit potential application speedup with future HMC device specifications. In doing so, we present a model to replace traditional thread mutexes with custom HMC mutex commands.

Similar Papers
  • Research Article

PerfSuite: An Automatic Performance Suite for HPC Applications

  • Jan 14, 2026
  • Journal of Circuits, Systems and Computers
  • Qiong Jiang +5
  • Conference Article
  • Citations4

AWGR-based optical processor-to-memory communication for low-latency, low-energy vault accesses

  • Oct 01, 2018
  • Proceedings of the International Symposium on Memory Systems
  • Sebastian Werner +3
  • Conference Article
  • Citations3

Enabling energy efficient Hybrid Memory Cube systems with erasure codes

  • Jul 01, 2015
  • Shibo Wang +3
  • Conference Article
  • Citations22

A 5GHz 7nm L1 cache memory compiler for high-speed computing and mobile applications

  • Feb 01, 2018
  • Michael Clinton +5
  • Conference Article
  • Citations5

Multimode and single-mode fibers for data center and high-performance computing applications

  • Mar 15, 2016
  • Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE
  • Scott R Bickham
  • Conference Article
  • Citations1

Delegato: Locality-Aware Atomic Memory Operations on Chiplets

  • Oct 17, 2025
  • Víctor Soria-Pardos +5
  • Conference Article
  • Citations11

A Case Study of Designing Efficient Algorithm-based Fault Tolerant Application for Exascale Parallelism

  • May 01, 2012
  • Erlin Yao +4
  • Conference Article
  • Citations7

Performance and Cost-aware HPC in Clouds: A Network Interconnection Assessment

  • Jul 01, 2020
  • Anderson M Maliszewski +5
  • Conference Article
  • Citations2

Simulative analysis of dynamic scheduling heuristics for reconfigurable computing of parallel applications

  • Feb 22, 2006
  • Rajagopal Subramaniyan +3
  • PDF
  • Research Article
  • Citations5

On the Use of Probabilistic Worst-Case Execution Time Estimation for Parallel Applications in High Performance Systems

  • Mar 01, 2020
  • Mathematics
  • Matteo Fusi +6
  • Research Article
  • Citations10

SOMA: Observability, monitoring, and in situ analytics for exascale applications

  • Jun 02, 2024
  • Concurrency and Computation: Practice and Experience
  • Dewi Yokelson +9
  • Conference Article
  • Citations10

SciDP: Support HPC and Big Data Applications via Integrated Scientific Data Processing

  • Sep 01, 2018
  • Kun Feng +3
  • Conference Article
  • Citations1

FLAMA: Architecting Floating-Point Atomic Memory Operations for Heterogeneous HPC Systems

  • Sep 10, 2025
  • Vı́Ctor Soria-Pardos +5
  • Conference Article

On Optimizing Checkpoint Restoration for HPC Applications: Leveraging Merkle Trees and Asynchronous I/O

  • Jul 20, 2025
  • Zackary Malkmus +6
  • Conference Article
  • Citations36

Large Vector Extensions Inside the HMC

  • Jan 01, 2016
  • Marco A Z Alves +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.