• Home
  • Search
  • Delegato: Locality-Aware Atomic Memory Operations on Chiplets
  • Cite Icon1
  • https://doi.org/10.1145/3725843.3756030Copy DOI Icon

Delegato: Locality-Aware Atomic Memory Operations on Chiplets

  • Oct 17, 2025
  • Víctor Soria-Pardos +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The irruption of chiplet-based architectures has been a game changer, enabling higher transistor integration and core counts in a single socket. However, chiplets impose higher and non-uniform memory access (NUMA) latencies than monolithic integration. This harms the efficiency of atomic memory operations (AMOs), which are fundamental to implementing fine-grained synchronization and concurrent data structures on large systems. AMOs are executed either near the core (near) or at a remote location within the cache hierarchy (far). On near AMOs, the core’s private cache fetches the target cache line in exclusiveness to modify it locally. Near AMOs cause significant data movement between private caches, especially harming parallel applications’ performance on chiplet-based architectures. Alternatively, far AMOs can alleviate the communication overhead by reducing data movement between processing elements. However, current multicore architectures only support one type of far AMO, which sends all updates to a single serialization point (centralized AMOs). This work introduces two new types of far AMOs, delegated and migrating, that execute AMOs remotely without centralizing updates in a single point of the cache hierarchy. Combining centralized, delegated, and migrating AMOs allows the directory to select the best location to execute AMOs. Moreover, we propose Delegato, a tracing optimization to effectively transport usage information from private caches to the directory to predict the best atomic type to issue accurately. Additionally, we design a simple predictor on top of Delegato that seamlessly selects the best placement to perform AMOs based on the data access pattern and usage activity of cores. Our evaluation using gem5 shows that Delegato can speed up applications on average by 1.07 × over centralized AMOs and by 1.13 × over the state-of-the-art AMO predictor.

Similar Papers
  • Conference Article
  • Citations1

FLAMA: Architecting Floating-Point Atomic Memory Operations for Heterogeneous HPC Systems

  • Sep 10, 2025
  • Vı́Ctor Soria-Pardos +5
  • Conference Article
  • Citations7

Optimizing a FIFO, scalable spin lock using consistent memory

  • Dec 04, 1996
  • I Rhee
  • Research Article
  • Citations289

Fast analysis of molecular dynamics trajectories with graphics processing units—Radial distribution function histogramming

  • Feb 06, 2011
  • Journal of Computational Physics
  • Benjamin G Levine +2
  • Conference Article
  • Citations35

HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations

  • May 01, 2016
  • John D Leidel +1
  • Conference Article
  • Citations9

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

  • May 01, 2021
  • Vito Giovanni Castellana +2
  • Single Report

Milestone M6 Report: Reducing Excess Data Movement Part 1

  • Mar 01, 2021
  • Ivy Peng +6
  • Conference Article
  • Citations3

Opportunity for compute partitioning in pursuit of energy-efficient systems

  • Jun 13, 2016
  • Prasenjit Chakraborty +3
  • Conference Article
  • Citations52

Reducing contention through priority updates

  • Jul 23, 2013
  • Julian Shun +3
  • Conference Article
  • Citations17

Reducing contention through priority updates

  • Feb 23, 2013
  • Julian Shun +3
  • Conference Article
  • Citations129

In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT Systems

  • Feb 01, 2018
  • Mingcong Song +7
  • Book Chapter
  • Citations9

Optimizing Fortran 90 shift operations on distributed-memory multicomputers

  • Jan 01, 1996
  • Ken Kennedy +2
  • Conference Article
  • Citations6

Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference Workloads

  • May 01, 2020
  • Ben Feinberg +5
  • Research Article
  • Citations8

PIM-GraphSCC: PIM-Based Graph Processing Using Graph’s Community Structures

  • Jul 01, 2020
  • IEEE Computer Architecture Letters
  • Newton +2
  • Research Article
  • Citations60

A Residual Replacement Strategy for Improving the Maximum Attainable Accuracy of $s$-Step Krylov Subspace Methods

  • Jan 01, 2014
  • SIAM Journal on Matrix Analysis and Applications
  • Erin Carson +1
  • Conference Article
  • Citations5

PufferFish: NUMA-Aware Work-stealing Library using Elastic Tasks

  • Dec 01, 2020
  • Vivek Kumar
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.