• Home
  • Search
  • FLAMA: Architecting Floating-Point Atomic Memory Operations for Heterogeneous HPC Systems
  • Cite Icon1
  • https://doi.org/10.1109/dsd67783.2025.00066Copy DOI Icon

FLAMA: Architecting Floating-Point Atomic Memory Operations for Heterogeneous HPC Systems

  • Sep 10, 2025
  • Vı́ctor Soria-Pardos +5 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Current heterogeneous systems integrate generalpurpose Central Processing Units (CPUs), Graphics Processing Units (GPUs), and Neural Processing Units (NPUs). The efficient use of such systems requires a significant programming effort to distribute computation and synchronize across devices, which usually involves using Atomic Memory Operations (AMOs). Arm recently launched a floating-point Atomic Memory Operations (FAMOs) extension to perform atomic updates on floating-point data types specifically. This work characterizes and models heterogeneous architectures to understand how floating-point AMOs impact graph, Machine Learning (ML), and high-performance computing (HPC) workloads. Our analysis shows that many AMOs are performed on floating-point data, which modern systems execute using inefficient compare-and-swap (CAS) constructs. Therefore, replacing CASbased constructs with FAMOs can improve a wide range of workloads. Moreover, we analyze the trade-offs of executing FAMOs at different memory hierarchy levels, either in private caches (near) or remotely in shared caches (far). We have extended the widely used AMBA CHI protocol to evaluate such FAMO support on a simulated chiplet-based heterogeneous architecture. While near FAMOs achieve an average $1.34 \times$ speed-up, far FAMOs reach an average $1.58 \times$ speed-up. We conclude that FAMOs can bridge the gap between CPU architecture and accelerators and enabling synchronization in key application domains.

Similar Papers
  • Conference Article
  • Citations1

Delegato: Locality-Aware Atomic Memory Operations on Chiplets

  • Oct 17, 2025
  • Víctor Soria-Pardos +5
  • Research Article
  • Citations289

Fast analysis of molecular dynamics trajectories with graphics processing units—Radial distribution function histogramming

  • Feb 06, 2011
  • Journal of Computational Physics
  • Benjamin G Levine +2
  • Conference Article
  • Citations7

Optimizing a FIFO, scalable spin lock using consistent memory

  • Dec 04, 1996
  • I Rhee
  • Research Article
  • Citations8

Fast parallel cutoff pair interactions for molecular dynamics on heterogeneous systems

  • Jun 01, 2012
  • Tsinghua Science and Technology
  • Qiang Wu +3
  • Research Article
  • Citations12

HPMaX: heterogeneous parallel matrix multiplication using CPUs and GPUs

  • Oct 11, 2020
  • Computing
  • Homin Kang +2
  • Research Article
  • Citations13

Practical parallel AES algorithms on cloud for massive users and their performance evaluation

  • Dec 17, 2015
  • Concurrency and Computation: Practice and Experience
  • Xiongwei Fei +3
  • Research Article
  • Citations26

Efficient methods for implementation of multi-level nonrigid mass-preserving image registration on GPUs and multi-threaded CPUs

  • Jan 06, 2016
  • Computer Methods and Programs in Biomedicine
  • Nathan D Ellingwood +3
  • Research Article

The ocean model for E3SM global applications: Omega version 0.1.0 – a new high-performance computing code for exascale architectures

  • May 04, 2026
  • Geoscientific Model Development
  • Mark R Petersen +16
  • PDF
  • Research Article
  • Citations25

DdcMD: A fully GPU-accelerated molecular dynamics program for the Martini force field.

  • Jul 23, 2020
  • The Journal of Chemical Physics
  • Xiaohua Zhang +8
  • Research Article
  • Citations7

Occamy: A 432-Core Dual-Chiplet Dual-HBM2E 768-DP-GFLOP/s RISC-V System for 8-to-64-bit Dense and Sparse Computing in 12-nm FinFET

  • Apr 01, 2025
  • IEEE Journal of Solid-State Circuits
  • Paul Scheffler +14
  • Research Article

Improving MSAProbs Algorithm performance and Parallel Computing using GPU

  • Apr 18, 2023
  • International Journal of Computer Applications
  • Sally Zaki El-Hadary +2
  • Research Article
  • Citations12

High-Performance and Parallel Computing Techniques Review: Applications, Challenges and Potentials to Support Net-Zero Transition of Future Grids

  • Nov 18, 2022
  • Energies
  • Ahmed Al-Shafei +2
  • Research Article
  • Citations4

A systematic parallel strategy for generating contours from large-scale DEM data using collaborative CPUs and GPUs

  • Feb 18, 2021
  • Cartography and Geographic Information Science
  • Chen Zhou +1
  • Research Article
  • Citations31

GPU acceleration of MPAS microphysics WSM6 using OpenACC directives: Performance and verification

  • Oct 14, 2020
  • Computers & Geosciences
  • Jae Youp Kim +2
  • Conference Article
  • Citations35

HMC-Sim-2.0: A Simulation Platform for Exploring Custom Memory Cube Operations

  • May 01, 2016
  • John D Leidel +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.