• Home
  • Search
  • Taking GPU Programming Models to Task for Performance Portability
  • Cite Icon8
  • https://doi.org/10.1145/3721145.3730423Copy DOI Icon

Taking GPU Programming Models to Task for Performance Portability

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Portability is critical to ensuring high productivity in developing and maintaining scientific software as the diversity in on-node hardware architectures increases. While several programming models provide portability for diverse GPU systems, they don't make any guarantees about performance portability. In this work, we explore several programming models -- CUDA, HIP, Kokkos, RAJA, OpenMP, OpenACC, and SYCL, to assess the consistency of their performance across NVIDIA and AMD GPUs. We use five proxy applications from different scientific domains, create implementations where missing, and use them to present a comprehensive comparative evaluation of the performance portability of these programming models. We provide a Spack scripting-based methodology to ensure reproducibility of experiments conducted in this work. Finally, we analyze the reasons for why some programming models underperform in certain scenarios and in some cases, present performance optimizations to the proxy applications.

Similar Papers
  • Conference Article
  • Citations17

Case Study of Using Kokkos and SYCL as Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

  • Nov 01, 2021
  • Amanda S Dufek +6
  • Conference Article
  • Citations21

First Experiences in Performance Benchmarking with the New SPEChpc 2021 Suites

  • May 01, 2022
  • Holger Brunst +9
  • Conference Article
  • Citations6

Performance Portability of Sparse Block Diagonal Matrix Multiple Vector Multiplications on GPUs

  • Nov 01, 2022
  • Khaled Z Ibrahim +2
  • Book Chapter
  • Citations2

Portability of Performance in the BSP Model

  • Jan 01, 1999
  • Jonathan Hill
  • Research Article
  • Citations7

Portable performance on heterogeneous architectures

  • Mar 16, 2013
  • ACM SIGPLAN Notices
  • Phitchaya Mangpo Phothilimthana +3
  • Research Article
  • Citations20

New capabilities of the Monte Carlo dose engine ARCHER-RT: Clinical validation of the Varian TrueBeam machine for VMAT external beam radiotherapy.

  • Apr 13, 2020
  • Medical Physics
  • David P Adam +4
  • Research Article
  • Citations39

A comparative study of GPU programming models and architectures using neural networks

  • May 31, 2011
  • The Journal of Supercomputing
  • Vivek K Pallipuram +2
  • Book Chapter
  • Citations10

Evaluating Performance Portability of OpenMP for SNAP on NVIDIA, Intel, and AMD GPUs Using the Roofline Methodology

  • Jan 01, 2021
  • Neil A Mehta +4
  • Research Article
  • Citations16

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

  • Sep 06, 2021
  • Computer Physics Communications
  • Saaketh Desai +2
  • Conference Article

Beyond Guess and Check: Quantifying the Fidelity of Proxy Applications

  • Nov 07, 2025
  • Si Chen +4
  • Book Chapter
  • Citations8

The Parallel Boost Graph Library 2.0

  • May 26, 2022
  • Nicholas Edmonds +1
  • Book Chapter
  • Citations2

Examining Performance Portability with Kokkos for an Ewald Sum Coulomb Solver

  • Jan 01, 2020
  • Rene Halver +2
  • PDF
  • Research Article
  • Citations48

The ESCAPE project: Energy-efficient Scalable Algorithms for Weather Prediction at Exascale

  • Oct 22, 2019
  • Geoscientific Model Development
  • Andreas Müller +58
  • Conference Article
  • Citations8

Performance Portable Applications for Hardware Accelerators: Lessons Learned from SPEC ACCEL

  • May 01, 2015
  • Guido Juckeland +2
  • Single Report

Proxy Applications for Converged Workloads: DMC LDRD Initiative

  • Sep 26, 2023
  • Sayan Ghosh +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.