• Home
  • Search
  • A Python/Fortran Implementation of the Lattice‐Boltzmann Kernel on Multiple GPU Using the OpenACC Framework
  • https://doi.org/10.1002/cpe.70518Copy DOI Icon

A Python/Fortran Implementation of the Lattice‐Boltzmann Kernel on Multiple GPU Using the OpenACC Framework

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

ABSTRACT The increasing availability of GPU accelerated architectures for high‐performance computing presents new opportunities for scientific software but also challenges due to the complexity of porting legacy codes to accelerator platforms. Directive‐based programming models such as OpenACC offer a minimally intrusive pathway to exploit GPU acceleration without requiring extensive rewriting of existing codes. The current work presents a comprehensive performance and portability study of a LatticeBoltzmann Method solver (PyLB) originally written in Python, Mpi4Py, and Fortran for CPU architectures, which is ported to GPUs using OpenACC directives applied to the Fortran routines. The performance of the solver is evaluated on NVIDIA V100, A100, and H100 GPUs available on the Jean Zay supercomputer from Institute for Development and Resources in Intensive Scientific Computing (IDRIS) in France. Roofline analysis and extensive strong and weak scalability tests are conducted, showing that the GPU‐enabled version of PyLB scales efficiently across multiple GPUs. The solver achieves performance on the H100 GPU equivalent to thousands of CPU cores and shows strong energy and carbon efficiency advantages over traditional CPU‐based simulations. The implementation is validated using classical benchmarks, including the decaying Taylor‐Green vortex and the flow over a 3‐D sphere. The results confirm the physical accuracy of the GPU port while highlighting its computational and environmental advantages.

Similar Papers
  • Conference Article

CO2-Filled Vertical Geothermal Boreholes: Modeling and Application

  • Nov 15, 2013
  • Parham Eslami-Nejad +2
  • Research Article
  • Citations24

Performance of preconditioned iterative linear solvers for cardiovascular simulations in rigid and deformable vessels.

  • Feb 06, 2019
  • Computational Mechanics
  • Jongmin Seo +2
  • Conference Article
  • Citations8

HPIPE NX: Boosting CNN Inference Acceleration Performance with AI-Optimized FPGAs

  • Dec 05, 2022
  • Marius Stan +3
  • Research Article
  • Citations3

Optimal Decision Trees for Feature Based Parameter Tuning: Integer Programming Model and VNS Heuristic

  • Apr 01, 2018
  • Electronic Notes in Discrete Mathematics
  • Matheus Guedes Vilas Boas +2
  • Conference Article
  • Citations12

Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance Assurance

  • Apr 22, 2024
  • EuroSys 2024 - Proceedings of the 2024 European Conference on Computer Systems
  • Yijia Zhang +4
  • PDF
  • Research Article
  • Citations16

Multi-GPU Support on Single Node Using Directive-Based Programming Model

  • Jan 01, 2015
  • Scientific Programming
  • Rengan Xu +3
  • Research Article
  • Citations20

Optimizing makespan and resource utilization for multi-DNN training in GPU cluster

  • Jun 24, 2021
  • Future Generation Computer Systems
  • Zhongjin Li +5
  • Conference Article
  • Citations2

GPU Acceleration of a High-Order CFD Program

  • Jun 27, 2020
  • Shengxiang Wang +2
  • Book Chapter
  • Citations8

Chapter 40 - GPU Acceleration of Iterative Digital Breast Tomosynthesis

  • Jan 01, 2011
  • GPU Computing Gems Emerald Edition
  • Dana Schaa +7
  • Research Article

Leveraging AI in Golang: Building Intelligent Applications with Go

  • Apr 15, 2025
  • European Journal of Computer Science and Information Technology
  • Sruthi Deva
  • PDF
  • Book Chapter
  • Citations2

PHINEAS: An Embedded Heterogeneous Parallel Platform

  • Jan 01, 2019
  • Nikhil Khatri +2
  • Book Chapter

A High-Level Programming Approach for Distributed Systems with Accelerators

  • Jan 01, 2012
  • Frontiers in artificial intelligence and applications
  • Steuwer Michel +2
  • Book Chapter
  • Citations110

Hierarchical Place Trees: A Portable Abstraction for Task Parallelism and Data Movement

  • Jan 01, 2010
  • Yonghong Yan +3
  • Research Article
  • Citations18

Validating large-scale quantum machine learning: efficient simulation of quantum support vector machines using tensor networks

  • Feb 20, 2025
  • Machine Learning: Science and Technology
  • Kuan-Cheng Chen +8
  • PDF
  • Research Article
  • Citations17

Parametric Analysis of Design Parameter Effects on the Performance of a Solar Desiccant Evaporative Cooling System in Brisbane, Australia

  • Jun 25, 2017
  • Energies
  • Yunlong Ma +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.