- Research Article
49
- 10.1016/j.jpdc.2014.07.006
Regularizing graph centrality computations
- Aug 07, 2014
- Journal of Parallel and Distributed Computing
- Ahmet Erdem Sarıyüce + 3 more +3
Regularizing graph centrality computations
The HPCMP CREATE[TM]-AV Helios overset CFD framework has for long used patch-based Cartesian solver SAMCart for adaptive mesh refinement in off-body meshes. This work presents the development of a new solver, Orchard, designed to improve upon three primary aspects of the current approach. First, Orchard uses structured sets of linear octrees to improve scalability and enable growth to exa-scale simulations. Second, Orchard sub-divides each octant into fixed-dimension Cartesian blocks for cache-efficient algorithms at the block-level and global GMRES algorithm over all blocks. Finally, the framework is designed to run on either CPUs or GPUs to accelerate computation. Vortex convection-diffusion, forward-flight PSP rotor, and forward-flight HART-II rotor cases are presented to analyze accuracy and performance of the new Cartesian solver. On CPUs, Orchard out-performs SAMCart timing for each case. Orchard on GPU-accelerated nodes runs 12.8X faster than Orchard on typical CPU cluster nodes.
Regularizing graph centrality computations
Regularizing graph centrality computations
A Python/Fortran Implementation of the Lattice‐Boltzmann Kernel on Multiple GPU Using the OpenACC Framework
The increasing availability of GPU accelerated architectures for high‐performance computing presents new opportunities for scientific software but also challenges due to the complexity of porting legacy codes to accelerator platforms. Directive‐based programming models such as OpenACC offer a minimally intrusive pathway to exploit GPU acceleration without requiring extensive rewriting of existing codes. The current work presents a comprehensive performance and portability study of a LatticeBoltzmann Method solver (PyLB) originally written in Python, Mpi4Py, and Fortran for CPU architectures, which is ported to GPUs using OpenACC directives applied to the Fortran routines. The performance of the solver is evaluated on NVIDIA V100, A100, and H100 GPUs available on the Jean Zay supercomputer from Institute for Development and Resources in Intensive Scientific Computing (IDRIS) in France. Roofline analysis and extensive strong and weak scalability tests are conducted, showing that the GPU‐enabled version of PyLB scales efficiently across multiple GPUs. The solver achieves performance on the H100 GPU equivalent to thousands of CPU cores and shows strong energy and carbon efficiency advantages over traditional CPU‐based simulations. The implementation is validated using classical benchmarks, including the decaying Taylor‐Green vortex and the flow over a 3‐D sphere. The results confirm the physical accuracy of the GPU port while highlighting its computational and environmental advantages.
Read moreEnabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures
Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.
Read moreConvis: A Toolbox to Fit and Simulate Filter-Based Models of Early Visual Processing
We developed Convis, a Python simulation toolbox for large scale neural populations which offers arbitrary receptive fields by 3D convolutions executed on a graphics card. The resulting software proves to be flexible and easily extensible in Python, while building on the PyTorch library (The Pytorch Project, 2017), which was previously used successfully in deep learning applications, for just-in-time optimization and compilation of the model onto CPU or GPU architectures. An alternative implementation based on Theano (Theano Development Team, 2016) is also available, although not fully supported. Through automatic differentiation, any parameter of a specified model can be optimized to approach a desired output which is a significant improvement over e.g., Monte Carlo or particle optimizations without gradients. We show that a number of models including even complex non-linearities such as contrast gain control and spiking mechanisms can be implemented easily. We show in this paper that we can in particular recreate the simulation results of a popular retina simulation software VirtualRetina (Wohrer and Kornprobst, 2009), with the added benefit of providing (1) arbitrary linear filters instead of the product of Gaussian and exponential filters and (2) optimization routines utilizing the gradients of the model. We demonstrate the utility of 3d convolution filters with a simple direction selective filter. Also we show that it is possible to optimize the input for a certain goal, rather than the parameters, which can aid the design of experiments as well as closed-loop online stimulus generation. Yet, Convis is more than a retina simulator. For instance it can also predict the response of V1 orientation selective cells. Convis is open source under the GPL-3.0 license and available from https://github.com/jahuth/convis/ with documentation at https://jahuth.github.io/convis/.
Read moreA GPU-accelerated pseudo-3D vortex method for aerodynamic analysis
A GPU-accelerated pseudo-3D vortex method for aerodynamic analysis
Versatile regularisation toolkit for iterative image reconstruction with proximal splitting algorithms
Ill-posed image recovery requires regularisation to ensure stability. The presented open-source regularisation toolkit consists of state-of-the-art variational algorithms which can be embedded in a plug-and-play fashion into the general framework of proximal splitting methods. The packaged regularisers aim to satisfy various prior expectations of the investigated objects, e.g., their structural characteristics, smooth or non-smooth surface morphology. The flexibility of the toolkit helps with the design of more advanced model-based iterative reconstruction methods for different imaging modalities while operating with simpler building blocks. The toolkit is written for CPU and GPU architectures and wrapped for Python/MATLAB. We demonstrate the functionality of the toolkit in application to Positron Emission Tomography (PET) and X-ray synchrotron computed tomography (CT).
Read moreHigh-Performance Matrix-Matrix Multiplications of Very Small Matrices
The use of the general dense matrix-matrix multiplication GEMM is fundamental for obtaining high performance in many scientific computing applications. GEMMs for small matrices of sizes less than 32 however, are not sufficiently optimized in existing libraries. In this paper we consider the case of many small GEMMs on either CPU or GPU architectures. This is a case that often occurs in applications like big data analytics, machine learning, high-order FEM, and others. The GEMMs are grouped together in a single batched routine. We present specialized for these cases algorithms and optimization techniques to obtain performance that is within 90i¾ź% of the optimal. We show that these results outperform currently available state-of-the-art implementations and vendor-tuned math libraries.
Read moreKNN-Joins Using a Hybrid Approach
K Nearest Neighbor (KNN) joins are used in many scientific domains for data analysis, and are building blocks of several well-known algorithms. KNN-joins find the KNN of all points in a dataset. However, KNN searches are computationally expensive, and many GPU KNN algorithms focus on the high-dimensional case that plainly gives a performance advantage to the GPU rather than the CPU. Consequently, in this work, we focus on a hybrid CPU/GPU approach for the low-dimensional KNN-join problem. In particular, we utilize a work queue that prioritizes computing data points in high density regions on the GPU, and low density regions on the CPU, thereby taking advantage of each architecture's relative strengths. Our approach, HybridKNN-Join, is shown to effectively augment a state-of-the-art multi-core CPU algorithm. We propose optimizations that (i) maximize GPU query throughput by assigning the GPU larger batches of work than the CPU; (ii) increase workload granularity to optimize GPU resource utilization; and, (iii) limit load imbalance between CPU and GPU architectures. Furthermore, the work queue utilized in our approach shows promise for the general purpose division of work for other hybrid CPU/GPU algorithms.
Read moreAn Algorithm Architecture for Radio Interferometric Data Processing
We present a foundational, scalable algorithm architecture for processing data from aperture synthesis radio telescopes. The analysis leading to the architecture is rooted in the theory of aperture synthesis, signal processing, and numerical optimization, keeping it scalable for variations in computing load, algorithmic complexity, and accommodate the continuing evolution of algorithms. It also adheres to scientific software design principles and use of modern performance engineering techniques providing a stable foundation for long-term scalability, performance, and development cost. We first show that algorithms for both calibration and imaging share a common mathematical foundation and can be expressed as numerical optimization problems. We then decompose the resulting mathematical framework into fundamental conceptual architectural components, and assemble calibration and imaging algorithms from these foundational components. For a physical architectural view, we used a library of algorithms implemented in the LibRA software for the various architectural components, and used the Kokkos framework in the compute-intensive components for performance portable implementation. This was deployed on hardware ranging from desktop-class computers to multiple supercomputer class high-performance computing and high-throughput computing (HTC) platforms with a variety of CPU and GPU architectures, and job schedulers (HTCondor and Slurm). As a test, we imaged archival data from the Very Large Array telescope in the A-array configuration for the Hubble Ultra Deep Field. Using over 100 GPUs, we achieve a processing rate of ∼2 TB hr−1 to make one of the deepest images in the 2–4 GHz band with an rms noise of ∼1 μJy beam−1.
Read moreAutomatic CUDA Code Synthesis Framework for Multicore CPU and GPU Architectures
Recently, general purpose GPU (GPGPU) programming has spread rapidly after CUDA was first introduced to write parallel programs in high-level languages for NVIDIA GPUs. While a GPU exploits data parallelism very effectively, task-level parallelism is exploited as a multi-threaded program on a multicore CPU. For such a heterogeneous platform that consists of a multicore CPU and GPU, we propose an automatic code synthesis framework that takes a process network model specification as input and generates a multithreaded CUDA code. With the model based specification, one can explicitly specify both function-level and loop-level parallelism in an application and explore the wide design space in mapping of function blocks and selecting the communication methods between CPU and GPU. The proposed technique is complementary to other high-level methods of CUDA programming.KeywordsGPGPUCUDAmodel-based designautomatic code synthesis
Read moreMassively parallel skyline computation for processing-in-memory architectures
Processing-In-Memory (PIM) is an increasingly popular architecture aimed at addressing the 'memory wall' crisis by prioritizing the integration of processors within DRAM. It promotes low data access latency, high bandwidth, massive parallelism, and low power consumption. The skyline operator is a known primitive used to identify those multi-dimensional points offering optimal trade-offs within a given dataset. For large multidimensional dataset, calculating the skyline is extensively compute and data intensive. Although, PIM systems present opportunities to mitigate this cost, their execution model relies on all processors operating in isolation with minimal data exchange. This prohibits direct application of known skyline optimizations which are inherently sequential, creating dependencies and large intermediate results that limit the maximum parallelism, throughput, and require an expensive merging phase. In this work, we address these challenges by introducing the first skyline algorithm for PIM architectures, called DSky. It is designed to be massively parallel and throughput efficient by leveraging a novel work assignment strategy that emphasizes load balancing. Our experiments demonstrate that it outperforms the state-of-the-art algorithms for CPUs and GPUs, in most cases. DSky achieves 2× to 14× higher throughput compared to the state-of-the-art solutions on competing CPU and GPU architectures. Furthermore, we showcase DSky's good scaling properties which are intertwined with PIM's ability to allocate resources with minimal added cost. In addition, we showcase an order of magnitude better energy consumption compared to CPUs and GPUs.
Read moreA package for OpenCL based heterogeneous computing on clusters with many GPU devices
Heterogeneous systems provide new opportunities to increase the performance of parallel applications on clusters with CPU and GPU architectures. Currently, applications that utilize GPU devices run their device-executable code on local devices in their respective hosting-nodes. This paper presents a package for running OpenMP, C++ and unmodified OpenCL applications on clusters with many GPU devices. This Many GPUs Package (MGP) includes an implementation of the OpenCL specifications and extensions of the OpenMP API that allow applications on one hosting-node to transparently utilize cluster-wide devices (CPUs and/or GPUs). MGP provides means for reducing the complexity of programming and running parallel applications on clusters, including scheduling based on task dependencies and buffer management. The paper presents MGP and the performance of its internals.
Read moreGeoTaichi: A Taichi-powered high-performance numerical simulator for multiscale geophysical problems
GeoTaichi: A Taichi-powered high-performance numerical simulator for multiscale geophysical problems
Forecast Model for Scheduling an HPC Application on CPU and GPU Architecture
Process scheduling is an essential part of multiprogramming operating systems. Scheduling is a process that allows one process to use the processing unit while the execution of another process is on hold (in a waiting state) due to the unavailability of any resource like I/O, thereby making full use of CPU or GPU. The major issue of scheduling is to make the system efficient, fast, and fair. This work focuses on developing a Forecast model and constructing scheduling strategies to schedule parallel applications on CPU and GPU. During the design of the Forecast model, we will consider the history of the actual execution time set of processes, then we compute the average time of individual sets of processes by considering parameters such as complete execution time, the sum of processes, and the number of threads assigned to individual processes. Then we will evaluate the Prediction time of CPU and GPU for individual sets of processes. By considering parameters such as the average time of the previous set of processes, the weight of processes, and the number of processes. Then based on the prediction time we will develop a scheduling strategy. As the minimum prediction time required set of process resources is assigned to the CPU and the GPU is assigned by the maximum predicted timed resource of the process. In this work we utilized the CPU and GPU resources effectively for stream benchmark application, our experiment shows that less than 20% average percentage prediction error in all cases.
Read moreAutomatic Differentiation of C++ Codes on Emerging Manycore Architectures with Sacado
Automatic differentiation (AD) is a well-known technique for evaluating analytic derivatives of calculations implemented on a computer, with numerous software tools available for incorporating AD technology into complex applications. However, a growing challenge for AD is the efficient differentiation of parallel computations implemented on emerging manycore computing architectures such as multicore CPUs, GPUs, and accelerators as these devices become more pervasive. In this work, we explore forward mode, operator overloading-based differentiation of C++ codes on these architectures using the widely available Sacado AD software package. In particular, we leverage Kokkos, a C++ tool providing APIs for implementing parallel computations that is portable to a wide variety of emerging architectures. We describe the challenges that arise when differentiating code for these architectures using Kokkos, and two approaches for overcoming them that ensure optimal memory access patterns as well as expose additional dimensions of fine-grained parallelism in the derivative calculation. We describe the results of several computational experiments that demonstrate the performance of the approach on a few contemporary CPU and GPU architectures. We then conclude with applications of these techniques to the simulation of discretized systems of partial differential equations.
Read more