- Research Article
66
- 10.1016/s0045-7825(98)00374-0
Distributed parallel Delaunay mesh generation
- Jul 01, 1999
- Computer Methods in Applied Mechanics and Engineering
- R Said + 3 more +3
Distributed parallel Delaunay mesh generation
Scalable 3D Hybrid Parallel Delaunay Image-to-Mesh Conversion Algorithm for Distributed Shared Memory Architectures
Distributed parallel Delaunay mesh generation
Distributed parallel Delaunay mesh generation
Algorithm 995
A bottom-up approach to parallel anisotropic mesh generation is presented by building a mesh generator starting from the basic operations of vertex insertion and Delaunay triangles. Applications focusing on high-lift design or dynamic stall, or numerical methods and modeling test cases, still focus on two-dimensional domains. This automated parallel mesh generation approach can generate high-fidelity unstructured meshes with anisotropic boundary layers for use in the computational fluid dynamics field. The anisotropy requirement adds a level of complexity to a parallel meshing algorithm by making computation depend on the local alignment of elements, which in turn is dictated by geometric boundaries and the density functions— one-dimensional spacing functions generated from an exponential distribution. This approach yields computational savings in mesh generation and flow solution through well-shaped anisotropic triangles instead of isotropic triangles. The validity of the meshes is shown through solution characteristic comparisons to verified reference solutions. A 79% parallel weak scaling efficiency on 1,024 distributed memory nodes, and a 72% parallel efficiency over the fastest sequential isotropic mesh generator on 512 distributed memory nodes, is shown through numerical experiments.
Read morePerformance analysis of parallel hash join algorithms on a distributed shared memory machine implementation and evaluation on HP exemplar SPP 1600
The distributed shared memory (DSM) architecture is considered to be one of the most likely parallel computing environment candidate for the near future because of its ease of system scalability and facilitation for parallel programming. However, a naive program based on shared memory execution on a DSM machine often deteriorates performance, because of the overhead involved for maintaining cache coherency particularly with frequent remote memory accesses. We show that careful buffer management of parallel join processing on DSM can produce considerable performance improvements in comparison with a naive implementation. We propose four buffer management strategies for parallel hash join processing on the DSM architecture and actually implement them on the HP Exemplar SPP 1600. The basic strategy is to begin with the hash join algorithm for the shared everything architecture and then to consider the memory locality of DSM by distributing the hash table and data pool buffers among the nodes. The results of four buffering strategies are analyzed in detail. Consequently, we can conclude that, in order to achieve high performance on a DSM machine, our buffer management strategy in which the memory access pattern is extracted and buffers are allocated in the local memory of nodes to minimize memory access cost is very efficient.
Read moreA template for developing next generation parallel Delaunay refinement methods
A template for developing next generation parallel Delaunay refinement methods
Parallel Two-Dimensional Unstructured Anisotropic Delaunay Mesh Generation for Aerospace Applications
A bottom-up approach to parallel anisotropic mesh generation is presented by building a mesh generator from the principles of point-insertion, triangulation, and Delaunay refinement. Applications focusing on high-lift design or dynamic stall, or numerical methods and modeling test cases use two-dimensional domains. Our push-button parallel mesh generation approach, meaning the user only needs to start the program by specifying the initial geometry, anisotropic gradation, and ray angle constraint, can generate high-fidelity unstructured meshes with anisotropic boundary layers for use in the computational fluid dynamics field. © 2015 The Authors. Published by Elsevier Ltd. Peer-review under responsibility of organizing committee of the 24th International Meshing Roundtable (IMR24).
Read moreTowards implementation of a novel scheme for data prefetching on distributed shared memory systems
High speed networks and rapidly improving microprocessor performance make the network of workstations an extremely important tool for parallel computing in order to speedup the execution of scientific applications. Shared memory is an attractive programming model for designing parallel and distributed applications, where the programmer can focus on algorithmic development rather than data partition and communication. Based on this important characteristic, the design of systems to provide the shared memory abstraction on physically distributed memory machines has been developed, known as Distributed Shared Memory (DSM). DSM is built using specific software to combine a number of computer hardware resources into one computing environment. Such an environment not only provides an easy way to execute parallel applications, but also combines available computational resources with the purpose of speeding up execution of these applications. DSM systems need to maintain data consistency in memory, which usually leads to communication overhead. Therefore, there exists a number of strategies that can be used to overcome this overhead issue and improve overall performance. Strategies as prefetching have been proven to show great performance in DSM systems, since they can reduce data access communication latencies from remote nodes. On the other hand, these strategies also transfer unnecessary prefetching pages to remote nodes. In this research paper, we focus on the access pattern during execution of a parallel application, and then analyze the data type and behavior of parallel applications. We propose an adaptive data classification scheme to improve prefetching strategy with the goal to improve overall performance. Adaptive data classification scheme classifies data according to the accessing sequence of pages, so that the home node uses past history access patterns of remote nodes to decide whether it needs to transfer related pages to remote nodes. From experimental results, we can observe that our proposed method can increase the accuracy of data access in effective prefetch strategy by reducing the number of page faults and misprefetching. Experimental results using our proposed classification scheme show a performance improvement of about 9---25% over the same benchmark applications running on top of an original JIAJIA DSM system.
Read moreA Parallel Approach for the Generation of Unstructured Meshes with Billions of Elements on Distributed-Memory Supercomputers
This paper describes a parallel approach for the rapid generation of ultra-large-scale unstructured meshes on distributed-memory supercomputers. A medium-sized initial mesh is prepared first. Afterwards, a two-level domain decomposition (DD) strategy is used to split and distribute the initial mesh to different cores. Finally, the parallel mesh generation, comprising a recursive procedure which includes parallel surface recovery, parallel boundary updating, and parallel mesh multiplication, is performed. The two-level DD differentiates the intra-node and inter-node communication to reduce communication overheads. A global indexing and updating scheme is used to make the mesh multiplication devoid of communication. A new parallel surface recovery algorithm without communication is developed to maintain the fidelity of the resulting mesh model to the original geometric model. Tests of the parallel approach for some real-life problems on supercomputers (Dawning-5000A and Tianhe-2) are presented. Issues regarding the speedup, parallel efficiency, and mesh quality are discussed. Results show that the proposed parallel approach has a reasonably good scalability, that the quality of the resulting mesh is improved, and that ultra-large-scale meshes with billions of elements can be generated quickly.
Read moreOpenMP‐oriented applications for distributed shared memory architectures
The rapid rise of OpenMP as the preferred parallel programming paradigm for small‐to‐medium scale parallelism could slow unless OpenMP can show capabilities for becoming the model‐of‐choice for large scale high‐performance parallel computing in the coming decade.The main stumbling block for the adaptation of OpenMP to distributed shared memory (DSM) machines, which are based on architectures like cc‐NUMA, stems from the lack of capabilities for data placement among processors and threads for achieving data locality. The absence of such a mechanism causes remote memory accesses and inefficient cache memory use, both of which lead to poor performance.This paper presents a simple software programming approach called copy‐inside–copy‐back (CC) that exploits the data privatization mechanism of OpenMP for data placement and replacement. This technique enables one to distribute data manually without taking away control and flexibility from the programmer and is thus an alternative to the automat and implicit approaches. Moreover, the CC approach improves on the OpenMP‐SPMD style of programming that makes the development process of an OpenMP application more structured and simpler.The CC technique was tested and analyzed using the NAS Parallel Benchmarks on SGI Origin 2000 multiprocessor machines. This study shows that OpenMP improves performance of coarse‐grained parallelism, although a fast copy mechanism is essential. Copyright © 2004 John Wiley & Sons, Ltd.
Read moreParallel adaptive tetrahedral mesh generation by the advancing front technique
Parallel adaptive tetrahedral mesh generation by the advancing front technique
Massively parallel mesh generation for physics codes
Massively parallel processors (MPPs) will soon enable realistic 3-D physical modeling of complex objects and systems. Work is planned or presently underway to port many of LLNL`s physical modeling codes to MPPs. LLNL`s DSI3D electromagnetics code already can solve 40+ million zone problems on the 256 processor Meiko. However, the author lacks the software necessary to generate and manipulate the large meshes needed to model many complicated 3-D geometries. State-of-the-art commercial mesh generators run on workstations and have a practical limit of several hundred thousand elements. In the foreseeable future MPPs will solve problems with a billion mesh elements. The objective of the Parallel Mesh Generation (PMESH) Project is to develop a unique mesh generation system that can construct large 3-D meshes (up to a billion elements) on MPPs. Such a capability will remove a critical roadblock to unleashing the power of MPPs for physical analysis and will put LLNL at the forefront of mesh generation technology. PMESH will ``front-end`` a variety of LLNL 3-D physics codes, including those in the areas of electromagnetics, structural mechanics, thermal analysis, and hydrodynamics. The DSI3D and DYNA3D codes are already running on MPPs. The primary goal of the PMESH project is to provide the robust generation of large meshes for complicated 3-D geometries through the appropriate distribution of the generation task between the user`s workstation and the MPP. Secondary goals are to support the unique features of LLNL physics codes (e.g., unusual elements) and to minimize the user effort required to generate different meshes for the same geometry. PMESH`s capabilities are essential because mesh generation is presently a major limiting factor in simulating larger and more complex 3-D geometries. PMESH will significantly enhance LLNL`s capabilities in physical simulation by advancing the state-of-the-art in large mesh generation by 2 to 3 orders of magnitude.
Read moreAUTOMESH‐2D/3D: robust automatic mesh generator for metal forming simulation
Meshing and remeshing are often needed to complete a metal forming simulation with a finite element method programme. In order to automate the simulation process and improve the simulation accuracy, an automatic mesh generator, AUTOMESH‐2D/3D, is developed in this paper. The generator can adaptively produce quadrilateral, tetrahedral and hexahedral element meshes based on geometry curvature, thickness, interpolation errors, etc. An improved looping method is used for two‐dimensional quadrilateral and three‐dimensional triangle mesh generation. The tetrahedral and hexahedral element meshes are generated based on the advancing front method and improved grid based method respectively. The mesh generation processes and key technologies are presented in this paper. Many examples of mesh generation with AUTOMESH‐2D/3D are given to demonstrate the robustness of the mesh generator.
Read moreA 3-D Unstructured Multigrid Navier-Stokes Solver
Unstructured mesh methods have proven to be a useful practical tool for aeronautical design. Currently, the process of mesh generation and computation of inviscid flows past complex configurations such as a full aircraft can be completed in a matter of days. The meshes used for inviscid computations consist, in general, of tetrahedral elements of aspect ratio close to unity [1]. The use of such meshes for the simulation of high Reynolds number flows is computationally unfeasible due to the very small mesh sizes required in the near-wall regions where viscous layers are present. A compromise between economy and accuracy can be achieved by the use of elements that are stretched along the flow direction in those regions. Current mesh generation methods can be used to produce these stretched elements ([2],[4]) but the quality of the generated elements is very difficult to control. An alternative method is the use of an hybrid approach in which the mesh near the solid surfaces is generated by an hyperbolic type mesh generator whilst the region outside the viscous layer is discretized using the conventional approach [5].
Read moreParallel and distributed adaptive quadrilateral mesh generation
Parallel and distributed adaptive quadrilateral mesh generation
A case study of optimistic computing on the grid: parallel mesh generation
This paper describes our progress in creating a case study on optimistic computing for the Grid using parallel mesh generation. For the implementation of both methods we use a portable runtime environment for mobile applications (PREMA) which is extended to provide support for optimistic control using grid performance monitoring and prediction. Based on the observed performance of a world-wide grid testbed, we use this case study to develop a methodology for estimating target operating regions for grid applications. The goal of this project is to generalize the experience and knowledge of optimistic grid computing gained through mesh generation into a tool that can be applied to tightly coupled computations in other application domains.
Read moreParallel generalized delaunay mesh refinement
The modeling of physical phenomena in computational fracture mechanics, computational fluid dynamics and other fields is based on solving systems of partial differential equations (PDEs). When PDEs are defined over geometrically complex domains, they often do not admit closed form solutions. In such cases, they are solved approximately using discretizations of domains into simple elements like triangles and quadrilaterals in two dimensions (2D), and tetrahedra and hexahedra in three dimensions (3D). These discretizations are called finite element meshes. Many applications, for example, real-time computer assisted surgery, or crack propagation from fracture mechanics, impose time and/or mesh size constraints that cannot be met on a single sequential machine. As a result, the development of parallel mesh generation algorithms is required. In this dissertation, we describe a complete solution for both sequential and parallel construction of guaranteed quality Delaunay meshes for 2D and 3D geometries. First, we generalize the existing 2D and 3D Delaunay refinement algorithms along with theoretical proofs of mesh quality in terms of element shape and mesh gradation. Existing algorithms are constrained by just one or two specific positions for the insertion of a Steiner point inside a circumscribed disk of a poorly shaped element. We derive an entire 2D or 3D region for the selection of a Steiner point (i.e., infinitely many choices) inside the circumscribed disk. Second, we develop a novel theory which extends both the 2D and the 3D Generalized Delaunay Refinement methods for the concurrent and mathematically guaranteed independent insertion of Steiner points. Previous parallel algorithms are either reactive relying on implementation heuristics to resolve dependencies in parallel mesh generation computations or require the solution of a very difficult geometric optimization problem (the domain decomposition problem) which is still open for general 3D geometries. Our theory solves both of these drawbacks. Third, using our generalization of both the sequential and the parallel algorithms we implemented prototypes of practical and efficient parallel generalized guaranteed quality Delaunay refinement codes for both 2D and 3D geometries using existing state-of-the-art sequential codes for traditional Delaunay refinement methods. On a heterogeneous cluster of more than 100 processors our implementation can generate a uniform mesh with about a billion elements in less than 5 minutes. Even on a workstation with a few cores, we achieve a significant performance improvement over the corresponding state-of-the-art sequential 3D code, for graded meshes.
Read more