- Research Article
41
- 10.1016/s0898-1221(98)00191-6
Efficient approximate solution of sparse linear systems
- Nov 01, 1998
- Computers & Mathematics with Applications
- J.H Reif
Efficient approximate solution of sparse linear systems
SummaryWe propose an adaptive scheme to reduce communication overhead caused by data movement by selectively storing the diagonal blocks of a block‐Jacobi preconditioner in different precision formats (half, single, or double). This specialized preconditioner can then be combined with any Krylov subspace method for the solution of sparse linear systems to perform all arithmetic in double precision. We assess the effects of the adaptive precision preconditioner on the iteration count and data transfer cost of a preconditioned conjugate gradient solver. A preconditioned conjugate gradient method is, in general, a memory bandwidth‐bound algorithm, and therefore its execution time and energy consumption are largely dominated by the costs of accessing the problem's data in memory. Given this observation, we propose a model that quantifies the time and energy savings of our approach based on the assumption that these two costs depend linearly on the bit length of a floating point number. Furthermore, we use a number of test problems from the SuiteSparse matrix collection to estimate the potential benefits of the adaptive block‐Jacobi preconditioning scheme.
Efficient approximate solution of sparse linear systems
Efficient approximate solution of sparse linear systems
Iterative signal restoration by sine transform based preconditioners
A novel sine transform based preconditioned conjugate gradient (PCG) method for signal restoration is presented. If the bandwidth 2? +1 in the model matrix is a constant, then the asymptotic complexity for the new algorithm is O(N) per iterative step, where Nis the order of the model. This complexity is reduced by an order of magnitude in comparison with the known fast Fourier transform (FFT) technique proposed recently in the literature. Moreover, if ? has a size similar to N, then the new PCG method is accomplished by the fast sine transform (FST) and the fast Hartley transform (FHT). The computational and storage cost per preconditioning step for the new PCG method is reduced by 50% as compared to the FFT approach. To stabilize the computation, a special Tikhonov regularization is introduced. Numerical experimentations show that the new PCG method is of a speedy convergence rate. The new PCG methods maintain less computational complexity per iterative step and have a better convergence rate than the other known PCG methods.
Read moreThe Efficient Parallel Iterative Solution of Large Sparse Linear Systems
The development of efficient, general-purpose software for the iterative solution of sparse linear systems on parallel MIMD computers depends on recent results from a wide variety of research areas. Parallel graph heuristics, convergence analysis, and basic linear algebra implementation issues must all be considered.
Read moreA parallel preconditioned conjugate gradient package for solving sparse linear systems on a Cray Y-MP
A parallel preconditioned conjugate gradient package for solving sparse linear systems on a Cray Y-MP
Parallel Sub-structuring Methods for Solving Sparse Linear Systems on a Cluster of GPUs
The main objective of this work consists in analyzing sub-structuring method for the parallel solution of sparse linear systems with matrices arising from the discretization of partial differential equations such as finite element, finite volume and finite difference. With the success encountered by the general-purpose processing on graphics processing units (GPGPU), we develop an hybrid multiGPUs and CPUs sub-structuring algorithm. GPU computing, with CUDA, is used to accelerate the operations performed on each processor. Numerical experiments have been performed on a set of matrices arising from engineering problems. We compare C+MPI implementation on classical CPU cluster with C+MPI+CUDA on a cluster of GPU. The performance comparison shows a speed-up for the sub-structuring method up to 19 times in double precision by using CUDA.
Read moreParallel Implementation of the Coordinates-partitioning Based Aggregation-type Algebraic Multigrid Preconditioners
The coordinates-partitioning based aggregation-type algebraic multigrid preconditioners have been proven to be very efficient in the solution of sparse linear systems with conjugate gradient iterations. In this paper, a parallel algorithm for the setup is provided and the parallelization of the preconditioning process is also considered. The parallel algorithm is based on a good property of the coordinates-partitioning based aggregation, that is, the aggregation process is performed in a fashion from the coarsest level to the finest step by step. Thus, the original adjacent graph is partitioned into a number of sub-graphs, where each sub-graph is related to a node in the coarsest level and is assigned to a processor. The aggregation process can then proceed forward on each processor independently from the assigned sub-graph. When this kind of multigrid preconditioner is applied in Krylov subspace iterations, only the computation on the coarsest level and the matrix-vector multiplications related to the smoothing on each level require communication for V- and W-cycle versions, and require only some extra communication related to dot products for the K-cycle version. The size of the derived linear system on the coarsest level can be controlled by the adjustable arguments and this system can be solved again in parallel with some preconditioned Krylov subspace iterations. The structure information related to the matrix-vector multiplications is invariable to the iterations and can be derived in the setup and be stored, and is used in the latter iterations. Finally, the parallel algorithm is validated for the Gauss-Seidel smoother and some kinds of popular multigrid cycles in solving sparse linear systems from some two-dimensional model partial equations with preconditioned conjugate gradient iterations. The results show that the parallel efficiencies of both the setup and the iteration processes are satisfied.
Read moreA Barotropic Solver for High-Resolution Ocean General Circulation Models
High-resolution global ocean general circulation models (OGCMs) play a key role in accurate ocean forecasting. However, the models of the operational forecasting systems are still not in high resolution due to the subsequent high demand for large computation, as well as the low parallel efficiency barrier. Good scalability is an important index of parallel efficiency and is still a challenge for OGCMs. We found that the communication cost in a barotropic solver, namely, the preconditioned conjugate gradient (PCG) method, is the key bottleneck for scalability due to the high frequency of the global reductions. In this work, we developed a new algorithm—a communication-avoiding Krylov subspace method with a PCG (CA-PCG)—to improve scalability and then applied it to the Nucleus for European Modelling of the Ocean (NEMO) as an example. For PCG, inner product operations with global communication were needed in every iteration, while for CA-PCG, inner product operations were only needed every eight iterations. Therefore, the global communication cost decreased from more than 94.5% of the total execution time with PCG to less than 63.4% with CA-PCG. As a result, the execution time of the barotropic modes decreased from more than 17,000 s with PCG to less than 6000 s with CA-PCG, and the total execution time decreased from more than 18,000 s with PCG to less than 6200 s with CA-PCG. Besides, the ratio of the speedup can also be increased from 3.7 to 4.6. In summary, the high process count scalability when using CA-PCG was effectively improved from that using the PCG method, providing a highly effective solution for accurate ocean simulation.
Read moreProcess scheduling in DSC and the large sparse linear systems challenge
New features of our DSC system for distributing a symbolic computation task over a network of processors are described. A new scheduler sends parallel subtasks to those compute nodes that are best suited in handling the added load of CPU usage and memory. Furthermore, a subtask can communicate back to the process that spawned it by a co-routine style calling mechanism. Two large experiments are described in this improved setting. We have implemented an algorithm that can prove a number of more than 1,000 decimal digits prime in about 2 months elapsed time on some 20 computers. A parallel version of a sparse linear system solver is used to compute the solution of sparse linear systems over finite fields. We are able to find the solution of a 100,000 by 100,000 linear system with about 10.3 million non-zero entries over the Galois field with 2 elements using 3 computers in about 54 hours CPU time.
Read moreAnalysis of Coppersmith’s block Wiedemann algorithm for the parallel solution of sparse linear systems
By using projections by a block of vectors in place of a single vector it is possible to parallelize the outer loop of iterative methods for solving sparse linear systems. We analyze such a scheme proposed by Coppersmith for Wiedemann’s coordinate recurrence algorithm, which is based in part on the Krylov subspace approach. We prove that by use of certain randomizations on the input system the parallel speed up is roughly by the number of vectors in the blocks when using as many processors. Our analysis is valid for fields of entries that have sufficiently large cardinality. Our analysis also deals with an arising subproblem of solving a singular block Toeplitz system by use of the theory of Toeplitz-like matrices.
Read moreComparisons of two implementations for the solution of sparse linear systems - Part II
Comparisons of two implementations for the solution of sparse linear systems - Part II
Performance-Based Numerical Solver Selection in the Lighthouse Framework
Scientific and engineering computing rely heavily on linear algebra for large-scale data analysis, modeling and simulation, machine learning, and other applied problems. Sparse linear system solution often dominates the execution time of such applications, prompting the ongoing development of highly optimized iterative algorithms and high-performance parallel implementations. In the Lighthouse project, we enable application developers with varied backgrounds to readily discover and effectively apply the best available numerical software for their problems, aiming to maximize both developer productivity and application performance. Lighthouse is a search-based expert system built on a software taxonomy that combines expert knowledge, machine learning--based classification of existing numerical software collections, and automated code generation and optimization. In this paper we present the integration of PETSc and Trilinos iterative solvers for sparse linear systems into the Lighthouse framework. In addition to functional information in the taxonomy, we have created a comprehensive machine learning--based workflow for the automated classification of sparse solvers, which can be generalized to other types of rapidly evolving numerical methods. We present a comparative analysis of the solver classification results for a varied set of input problems and machine learning methods, achieving up to 93% accuracy in identifying the best-performing linear solution methods in PETSc and Trilinos.
Read morePreconditioned Legendre Spectral Galerkin Methods for the Non-separable Elliptic Equation
The Legendre spectral Galerkin method of self-adjoint second order elliptic equations usually results in a linear system with a dense and ill-conditioned coefficient matrix. In this paper, the linear system is solved by a preconditioned conjugate gradient (PCG) method where the preconditioner M is constructed by approximating the variable coefficients with a (T+1)-term Legendre series in each direction to a desired accuracy. A feature of the proposed PCG method is that the iteration step increases slightly with the size of the resulting matrix when reaching a certain approximation accuracy. The efficiency of the method lies in that the system with the preconditioner M is approximately solved by a one-step method based on the incomplete LU factorization technique with no fill-in, denoted by ILU(0). The ILU(0) factorization of \(M\in {\mathbb {R}}^{(N-1)^d\times (N-1)^d}\) can be computed using \({\mathcal {O}}(T^{2d} N^d)\) operations, and the number of nonzeros in the factorization factors is of \({\mathcal {O}}(T^{d} N^d)\), \(d=1,2,3\). A conclusion of the algorithm is to fast solve the resulting system from the Legendre Galerkin spectral method for Poisson equations with Dirichlet boundary conditions, which has a complexity of \({\mathcal {O}}(N^d)\). To further speed up the PCG method, an algorithm is developed for fast matrix-vector multiplications by the resulting matrix of Legendre-Galerkin spectral discretization, without the need to explicitly form it. The complexity of the fast matrix-vector multiplications is of \({\mathcal {O}}(N^d (\log _2 N)^{2})\). In view that T is independent of N in one dimension and is set to be of order \({\mathcal {O}}(\log _2 N)\) in two and three dimensions, the PCG method has a \({\mathcal {O}}(N^d (\log _2 N)^{2d})\) quasi-optimal complexity for a d dimensional domain with \((N-1)^d\) unknows, \(d=1,2,3\). In addition, a fast direct solver for the three-dimensional Poisson equation is developed, which is of \({\mathcal {O}}(N^{3} (\log _2 N)^{2})\) and improves the existing results on the computational complexity. Numerical examples are given to demonstrate the efficiency of proposed preconditioners and the algorithm for fast matrix-vector multiplications.
Read moreAdaptive solution of linear systems of equations based on a posteriori error estimators
In this paper, we discuss a new adaptive approach for iterative solution of sparse linear systems arising from partial differential equations (PDEs) with self-adjoint operators. The idea is to use the a posteriori estimated local distribution of the algebraic error in order to steer and guide the solve process in such way that the algebraic error is reduced more efficiently in the consecutive iterations. We first explain the motivation behind the proposed procedure and show that it can be equivalently formulated as constructing a special combination of preconditioner and initial guess for the original system. We present several numerical experiments in order to identify when the adaptive procedure can be of practical use.
Read moreAnalysis and optimization of power consumption in the iterative solution of sparse linear systems on multi-core and many-core platforms
Energy efficiency is a major concern in modern high-performance-computing. Still, few studies provide a deep insight into the power consumption of scientific applications. Especially for algorithms running on hybrid platforms equipped with hardware accelerators, like graphics processors, a detailed energy analysis is essential to identify the most costly parts, and to evaluate possible improvement strategies. In this paper we analyze the computational and power performance of iterative linear solvers applied to sparse systems arising in several scientific applications. We also study the gains yield by dynamic voltage/frequency scaling (DVFS), and illustrate that this technique alone cannot to reduce the energy cost to a considerable amount for iterative linear solvers. We then apply techniques that set the (multi-core processor in the) host system to a low-consuming state for the time that the GPU is executing. Our experiments conclusively reveal how the combination of these two techniques deliver a notable reduction of energy consumption without a noticeable impact on computational performance.
Read moreComparison of Iterative Methods for the Solution of Compressible-Fluid Reynolds Equation
This study presents an efficacy comparison of iterative solution methods for solving the compressible-fluid Reynolds equation in modeling air- or gas-lubricated bearings. A direct fixed-point iterative (DFI) method and Newton’s method are employed to transform the Reynolds equation in a form that can be solved iteratively. The iterative solution methods examined are the Gauss–Seidel method, the successive over-relaxation (SOR) method, the preconditioned conjugate gradient (PCG) method, and the multigrid method. The overall solution time is affected by both the transformation method and the iterative method applied. In this study, Newton’s method shows its effectiveness over the straightforward DFI method when the same iterative method is used. It is demonstrated that the use of an optimal relaxation factor is of vital importance for the efficiency of the SOR method. The multigrid method is an order faster than the PCG and optimal SOR methods. Also, the multigrid and PCG methods involve an extended coding work and are less flexible in dealing with gridwork and boundary conditions. Consequently, a compromise has to be made in terms of ease of use as well as programming effort for the solution of the compressible-fluid Reynolds equation.
Read more