- Research Article
41
- 10.1016/j.neucom.2009.04.006
Self-scaled conjugate gradient training algorithms
- May 10, 2009
- Neurocomputing
- A.E Kostopoulos + 1 more +1
Self-scaled conjugate gradient training algorithms
The left-preconditioned communication avoiding conjugate gradient (LP-CA-CG) method is applied to the pressure Poisson equation in the multiphase CFD code JUPITER. The arithmetic intensity of the LP-CA-CG method is analyzed, and is dramatically improved by loop splitting for inner product operations and for three term recurrence operations. Two LPCA-CG solvers with block Jacobi preconditioning and with underlap preconditioning are developed. The former is developed based on a hybrid CA approach, in which CA is applied only to global collective communications for inner product operations. The latter is a full CA approach, in which CA is applied also to local point-to-point communications in sparse matrix-vector (SpMV) operations and preconditioning. CA-SpMV requires additional computation for overlapping regions. CA-preconditiong is enabled by underlap preconditioning, which approximates preconditioning for overlapping regions by point Jacobi preconditioning. It is shown that on the K computer, the former is faster, because the performance of local point-to-point communications scales well, and the convergence property becomes worse with underlap preconditioning. The LP-CA-CG solver shows good strong scaling up to 30,000 nodes, where the LP-CA-CG solver achieved higher performance than the original CG solver by reducing the cost of global collective communications by 69 percent.
Self-scaled conjugate gradient training algorithms
Self-scaled conjugate gradient training algorithms
Preconditioning of variational data assimilation and the use of a bi‐conjugate gradient method
Presently, a preferred minimization for strong‐constraint four‐dimensional variational (4D‐Var) assimilation uses a Lanczos‐based conjugate gradient (CG) algorithm. This requires the availability of a square‐root of the background‐error covariance matrix (B). In the context of weak‐constraint 4D‐Var, this requirement might be too restrictive for the formulations of the model error term. It might therefore be desirable to avoid a square‐root decomposition of the augmented background term. An appealing minimization scheme is the double CG minimization employed, for example, in the grid‐point statistical interpolation (GSI) analysis. Realizing the double CG algorithm is a special case of the more general bi‐conjugate gradient (BiCG) method for solving non‐symmetric problems, the present work introduces a Lanczos‐based preconditioning strategy when B, instead of its square‐root, is used initially. Implementation of the scheme is done in the context of the GSI analysis system, and preliminary experiments are presented using its 3D‐Var version. Comparison of the Lanczos‐based CG and the BiCG shows that the algorithms converge at the same rate and to the same solution. Despite the additional computational cost, the importance of the re‐orthogonalization step is also shown to be fundamental to any of these CG algorithms. Furthermore, when using the Hessian eigenvectors for preconditioning, the BiCG behaviour is shown to be comparable to that of the Lanczos‐CG algorithm. Both schemes construct the same approximation of the Hessian with the same number of eigenvectors, and benefit in the same way from the reduction of the condition number. The efficiency, computational cost, and stability of the three algorithms are discussed. Copyright © 2012 Royal Meteorological Society
Read moreApplication of a Preconditioned Chebyshev Basis Communication-Avoiding Conjugate Gradient Method to a Multiphase Thermal-Hydraulic CFD Code
A preconditioned Chebyshev basis communication-avoiding conjugate gradient method (P-CBCG) is applied to the pressure Poisson equation in a multiphase thermal-hydraulic CFD code JUPITER, and its computational performance and convergence properties are compared against a preconditioned conjugate gradient (P-CG) method and a preconditioned communication-avoiding conjugate gradient (P-CACG) method on the Oakforest-PACS, which consists of 8,208 KNLs. The P-CBCG method reduces the number of collective communications with keeping the robustness of convergence properties. Compared with the P-CACG method, an order of magnitude larger communication-avoiding steps are enabled by the improved robustness. It is shown that the P-CBCG method is 1.38 $$\times $$ and 1.17 $$\times $$ faster than the P-CG and P-CACG methods at 2,000 processors, respectively.
Read moreSolving Unconstrained Optimization Problem Using Hybrid CG Method with Exact Line Search
Abstract: The conjugate gradient (CG) method is a one of the common approaches for solving unconstrained optimization problems, notably known for its suitability for large scale problems. Many recent studies show that this method is also useful for problems of smaller scale. One of the methods used for improving the performance of CG method is hybrid approach, where a CG method is combined with another method. In this study, the ARM CG method is combined with the SMR CG method and tested under exact line search. The resulting hybrid algorithm is globally convergent under exact line search and shown to perform well numerically in comparison to other tested CG methods.
Read moreLocal heat flux estimation inside tubes through conjugate gradient method with adjoint operator: application to the pulsating heat pipes case
PurposeThe purpose of this paper is to apply the conjugate gradient (CG) method, together with the adjoint operator (AO) to the pulsating heat pipe problem, including some quite interesting experimental results. The CG method, together with the AO, was able to estimate the unknown functions more efficiently than the other techniques presented in this paper. The estimation of local heat transfer coefficients, rather than the global ones, in pulsating heat pipes is a relatively new subject and presenting a robust, efficient and self-regularized inverse tool to estimate it, supported also by some experimental results, is the main purpose of this paper. To also increase the visibility and the general use of the paper to the heat transfer community, the authors include, as supplemental material, all numerical and experimental data used in this paper.Design/methodology/approachThe approach was established on the solution of the inverse heat conduction problem in the wall by using as starting data the temperature measurements on the outer surface. The procedure is based on the CG method with AO. The here proposed approach was first verified adopting synthetic data and then it was validated with real cases regarding pulsating heat pipes.FindingsAn original fast methodology to estimate local convective heat flux is proposed. The procedure has been validated both numerically and experimentally. The procedure has been compared to other classical methods presenting some peculiar benefits.Practical implicationsThe approach is suitable for pulsating heat pipes performance evaluation because these devices present a local heat flux distribution characterized by an important variation both in time and in space as a result of the complex flow patterns that are generated in this type of devices.Originality/valueThe procedure here proposed shows these benefits: it affords a general model of the heat conduction problem that is effortlessly customized for the particular case, it can be applied also to large datasets and it presents reduced computational expense.
Read moreTwo Modified Hager and Zhang's Conjugate Gradient Algorithms For Solving Large-Scale Optimization Problems
At present, the conjugate gradient (CG) method of Hager and Zhang (Hager and Zhang, SIAM Journal on Optimization, 16(2005)) is regarded as one of the most effective CG methods for optimization problems. In order to further study the CG method, we develop the Hager and Zhang's CG method and present two modified CG formulas, where the given formulas possess the value information of not only the gradient but also the function. Moreover, the sufficient descent condition will be holden without any line search. The global convergence is established for nonconvex function under suitable conditions. Numerical results show that the proposed methods are competitive to the normal conjugate gradient method.
Read moreBilinear Optimal Control of an Advection-Reaction-Diffusion System
We consider the bilinear optimal control of an advection-reaction-diffusion system, where the control arises as the velocity field in the advection term. Such a problem is generally challenging from both theoretical analysis and algorithmic design perspectives, mainly because the state variable depends nonlinearly on the control variable and, an additional divergence-free constraint on the control is coupled together with the state equation. Mathematically, the proof of the existence of optimal solutions is delicate, and, up to now, only some results have been known for a few special cases where additional restrictions are imposed on the space dimension and the regularity of the control. We prove the existence of optimal controls and derive the first-order optimality conditions in general settings without any extra assumptions. Computationally, the well-known conjugate gradient (CG) method can be applied conceptually. However, due to the additional divergence-free constraint on the control variable and the nonlinear relation between the state and control variables, it is challenging to compute the gradient and the optimal stepsize at each CG iteration, and thus nontrivial to implement the CG method. To address these issues, we advocate a fast inner preconditioned CG method to ensure the divergence-free constraint and an efficient inexactness strategy to determine an appropriate stepsize. An easily implementable nested CG method is thus proposed for solving such a complicated problem. For the numerical discretization, we combine finite difference methods for the time discretization and finite element methods for the space discretization. Efficiency of the proposed nested CG method is promisingly validated by the results of some preliminary numerical experiments.
Read moreNeural network model of heat and fluid flow in gas metal arc fillet welding based on genetic algorithm and conjugate gradient optimisation
Although numerical calculations of heat transfer and fluid flow can provide detailed insights into welding processes and welded materials, these calculations are complex and unsuitable in situations where rapid calculations are needed. A recourse is to train and validate a neural network, using results from a well tested heat and fluid flow model to significantly expedite calculations and ensure that the computed results conform to the basic laws of conservation of mass, momentum and energy. Seven feedforward neural networks were developed for gas metal arc (GMA) fillet welding, one each for predicting penetration, leg length, throat, weld pool length, cooling time between 800°C and 500°C, maximum velocity and peak temperature in the weld pool. Each model considered 22 inputs that included all the welding variables, such as current, voltage, welding speed, wire radius, wire feed rate, arc efficiency, arc radius, power distribution, and material properties such as thermal conductivity, specific heat and temperature coefficient of surface tension. The weights in the neural network models were calculated using the conjugate gradient (CG) method and by a hybrid optimisation scheme involving the CG method and a genetic algorithm (GA). The neural network produced by the hybrid optimisation model produced better results than the networks based on the CG method with various sets of randomised initial weights. The CG method alone was unable to find the best optimal weights for achieving low errors. The hybrid optimisation scheme helped in finding optimal weights through a global search, as evidenced by good agreement between all the outputs from the neural networks and the corresponding results from the heat and fluid flow model.
Read moreA three-term conjugate gradient descent method with some applications
The stationary point of optimization problems can be obtained via conjugate gradient (CG) methods without the second derivative. Many researchers have used this method to solve applications in various fields, such as neural networks and image restoration. In this study, we construct a three-term CG method that fulfills convergence analysis and a descent property. Next, in the second term, we employ a Hestenses-Stiefel CG formula with some restrictions to be positive. The third term includes a negative gradient used as a search direction multiplied by an accelerating expression. We also provide some numerical results collected using a strong Wolfe line search with different sigma values over 166 optimization functions from the CUTEr library. The result shows the proposed approach is far more efficient than alternative prevalent CG methods regarding central processing unit (CPU) time, number of iterations, number of function evaluations, and gradient evaluations. Moreover, we present some applications for the proposed three-term search direction in image restoration, and we compare the results with well-known CG methods with respect to the number of iterations, CPU time, as well as root-mean-square error (RMSE). Finally, we present three applications in regression analysis, image restoration, and electrical engineering.
Read moreRestrictively Preconditioned Conjugate Gradient Method for a Series of Constantly Augmented Least Squares Problems
In this study, we analyze the real-time solution of a series of augmented least squares problems, which are generated by adding information to an original least squares model repetitively. Instead of solving the least squares problems directly, we transform them into a batch of saddle point linear systems and subsequently solve the linear systems using restrictively preconditioned conjugate gradient (RPCG) methods. Approximation of the new Schur complement is generated effectively based on a previously approximated Schur complement. Owing to the variations of the preconditioned conjugate gradient method, the proposed methods generate convergence results similar to the conjugate gradient method and achieve a very fast convergent iterative sequence when the coefficient matrix is well preconditioned. Numerical tests show that the new methods are more effective than some standard Krylov subspace methods. Updated RPCG methods meet the requirement of real-time computing successfully for multifactor models.
Read moreA new conjugate gradient method with exact line search
Conjugate gradient (CG) methods have been practically used to solve large-scale unconstrained optimization problems due to their simplicity and low memory storage. In this paper, we proposed a new type of CG coefficients
Read morePolarity for quadratic hypersurfaces and conjugate gradient method: Relation between degenerate and nondegenerate cases
In this paper we consider a geometric viewpoint to analyze the behaviour of the Conjugate Gradient (CG) method, for the solution of a symmetric linear system, when at current step a pivot breakdown possibly occurs (degenerate case). As well known this can occur when the system matrix is indefinite or singular. In the latter case the CG gets stuck, since the steplength along the current search direction cannot be computed. We show here that a simple geometric interpretation can be provided for the degenerate case, as long as some basics on projective geometry in the Euclidean space are considered.
Read moreA Hybrid Conjugate Gradient Method with Trust Region for Large-Scale Unconstrained Optimization Problems
In this work, we modify a conjugate gradient (CG) method recently proposed in the literature, where a PRP conjugate gradient method is modified using trust region. Particularly, we propose a hybrid CG method that incorporates the parameters $β^{PRP},$ $β^{FR}$ and $β^{CD},$ and this new search direction satisfies both the trust region feature and the sufficient descent conditions. Furthermore, under suitable conditions the developed method is proved to be globally convergent. The method is tested on some benchmark problems from the literature and numerical results show that it is quite efficient in solving large scale problems.
Read moreConjugate Gradient and L-curve like methods for large inverse problem
Conjugate Gradient and L-curve like methods for large inverse problem
Three-dimensional time-harmonic eddy current problems solved by the geometric multigrid preconditioned conjugate gradient method
The focus of this study is on the efficient solution for three-dimensional (3D) time-harmonic eddy current problems discretised by the finite-element method (FEM) with thin elements. The systems of equations arising from the finite-element formulations are solved by the geometric multigrid preconditioned conjugate gradient (MGCG) method. Numerical examples show that the MGCG method has stable convergence, which does not deteriorate for elements with high aspect ratios and its computation time is much less than that of the conjugate gradient method with incomplete Cholesky factorisation as the preconditioner (ICCG).
Read more