SEARCH · Search NASA
Results for “ITERATIVE METHODS”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization
Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.
An iterative method to deblend AGN-Host contributions for Integral Field spectroscopic observations
ABSTRACT We present a new iterative deblending method to separate the host galaxy (HG) and their Active Galactic Nuclei (AGNs) emission with the use of Integral Field spectroscopic (IFS) data. The method decomposes the resolved HG emission from the unresolved AGN emission by modelling the two-dimensional surface brightness (SB) profile of the point-spread function (PSF) and the two-dimensional SB HG continuum simultaneously per each monochromatic slide. Our method does not require any prior information about the observed SB profile or a detailed fitting of the PSF, making it ideal for the automatic analysis of large galaxy samples. In this work, we test the quality of our method, its advantages, and its disadvantages. We test our method by using a set of IFS mock data cubes to quantify the reliability of our deblending process and further compare our method with the qdblend3d analysis tool. Furthermore, we applied our method to three data cubes selected from the MaNGA survey according to the dominance of either its HG or its AGN. We show that our deblending method is capable of disengaging the bright, non-resolved AGN emission from the HG continuum and its narrow emission lines. However, the decoupling depends on how well the IFS spatially resolves the PSF, and on the relative flux intensity of the HG-AGN. Therefore, the method is ideal for disentangling the bright-flux contribution from AGN-dominated spectra.
Robust Iterative Method for Symmetric Quantum Signal Processing in All Parameter Regimes
Here, this paper addresses the problem of solving nonlinear systems in the context of symmetric quantum signal processing (QSP), a powerful technique for implementing matrix functions on quantum computers. Symmetric QSP focuses on representing target polynomials as products of matrices in SU(2) that possess symmetry properties. We present a novel Newton’s method tailored for efficiently solving the nonlinear system involved in determining the phase factors within the symmetric QSP framework. Our method demonstrates rapid and robust convergence in all parameter regimes, including the challenging scenario with ill-conditioned Jacobian matrices, using standard double precision arithmetic operations. For instance, solving symmetric QSP for a highly oscillatory target function α cos(1000x) (polynomial degree ≈ 1433) takes 6 iterations to converge to machine precision when α = 0.9, and the number of iterations only increases to 18 iterations when α = 1 – 10 -9 with a highly ill-conditioned Jacobian matrix. Leveraging the matrix product state structure of symmetric QSP, the computation of the Jacobian matrix incurs a computational cost comparable to a single function evaluation. Moreover, we introduce a reformulation of symmetric QSP using real-number arithmetics, further enhancing the method’s efficiency. Extensive numerical tests validate the effectiveness and robustness of our approach, which has been implemented in the QSPPACK software package.
An efficient iterative method for dynamical Ginzburg-Landau equations
Here we propose a new finite element approach to solving the time-dependent Ginzburg-Landau equations for superconductivity under the temporal gauge and design an efficient preconditioner for the Newton iteration of the resulting discrete system. It solves the vector magnetic potential in $H$(curl) space by the lowest order of the second kind N´ed´elec element. This approach offers a simple way to deal with the boundary conditions, leading to a stable and reliable performance for superconductor geomtry with reentrant corners. The bench-marking in numerical simulations verifies the efficiency of the proposed preconditioner, which can be employed to significantly speed up in large-scale computations.
HyKKT: a hybrid direct-iterative method for solving KKT linear systems
Here, we propose a solution strategy for the large indefinite linear systems arising in interior methods for nonlinear optimization. The method is suitable for implementation on hardware accelerators such as graphical processing units (GPUs). The current gold standard for sparse indefinite systems is the LBLT factorization where L is a lower triangular matrix and B is 1×1 or 2×2 block diagonal. However, this requires pivoting, which substantially increases communication cost and degrades performance on GPUs. Our approach solves a large indefinite system by solving multiple smaller positive definite systems, using an iterative solver on the Schur complement and an inner direct solve (via Cholesky factorization) within each iteration. Cholesky is stable without pivoting, thereby reducing communication and allowing reuse of the symbolic factorization. We demonstrate the practicality of our approach on large optimal power flow problems and show that it can efficiently utilize GPUs and outperform LBL T factorization of the full system.
Simultaneous retrieval of temperature and humidity profiles from ground‐based infrared hyperspectral spectrometer by combined machine learning and physical iterative method
Explore the source record for details and available documents.
An Efficient High-to-Low Iterative Method for Light Water Reactor Analysis Based on NEAMS Tools
Not provided.
Towards Verified Rounding-Error Analysis for Stationary Iterative Methods.
Abstract not provided.
PyAMG: Algebraic Multigrid Solvers in Python
PyAMG is a Python package of algebraic multigrid (AMG) solvers and supporting tools for approximating the solution to large, sparse linear systems of algebraic equations, Ax = b, where A is an n × n sparse matrix. Sparse linear systems arise in a range of problems in science, from fluid flows to solid mechanics to data analysis. While the direct solvers available in SciPy’s sparse linear algebra package (scipy.sparse.linalg) are highly efficient, in many cases iterative methods are preferred due to overall complexity. However, the iterative methods in SciPy, such as CG and GMRES, often require an efficient preconditioner in order to achieve a lower complexity. Preconditioning is a powerful tool whereby the conditioning of the linear system and convergence rate of the iterative method are both dramatically improved. PyAMG constructs multigrid solvers for use as a preconditioner in this setting. A summary of multigrid and algebraic multigrid solvers can be found in Olson (2015a), in Olson (2015b), and in Falgout (2006); a detailed description can be found in Briggs et al. (2000) and Trottenberg et al. (2001).
Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers
The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.
A simulation study of the ability to detect power distribution perturbations in the texas A&M TRIGA reactor with self-powered neutron detectors
Given the variety of ways that nuclear reactor core power may be perturbed, reactor operators and developers are keen on understanding the accuracy and convergence time during which perturbations in reactor power distribution may be synthesized (i.e., inferred) from an array of in-core radiation detectors. A simulation study was conducted as described herein using a highly detailed model of the Texas A&M Training, Research, Isotopes, General Atomics Reactor, in which an array of self-powered neutron detectors (SPNDs) was considered for input to the power synthesis methodology. The core power synthesis is conducted using a point-based iterative method with an iterative loop built in to ensure working equation consistency. The forward problem of SPND response to simulated perturbations in reactor power was solved for Gaussian peak-type perturbations in the reactor power distribution. These perturbations varied in variance, amplitude, and core location to assess their impact on synthesis error and to determine the number of iterations required for convergence. A relation between the unique resolvability limit and perturbation width was identified such that the maximum synthesis error increased rapidly when the peak width went beneath this limit (a width approximating half the reactor’s fuel pin-to-pin pitch); this resolvability limit is specific to the SPND configuration and fuel segmentation considered herein. The synthesis error increased linearly with perturbation peak amplitude, whereas the convergence time increased nonlinearly. Perturbations located closer to the center of the core were synthesized more accurately, albeit with a higher number of required iterations. These findings provide a qualitative and quantitative understanding of the accuracy and speed at which different types of spatial power perturbations can be resolved in light-water reactors.
Linear convergence of accelerated conditional gradient algorithms in spaces of measures
A class of generalized conditional gradient algorithms for the solution of optimization problem in spaces of Radon measures is presented. The method iteratively inserts additional Dirac-delta functions and optimizes the corresponding coefficients. Under general assumptions, a sub-linear [see formula in PDF] rate in the objective functional is obtained, which is sharp in most cases. To improve efficiency, one can fully resolve the finite-dimensional subproblems occurring in each iteration of the method. We provide an analysis for the resulting procedure: under a structural assumption on the optimal solution, a linear [see formula in PDF] convergence rate is obtained locally.
Hybrid eigensolvers for nuclear configuration interaction calculations
We examine and compare several iterative methods for solving large-scale eigenvalue problems arising from nuclear structure calculations. In particular, we discuss the possibility of using block Lanczos method, a Chebyshev filtering based subspace iterations and the residual minimization method accelerated by direct inversion of iterative subspace (RMM-DIIS) and describe how these algorithms compare with the standard Lanczos algorithm and the locally optimal block preconditioned conjugate gradient (LOBPCG) algorithm. Although the RMM-DIIS method does not exhibit rapid convergence when the initial approximations to the desired eigenvectors are not sufficiently accurate, it can be effectively combined with either the block Lanczos or the LOBPCG method to yield a hybrid eigensolver that has several desirable properties. We will describe a few practical issues that need to be addressed to make the hybrid solver efficient and robust.
Stabilization of the 81-channel coherent beam combination using machine learning
We develop a rapidly converging algorithm for stabilizing a large channel-count diffractive optical coherent beam combination. An 81-beam combiner is controlled by a novel, machine-learning based, iterative method to correct the optical phases, operating on an experimentally calibrated numerical model. A neural-network is trained to detect phase errors based on interference pattern recognition of uncombined beams adjacent to the combined one. Due to the non-uniqueness of solutions in the full space of possible phases, the network is trained within a limited phase perturbation/error range. This also reduces the number of samples needed for training. Simulations have proven that the network can converge in one step for small phase perturbations. When the trained neural-network is applied to a realistic case of 360 degree full range, an iterative scheme exploits random walking at the beginning, with the accuracy of prediction on phase feedback direction, to allow the neural-network to step into the training range for fast convergence. This neural-network-based iterative method of phase detection works tens of times faster than the commonly used stochastic parallel gradient descent approach (SPGD) using a single-detector and random dither when both are tested with random phase perturbations.
The NILO-CMFD Method for Iteratively Solving Coupled Neutron Transport–Thermal Hydraulics Problems
Not provided.
An Efficient High-Order Solver for Diffusion Equations with Strong Anisotropy on Non-Anisotropy-Aligned Meshes
This paper concerns numerical solution of the diffusion equation with strong anisotropy on meshes not aligned with the anisotropic vector field. In order to resolve the numerical pollution for simulations on a non-anisotropy-aligned mesh and reduce the associated high computational cost we propose an effective preconditioner, extending our previous work. Similar to the anisotropy-aligned mesh case, we apply the auxiliary space preconditioning framework to design a preconditioner where a continuous finite element space is used as the auxiliary space for the discontinuous finite element space. The key component is an effective line smoother that can mitigate the high-frequency errors perpendicular to the magnetic field. We design a graph-based approach to find such a line smoother that is approximately perpendicular to the vector fields when the mesh does not align with the anisotropy. Finally, numerical experiments for several benchmark problems are presented, demonstrating the effectiveness and robustness of the proposed preconditioner when applied to Krylov iterative methods.
Considering nonlocality in the optical potentials within eikonal models
Background: For its simplicity, the eikonal method is the tool of choice to analyze nuclear reactions at high energies (E > 100 MeV/nucleon), including knockout reactions. However, so far, the effective interactions used in this method are assumed to be fully local. Purpose: Given the recent studies on nonlocal optical potentials, in this work we assess whether nonlocality in the optical potentials is expected to impact reactions at high energies and then explore different avenues for extending the eikonal method to include nonlocal interactions. Method: We compare angular distributions obtained for nonlocal interactions (using the exact R-matrix approach for elastic scattering and the adiabatic distorted wave approximation for transfer) with those obtained using their local-equivalent interactions. Results: Our results show that transfer observables are significantly impacted by nonlocality in the high-energy regime. Because knockout reactions are dominated by stripping (transfer to inelastic channels), nonlocality is expected to have a large effect on knockout observables too. Three approaches are explored for extending the eikonal method to nonlocal interactions, including an iterative method and a perturbation theory. Conclusions: None of the derived extensions of the eikonal model provide a good description of elastic scattering. Here, this paper suggests that nonlocality removes the formal simplicity associated with the eikonal model.