Search NASA⌕ Search

SEARCH · Search NASA

Results for “Linear solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Alternative mixed integer linear programming optimization for joint job scheduling and data allocation in grid computing

This paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as a mixed integer quadratically constrained program. To tackle the nonlinearity in the constraint, we alternatively fix a subset of decision variables and optimize the remaining ones via Mixed Integer Linear Programming (MILP). We solve the MILP problem at each iteration via an off-the-shelf MILP solver. Our experimental results show that our method significantly outperforms existing heuristic methods, employing either independent optimization or joint optimization strategies. We have also verified the generalization ability of our method over grid environments with various sizes and its high robustness to the algorithm setting.

97 MATHEMATICS AND COMPUTING↗

Flutter analysis of supersonic axial flow cascades using a high resolution Euler solver. Part 1: Formulation and validation

This report presents, in two parts, a dynamic aeroelastic stability (flutter) analysis of a cascade of blades in supersonic axial flow. Each blade of the cascade is modeled as a typical section having pitching and plunging degrees of freedom. Aerodynamic forces are obtained from a time accurate, unsteady, two-dimensional cascade solver based on the Euler equations. The solver uses a time marching flux-difference splitting (FDS) scheme. Flutter stability is analyzed in the frequency domain. The unsteady force coefficients required in the analysis are obtained by harmonically oscillating (HO) the blades for a given flow condition, oscillation frequency, and interblade phase angle. The calculated time history of the forces is then Fourier decomposed to give the required unsteady force coefficients. An influence coefficient (IC) method and a pulse response (PR) method are also implemented to reduce the computational time for the calculation of the unsteady force coefficients for any phase angle and oscillation frequency. Part 1, this report, presents these analysis methods and their validation by comparison with results obtained from linear theory for a selected flat plate cascade geometry. A typical calculation for a rotor airfoil is also included to show the applicability of the present solver for airfoil configurations. The predicted unsteady aerodynamic forces for a selected flat plate cascade geometry and flow conditions correlated well with those obtained from linear theory for different interblade phase angles and oscillation frequencies. All the three methods of predicting unsteady force coefficients, namely, HO, IC, and PR, showed good correlations with each other. It was established that only a single calculation with four blade passages is required to calculate the aerodynamic forces for any phase angle for a cascade consisting of any number of blades, for any value of the oscillation frequency. Flutter results, including mistuning effects, for a cascade of stator airfoils are presented in Part 2 of the report.

Reddy, T. S. R.↗

Parallel Implicit Algorithms for CFD

The main goal of this project was efficient distributed parallel and workstation cluster implementations of Newton-Krylov-Schwarz (NKS) solvers for implicit Computational Fluid Dynamics (CFD.) "Newton" refers to a quadratically convergent nonlinear iteration using gradient information based on the true residual, "Krylov" to an inner linear iteration that accesses the Jacobian matrix only through highly parallelizable sparse matrix-vector products, and "Schwarz" to a domain decomposition form of preconditioning the inner Krylov iterations with primarily neighbor-only exchange of data between the processors. Prior experience has established that Newton-Krylov methods are competitive solvers in the CFD context and that Krylov-Schwarz methods port well to distributed memory computers. The combination of the techniques into Newton-Krylov-Schwarz was implemented on 2D and 3D unstructured Euler codes on the parallel testbeds that used to be at LaRC and on several other parallel computers operated by other agencies or made available by the vendors. Early implementations were made directly in Massively Parallel Integration (MPI) with parallel solvers we adapted from legacy NASA codes and enhanced for full NKS functionality. Later implementations were made in the framework of the PETSC library from Argonne National Laboratory, which now includes pseudo-transient continuation Newton-Krylov-Schwarz solver capability (as a result of demands we made upon PETSC during our early porting experiences). A secondary project pursued with funding from this contract was parallel implicit solvers in acoustics, specifically in the Helmholtz formulation. A 2D acoustic inverse problem has been solved in parallel within the PETSC framework.

Keyes, David E.↗

Efficient Kriging Algorithms

More efficient versions of an interpolation method, called kriging, have been introduced in order to reduce its traditionally high computational cost. Written in C++, these approaches were tested on both synthetic and real data. Kriging is a best unbiased linear estimator and suitable for interpolation of scattered data points. Kriging has long been used in the geostatistic and mining communities, but is now being researched for use in the image fusion of remotely sensed data. This allows a combination of data from various locations to be used to fill in any missing data from any single location. To arrive at the faster algorithms, sparse SYMMLQ iterative solver, covariance tapering, Fast Multipole Methods (FMM), and nearest neighbor searching techniques were used. These implementations were used when the coefficient matrix in the linear system is symmetric, but not necessarily positive-definite.

Memarsadeghi, Nargess↗

Modeling Boundary-Layer Transition in Subsonic Flow over a Swept Wing

Predicting the onset of boundary-layer transition is often more accurate using physics-based models that directly compute disturbance growth rather than phenomenological models often implemented into industrial CFD codes. The aim of this ongoing study is to calibrate linear, physics-based computations of transition in subsonic flows over swept wings against a large set of experimental data. Advancing the calibration of linear models of transition contributes to the CFD-Vision-2030 goal of automated boundary-layer transition prediction. This progress report uses the dual N-factor method to model transition over the swept NACA 64-2-015A wing. The flow conditions match selected test conditions from an extensive experimental dataset acquired from the NASA Ames 12-ft Pressure Tunnel. The OVERFLOW 2.4b flow solver is used to obtain laminar basic states based on an infinite-span assumption. Stability analyses are performed on 365 distinct configurations with linear stability theory (LST) and parabolized stability equations (PSE) from the Langley Stability and Transition Analysis Codes (LASTRAC), modeling the growth of Tollmien-Schlichting (TS) and stationary crossflow (SCF) disturbances. From a total of 67 data points for unswept, i.e., TS-dominant configurations, the critical N-factor based on PSE is found to be N_TS = 9. The SCF critical N-factor is found to be near 8 for the highly swept, SCF-dominant configurations. Dual N-factor curves for both LST and PSE computations demonstrate a high level of interaction between TS and SCF. It may be worthwhile to investigate an alternate metric to visualize maximal SCF amplification upstream of the transition location to account for the growth of SCF modes near the leading edge, which is not considered in the conventional applications of the dual N-factor criterion.

boundary-layer transition↗

Modeling Boundary-Layer Transition in Subsonic Flow over a Swept Wing

Predicting the onset of boundary-layer transition is often more accurate using physics-based models that directly compute disturbance growth rather than phenomenological models often implemented into industrial CFD codes. The aim of this ongoing study is to calibrate linear, physics-based computations of transition in subsonic flows over swept wings against a large set of experimental data. Advancing the calibration of linear models of transition contributes to the CFD-Vision-2030 goal of automated boundary-layer transition prediction. This progress report uses the dual N-factor method to model transition over the swept NACA 64-2-015A wing. The flow conditions match selected test conditions from an extensive experimental dataset acquired from the NASA Ames 12-ft Pressure Tunnel. The OVERFLOW 2.4b flow solver is used to obtain laminar basic states based on an infinite-span assumption. Stability analyses are performed on 365 distinct configurations with linear stability theory (LST) and parabolized stability equations (PSE) from the Langley Stability and Transition Analysis Codes (LASTRAC), modeling the growth of Tollmien-Schlichting (TS) and stationary crossflow (SCF) disturbances. From a total of 67 data points for unswept, i.e., TS-dominant configurations, the critical N-factor based on PSE is found to be N_TS = 9. The SCF critical N-factor is found to be near 8 for the highly swept, SCF-dominant configurations. Dual N-factor curves for both LST and PSE computations demonstrate a high level of interaction between TS and SCF. It may be worthwhile to investigate an alternate metric to visualize maximal SCF amplification upstream of the transition location to account for the growth of SCF modes near the leading edge, which is not considered in the conventional applications of the dual N-factor criterion.

computational modeling↗

Fast solvers for tokamak fluid models with PETSc

Multigrid (MG) is widely recognized as a highly effective solver for the model problem, the Laplacian, but textbook MG fails on most problems of interest. MG methods have been applied to complex, real-world applications with careful consideration of the physical model and discretization. In this work we develop the first step in applying MG methods to science and engineering relevant magnetohydrodynamics (MHD) tokamak models in the M3D-C1 (https://m3dc1.pppl.gov) fusion energy science code. The semi-implicit time integrator in M3D-C1 is composed of many linear solves. The implicit advance of the momentum equation is the most challenging and is the focus of this work. The current production solver in M3D-C1 is a block Jacobi (BJ) preconditioner within a Krylov solver, where blocks group degrees of freedom on planes of constant toroidal coordinate. BJ convergence degrades as the number of planes increases due to the spectral properties of the matrix preconditioned with BJ. The partially magnetic field-aligned, regular toroidal grid structure in M3D-C1 is amenable to semi-coarsening geometric MG in the toroidal direction. This paper develops such a solver and demonstrates competitive performance on a runaway electron model of a SPARC (https://cfs.energy/technology/sparc) disruption, and superior robustness on a stellarator model on which the BJ solver fails to converge.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Finite element analysis of periodic transonic flow problems

Flow about an oscillating thin airfoil in a transonic stream was considered. It was assumed that the flow field can be decomposed into a mean flow plus a periodic perturbation. On the surface of the airfoil the usual Neumman conditions are imposed. Two computer programs were written, both using linear basis functions over triangles for the finite element space. The first program uses a banded Gaussian elimination solver to solve the matrix problem, while the second uses an iterative technique, namely SOR. The only results obtained are for an oscillating flat plate.

Fix, G. J.↗

eddy Users Manual

eddy is a collection of tools - nonlinear solvers, meshing, post-processing, visualization, optimization, etc. - for performing scale-resolving simulations of multi-physics applications. The framework is designed to enable advanced R&D on a variety of topics by leveraging a mature capability for scale resolving simulations, and simultaneously be an appropriate tool for application analysis and support. Currently, eddy is at a relatively low technical readiness level (TRL), and users and developers should maintain appropriate expectations. The technical details behind eddy are outlined in several publications which can be consulted for more information [1–10]. The solvers are built around an unstructured high-order capability, and heavily utilize the tensor product sum-factorization approach for efficiency. The unsteady formulation utilizes a fully implicit space-time approach with a matrix-free Newton- Krylov method. A primitive steady-state solver is available for testing purposes, but is not expected to converge for all but simple verification cases. The Navier-Stokes fluid solvers do not support either RANS or hybrid-RANS capability, only LES and wall-modeled LES approaches. All of the solvers within eddy support three modes of operation: a primal solve of the full nonlinear problem, and two linearization approaches of the primal solve - the ad joint and the tangent solution. Details on how to select and use these three modes are outlined in Sec. 3.

Murman, Scott M.↗

Performance issues for iterative solvers in device simulation

Due to memory limitations, iterative methods have become the method of choice for large scale semiconductor device simulation. However, it is well known that these methods still suffer from reliability problems. The linear systems which appear in numerical simulation of semiconductor devices are notoriously ill-conditioned. In order to produce robust algorithms for practical problems, careful attention must be given to many implementation issues. This paper concentrates on strategies for developing robust preconditioners. In addition, effective data structures and convergence check issues are also discussed. These algorithms are compared with a standard direct sparse matrix solver on a variety of problems.

Fan, Qing↗

The semi-discrete Galerkin finite element modelling of compressible viscous flow past an airfoil

A method is developed to solve the two-dimensional, steady, compressible, turbulent boundary-layer equations and is coupled to an existing Euler solver for attached transonic airfoil analysis problems. The boundary-layer formulation utilizes the semi-discrete Galerkin (SDG) method to model the spatial variable normal to the surface with linear finite elements and the time-like variable with finite differences. A Dorodnitsyn transformed system of equations is used to bound the infinite spatial domain thereby permitting the use of a uniform finite element grid which provides high resolution near the wall and automatically follows boundary-layer growth. The second-order accurate Crank-Nicholson scheme is applied along with a linearization method to take advantage of the parabolic nature of the boundary-layer equations and generate a non-iterative marching routine. The SDG code can be applied to any smoothly-connected airfoil shape without modification and can be coupled to any inviscid flow solver. In this analysis, a direct viscous-inviscid interaction is accomplished between the Euler and boundary-layer codes, through the application of a transpiration velocity boundary condition. Results are presented for compressible turbulent flow past NACA 0012 and RAE 2822 airfoils at various freestream Mach numbers, Reynolds numbers, and angles of attack. All results show good agreement with experiment, and the coupled code proved to be a computationally-efficient and accurate airfoil analysis tool.

Meade, Andrew J., Jr.↗

A data-driven approach to real-time vertical position estimation for NSTX-U vertical stability control

In this paper, a database of 77 996 plasma equilibrium reconstructions from 727 discharges during the initial operation of the NSTX-U spherical tokamak is analyzed to develop a statistically robust model of the plasma vertical position for real-time control. A variety of regression models are developed and tested, ranging in complexity from linear models to deep neural networks, and including input signals ranging from the four pairs of flux loops used historically on NSTX-U up to the full set of 389 real-time signals available to the plasma control system. A linear model based on 140 real-time magnetics signals is found to offer excellent accuracy, with a coefficient of determination R 2 = 0.906. The robustness of this model to limited training data, new operating scenarios, and signal errors is tested, and a procedure is demonstrated to tune the model parameters to optimize its robustness. A time-dependent plasma equilibrium solver, TokaMaker, is used to simulate vertical stability control in NSTX-U, demonstrating that it should be possible to iteratively tune the parameters of a linear vertical position model to stabilize both positive and negative triangularity plasmas in future experiments.

magnetic diagnostics↗

The High-Resolution Wave-Propagation Method Applied to Meso- and Micro-Scale Flows

The high-resolution wave-propagation method for computing the nonhydrostatic atmospheric flows on meso- and micro-scales is described. The design and implementation of the Riemann solver used for computing the Godunov fluxes is discussed in detail. The method uses a flux-based wave decomposition in which the flux differences are written directly as the linear combination of the right eigenvectors of the hyperbolic system. The two advantages of the technique are: 1) the need for an explicit definition of the Roe matrix is eliminated and, 2) the inclusion of source term due to gravity does not result in discretization errors. The resulting flow solver is conservative and able to resolve regions of large gradients without introducing dispersion errors. The methodology is validated against exact analytical solutions and benchmark cases for non-hydrostatic atmospheric flows.

Ahmad, Nashat N.↗

Combined, nonlinear aerodynamic and structural method for the aeroelastic design of a three-dimensional wing in supersonic flow

An iterative procedure for the static aeroelastic design of a flexible wing at supersonic speeds has been developed. The procedure combines a nonlinear, full-potential solver (NCOREL) with an equivalent plate structural analysis method. The NCOREL method yields significantly improved aerodynamic estimates compared to linear theory. The equivalent plate structural analysis method demonstrates an order of magnitude reduction in computer memory and execution time compared to finite-element methods. A highly swept wing is analyzed at high lift using this aeroelastic procedure. The results indicate that the wing deforms favorably due to aerodynamic loading and, consequently, that the inviscid drag levels do not vary at the required lift coefficient although the angle of attack varies significantly. A sensitivity analysis of the type required for optimization studies was also performed with the aeroelastic design procedure.

Pittman, J. L.↗

A global multilevel atmospheric model using a vector semi-Lagrangian finite-difference scheme. I - Adiabatic formulation

An adiabatic global multilevel primitive equation model using a two time-level, semi-Lagrangian semi-implicit finite-difference integration scheme is presented. A Lorenz grid is used for vertical discretization and a C grid for the horizontal discretization. The momentum equation is discretized in vector form, thus avoiding problems near the poles. The 3D model equations are reduced by a linear transformation to a set of 2D elliptic equations, whose solution is found by means of an efficient direct solver. The model (with minimal physics) is integrated for 10 days starting from an initialized state derived from real data. A resolution of 16 levels in the vertical is used, with various horizontal resolutions. The model is found to be stable and efficient, and to give realistic output fields. Integrations with time steps of 10 min, 30 min, and 1 h are compared, and the differences are found to be acceptable.

Bates, J. R.↗

On the Development of an Efficient Parallel Hybrid Solver with Application to Acoustically Treated Aero-Engine Nacelles

A finite element solution to the convected Helmholtz equation in a nonuniform flow is used to model the noise field within 3-D acoustically treated aero-engine nacelles. Options to select linear or cubic Hermite polynomial basis functions and isoparametric elements are included. However, the key feature of the method is a domain decomposition procedure that is based upon the inter-mixing of an iterative and a direct solve strategy for solving the discrete finite element equations. This procedure is optimized to take full advantage of sparsity and exploit the increased memory and parallel processing capability of modern computer architectures. Example computations are presented for the Langley Flow Impedance Test facility and a rectangular mapping of a full scale, generic aero-engine nacelle. The accuracy and parallel performance of this new solver are tested on both model problems using a supercomputer that contains hundreds of central processing units. Results show that the method gives extremely accurate attenuation predictions, achieves super-linear speedup over hundreds of CPUs, and solves upward of 25 million complex equations in a quarter of an hour.

Watson, Willie R.↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

Simulations of bypass transition for spatially evolving disturbances

The spatial evolution of disturbances in plane Poiseuille flow and zero pressure gradient boundary layer flow is considered. For disturbances governed by the linearized equations, potential for significant transient growth of the amplitude is demonstrated. The maximum amplification occurs for disturbances with zero or near zero frequencies. Spatial numerical simulations of the transition scenario involving a pair of oblique waves has been conducted for both flows. A fully spectral solver using a simple but efficient fringe region technique allowed the flows to be computed with high resolution into the fully turbulent domain. A modal decomposition of the simulation results indicates that non-linear excitation of the transient growth is responsible for the rapid emergence of low-frequency structures. Physically, this corresponds to streaky flow structures, as seen from the results of a numerical amplitude expansion. Thus, this spatial transition scenario has been found to be similar to the corresponding temporal one. In the boundary layer simulations the streaks are seen to break down from what appears to be a secondary instability.

Lundbladh, A.↗