Search NASA⌕ Search

SEARCH · Search NASA

Results for “Numerical optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Two-level overlapping additive Schwarz preconditioner for training scientific machine learning applications

In this work we introduce a novel two-level overlapping additive Schwarz preconditioner for accelerating the training of scientific machine learning applications. The design of the proposed preconditioner is motivated by the nonlinear two-level overlapping additive Schwarz preconditioner. The neural network parameters are decomposed into groups (subdomains) with overlapping regions. In addition, the network’s feed-forward structure is indirectly imposed through a novel subdomain-wise synchronization strategy and a coarse-level training step. Through a series of numerical experiments, which consider physicsinformed neural networks and operator learning approaches, we demonstrate that the proposed two-level preconditioner significantly speeds up the convergence of the standard (LBFGS) optimizer while also yielding more accurate machine learning models. Moreover, the devised preconditioner is designed to take advantage of model-parallel computations, which can further reduce the training time.

97 MATHEMATICS AND COMPUTING↗

Development of a multi-layer canopy model for E3SM Land Model with support for heterogeneous computing

The vertical structure of vegetation canopies creates micro-climates. However, the land components of most Earth System Models, including the Energy Exascale Earth System Model (E3SM), typically neglect vertical canopy structure by using a single layer big-leaf representation to simulate water, CO 2 , and energy exchanges between the land and the atmosphere. In this study, we developed a Multi-Layer Canopy Model for the E3SM Land Model to resolve the micro-climate created by vegetation canopies. The model developed in this study re-implements the CLM-ml_v1 to support heterogeneous computing architectures consisting of CPUs and GPUs and includes three additional optimization-based stomatal conductance models. The use of Portable, Extensible Toolkit for Scientific Computation provides a speedup of 25–50 times on a GPU relative to a CPU. The numerical implementation of the model was verified against CLM-ml_v1 for a month-long simulation using data from the Ameriflux US-University of Michigan Biological Station site. Model structural uncertainty was explored by performing control simulations for five stomatal conductance models that exclude and include the control of plant hydrodynamics (PHD) on photosynthesis. The bias in simulated sensible and latent heat fluxes was lower when PHD was accounted for in the model. Additionally, six idealized simulations were performed to study the impact of three environmental variables (i.e. air temperature, atmospheric CO 2 , and soil moisture) on canopy processes (i.e. net CO 2 assimilation, leaf temperature, and leaf water potential). Increasing air temperature reduced net CO 2 assimilation and increased air temperature. Net CO 2 assimilation increased at higher atmospheric CO 2 , while decreasing soil moisture resulted in lower leaf water potential.

54 ENVIRONMENTAL SCIENCES↗

Faster Randomized Dynamical Decoupling

We present a randomized dynamical decoupling (DD) protocol that can substantially improve the performance of any given deterministic DD scheme for suppressing coherent noise by using no more than two additional pulses. Our construction is implemented by probabilistically applying sequences of pulses, which, when combined, effectively eliminate the error terms that scale linearly with the system-environment coupling strength. As a result, we show that a randomized protocol using a few pulses can outperform deterministic DD protocols that require considerably more pulses. Furthermore, we prove that the randomized protocol provides an improvement compared to deterministic DD sequences that aim to reduce the error in the system’s Hilbert space, such as Uhrig DD, which had been previously regarded to be optimal. To rigorously evaluate the performance, we introduce new analytical methods suitable for analyzing higher-order DD protocols that might be of independent interest. Here, we also present numerical simulations confirming the significant advantage of using randomized protocols compared to widely used deterministic protocols.

Quantum algorithms & computation↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

Development, Monitoring, and Control of Fracture Thermal Energy Storage (FTES) in Crystalline Rock Formations (DEMO-FTES) (CRADA Final Report)

The DEMO-FTES project sought to demonstrate the thermal efficiency of fracture thermal energy storage (FTES) through numerical simulations, laboratory and meso-scale field tests. A detailed dimensional and scaling analysis was performed to identify key parameters and how they can be most effectively scaled to the laboratory and decameter scale. Numerical models were developed and used for three purposes: 1. Before field experiments, numerical modelling can be used to estimate fracture properties based on previous data from the EGS Collab experiment and then predict thermal hydrological behaviors of the fracture system with hot water injection/withdrawal, therefore, help to design the experiments (e.g., to decide the duration of the cycles based on the flow rate the pump can provide, and the estimated fracture properties); 2. After the field experiment, to estimate the system properties during the experiment (as the size and shape of a fracture could change over time), and understand why system performance is different than what has been predicted, i.e., to help understand the meso-scale test; and, 3. To model the lab experiments and estimate fracture properties and storage efficiency. Ultimately, the experiment and numerical models could shed light on the processes and uncertainty happening during fracture activation and help understand the scaling between lab and field test, and finally, the design and optimization of potential fracture thermal energy storage systems.

25 ENERGY STORAGE↗

Ground and excited state gradients with end-to-end differentiable semiempirical quantum chemistry

Accurate and efficient gradients of molecular energy with respect to nuclear degrees of freedom are essential for geometry optimization and molecular dynamics, including simulations that go beyond the Born–Oppenheimer regime. A common approach involves deriving analytical formulas for new electronic structure methods, which is often conceptually difficult and requires tedious coding. Here, we implement analytical, semi-numerical, and automatic differentiation (AD)-based gradient pathways for semiempirical Hamiltonian models in the PYSEQM software package, leveraging both graphics processing unit (GPU) and central processing unit (CPU) architectures. We further extend these capabilities to excited states calculated using the configuration interaction singles and time-dependent Hartree–Fock ansätze. We benchmark wall time, peak memory usage, and accuracy across three molecular families of varying chemical complexity, including systems of up to a thousand atoms. For ground-state simulations, analytical and AD gradients achieve near-identical GPU runtimes, while semi-numerical gradients are slower on GPU but remain competitive on CPU. For excited states, both analytical and custom AD approaches using implicit differentiation show similar performance and low memory requirements, whereas gradients with full AD are memory-limited. AD gradients match analytical ones in accuracy across all tested systems, aided by a quaternion-based diatomic frame rotation for two-center quantities that ensures smooth energy surfaces. Overall, automatic differentiation emerges as a practical alternative to analytical gradients in semiempirical quantum chemistry, offering high accuracy while allowing seamless integration in AI-driven workflows and popular packages, such as PyTorch and JAX. Our results provide actionable guidance for selecting optimal gradient strategies in large-scale ground- and excited-state molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Self-consistent modeling of tokamak edge plasma transport with lithium sources

Magnetic confinement fusion devices require effective heat and particle exhaust solutions on the divertor plates to operate sustainably, especially under reactor-relevant conditions. Liquid lithium divertors have been proposed to address two major challenges: control of excessive heat flux to plasma-facing components through vapor shielding and minimization of core plasma contamination from impurities. The National Spherical Torus Experiment-Upgrade (NSTX-U) will explore lithium as a divertor material due to its potential to meet both objectives. We present a self-consistent coupling framework between the plasma boundary transport code UEDGE and the lithium wall transport code Wall–Li to evaluate the feasibility and operational limits of lithium-based divertors. The model aims to optimize lithium sourcing levels to prevent core plasma contamination via fuel dilution while ensuring divertor protection through vapor shielding. This integrated framework, applicable to any tokamak with lithium sources, dynamically adjusts lithium sourcing based on local plasma conditions and surface temperature. The coupled model is tested using NSTX-like geometry and plasma conditions to assess its performance and reliability. Wall–Li calculates lithium fluxes from plasma-facing components, incorporating physical sputtering, thermally enhanced sputtering, and evaporation driven by surface temperature and ion flux. These fluxes are reintroduced into UEDGE as neutral lithium atoms, enabling simulation of their transport and distribution within the plasma. UEDGE computes plasma and neutral transport, surface heat flux, and iteratively feeds this information back to Wall–Li. A small time step is employed to ensure numerical stability and convergence, enabling accurate simulations over typical tokamak discharge durations. This integrated modeling approach provides a robust tool for identifying operational regimes that balance effective lithium sourcing with minimal core plasma contamination, offering critical insights for optimizing lithium-based divertor systems in current and future fusion devices.

Magnetic confinement fusion↗

A scalable multidimensional fully implicit solver for Hall magnetohydrodynamics

We propose an optimally performant fully implicit algorithm for the Hall magnetohydrodynamics (HMHD) equations based on multigrid-preconditioned Jacobian-free Newton-Krylov methods. HMHD is a challenging system to solve numerically because it supports stiff fast dispersive waves. The preconditioner is formulated using an operator-split approximate block factorization (Schur complement), informed by physics insight. We use a vector-potential formulation (instead of a magnetic field one) to allow a clean segregation of the problematic $\nabla$ x $\nabla$ x operator in the electron Ohm's law subsystem. This segregation allows the formulation of an effective damped block-Jacobi smoother for multigrid. We demonstrate by analysis that our proposed block-Jacobi iteration is convergent and has the smoothing property. The resulting HMHD solver is verified linearly with wave propagation examples, and nonlinearly with the GEM challenge reconnection problem by comparison against another HMHD code. We demonstrate the excellent algorithmic and parallel performance of the algorithm up to 16384 MPI tasks in two dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Revenue-Maximizing Shared Parking and Electric Vehicle Charging Management in Multi-Unit Dwellings

In urban areas, searching for parking and electric vehicle (EV) charging can result in cruising, congestion, and environmental externalities. Recognizing the business opportunity of offering private parking and charging infrastructure access within multi-unit dwellings (MUDs) during daytime, we model a shared parking and EV charging management system. We maximize the revenue of MUD charging hubs in mixed land use, catering to public demand. Our approach accounts for the objectives of the two stakeholders involved: a demand model is fitted on the choices of EV charging users, and the supply model optimizes the allocation of parking and charging requests in an MUD parking lot. A binary integer linear programming model for the allocation of parking and charging spaces with a rolling horizon is integrated with matching rules that handle both parking and charging requests. In our numerical experiments in a neighborhood of Chicago, Illinois, we estimate the performance of the MUD parking and charging system with metrics that include revenue, number of matchings, and utilization rates. At any given time, MUDs with lower prices attract more charging requests, particularly those of longer duration, resulting in higher revenue and greater charging utilization. Dynamic pricing facilitates a more equitable distribution of requests; as MUD parking lots reach capacity and their fees increase, other MUDs become more competitive, attracting additional requests. Comparing our method against first-come-first-served and optimal-solution benchmarks, we demonstrate our model’s effectiveness in dynamically managing mixed parking and charging demand in MUD charging hubs.

electric vehicle, multi-unit dwelling, charging in↗

R-Adaptivity to Enable Compression of Elementary Computations in Extreme-Scale Finite Element Simulators

Modern computing systems are capable of exascale calculations, which are revolutionizing the development and application of high-fidelity numerical models in computational science and engineering. While these systems continue to grow in processing power, the available system memory has not increased commensurately, and electrical power consumption continues to grow. A predominant approach to limit the memory usage in large-scale applications is to exploit the abundant processing power and continually recompute many low-level simulation quantities, rather than storing them. However, this approach can adversely impact the throughput of the simulation and diminish the benefits of modern computing architectures. We present three novel contributions to reduce the memory burden while maintaining, and sometimes improving, performance in simulations based on finite element discretizations. The first contribution develops dictionary-based data compression schemes that detect and exploit the structure of the discretization, due to redundancies across the finite element mesh. While these schemes are shown to reduce memory requirements by more than 99% on meshes with large numbers of identical mesh cells, there are applications where this structure does not exist. The second contribution leverages a recently developed augmented Lagrangian optimization algorithm to enable r-adaptivity for meshes with the goal of enhancing the redundancies in the mesh. The third contribution extends these methods to patch-based linear solvers and preconditioners by compressing local matrices. Numerical results demonstrate the effectiveness of the proposed methods to detect, enhance and exploit mesh structure on a suite of examples inspired by large-scale applications.

97 MATHEMATICS AND COMPUTING↗

Revealing Decision Conservativeness Through Inverse Distributionally Robust Optimization

This paper introduces Inverse Distributionally Robust Optimization (I-DRO) as a method to infer the conservativeness level of a decision-maker, represented by the size of a Wasserstein metric-based ambiguity set, from the optimal decisions made using Forward Distributionally Robust Optimization (F-DRO). By leveraging the Karush-Kuhn-Tucker (KKT) conditions of the convex F-DRO model, we formulate I-DRO as a bi-linear program, which can be solved using off-the-shelf optimization solvers. Additionally, this formulation exhibits several advantageous properties. We demonstrate that I-DRO not only guarantees the existence and uniqueness of an optimal solution but also establishes the necessary and sufficient conditions for this optimal solution to accurately match the actual conservativeness level in F-DRO. Furthermore, we identify three extreme scenarios that may impact I-DRO effectiveness. Our case study applies F-DRO for power system scheduling under uncertainty and employs I-DRO to recover the conservativeness level of system operators. Numerical experiments based on an IEEE 5-bus system and a realistic NYISO 11-zone system demonstrate I-DRO performance in both normal and extreme scenarios. An extended version of this paper with additional analyses is available at li2024revealing.

distributionally robust optimization↗

On measuring stimulated photon–photon scattering using multiple ultraintense lasers

Stimulated photon–photon scattering is a predicted consequence of quantum electrodynamics that has yet to be measured directly. Measuring the cross section for stimulated photon–photon scattering is the aim of a flagship experiment for NSF OPAL, a proposed laser user facility with two, 25-PW beamlines. We present optimized experimental designs for achieving this challenging and canonical measurement. A family of experimental geometries is identified that satisfies the momentum- and energy-matching conditions for two selected laser frequency options. Numerical models predict a maximum signal exceeding 1000 scattered photons per shot at the experimental conditions envisaged at NSF OPAL. Experimental requirements on collision geometry, polarization, cotiming and copointing, background suppression, and diagnostic technologies are investigated numerically. These results confirm that a beam cotiming shorter than the pulse duration and control of the copointing on a scale smaller than the shortest laser wavelength are needed to robustly scatter photons on a per-shot basis. Finally, we assess the bounds that a successful execution of this experiment may place on the mass scale of Born–Infeld nonlinear electrodynamics beyond the standard model of physics.

Beyond the Standard Model↗

Localized Evaluation for Constructing Discrete Vector Fields

Topological abstractions offer a method to summarize the behavior of vector fields, but computing them robustly can be challenging due to numerical precision issues. One alternative is to represent the vector field using a discrete approach, which constructs a collection of pairs of simplices in the input mesh that satisfies criteria introduced by Forman's discrete Morse theory. While numerous approaches exist to compute pairs in the restricted case of the gradient of a scalar field, state-of-the-art algorithms for the general case of vector fields require expensive optimization procedures. This paper introduces a fast, novel approach for pairing simplices of two-dimensional, triangulated vector fields that do not vary in time. The key insight of our approach is that we can employ a local evaluation, inspired by the approach used to construct a discrete gradient field, where every simplex in a mesh is considered by no more than one of its vertices. Specifically, we observe that for any edge in the input mesh, we can uniquely assign an outward direction of flow. We can further expand this consistent notion of outward flow at each vertex, which corresponds to the concept of a downhill flow in the case of scalar fields. Working with outward flow enables a linear-time algorithm that processes the (outward) neighborhoods of each vertex one-by-one, similar to the approach used for scalar fields. Here, we couple our approach to constructing discrete vector fields with a method to extract, simplify, and visualize topological features. Empirical results on analytic and simulation data demonstrate drastic improvements in running time, produce features similar to the current state-of-the-art, and show the application of simplification to large, complex flows.

97 MATHEMATICS AND COMPUTING↗

Meshfree simulation and prediction of recrystallized grain size in friction stir processed 316L stainless steel

Friction stir processing (FSP) is a promising solid-phase microstructural modification technique that can repair and enhance damaged stainless steel surfaces exposed to harsh environments. The quality of the repaired material is closely correlated to the recrystallized grain size in the stir zone (SZ), which is influenced by the thermomechanical conditions dictated by FSP process parameters. Thus, establishing a reliable relationship between these parameters and recrystallized grain size in the SZ is crucial for optimizing repair quality. However, existing experimental approaches often rely on indirect temperatures measured far from the SZ, along with rough strain rate estimations, which are imprecise and time-consuming. Meanwhile, existing mesh-based modeling methods usually face numerical challenges when dealing with the large material deformations inherent in FSP. Here, to address these issues, this study introduces a meshfree process model for FSP based on the smoothed particle hydrodynamics (SPH) method, aimed at predicting process conditions under different parameters. The model is validated using experimental data from 11 combinations of tool traverse and rotation speeds on 316 L stainless steel. Correlations between process parameters, material flow, temperature, strain, strain rate, and recrystallized grain size are revealed through SPH simulations and electron backscatter diffraction (EBSD) imaging. The results show that in situ SZ temperatures range from 1071 to 1322°C, which exceed the tool temperature by over 300°C. Furthermore, SZ temperature, strain rate, and grain size increase monotonically with higher tool temperature and faster traverse speed. A relationship is then established between the model-predicted Zener-Hollomon parameter and the recrystallized grain size based on EBSD data, expressed as ln(d) = -0.364 ln(Z) + 14.673. Finally, this relationship exhibits satisfactory accuracy with errors of less than 26.9% in predicting grain sizes at various SZ locations, which offers valuable insights for optimizing FSP repair processes for 316 L stainless steel.

316L stainless steel↗

Mitigation of distortion of Al/steel part under simulated paint baking condition: Experiment and numerical model studies

Multi-material joining of lightweight structures is essential to reduce vehicle weight for more energy savings and less greenhouse gas emission. However, mismatch of thermal expansion coefficient for dissimilar materials during the paint baking process can induce part distortion and joint failure for adhesive bonding. Here, in the present work, a thermomechanical model based on contact mechanics and large deformation theory was developed for dissimilar high-strength Al alloy and steel components to study the distortion mechanism and influential factors of the residual gap. The established model was used to optimize joint conditions, such as pitch distance and part geometry. When a weld pitch is shorter than 100 mm, the maximum gap between Al and steel part can be greatly reduced to 0.1 mm, and the local stress and plastic strain around the joint during the oven heating and cooling cycle are also substantially reduced compared with the long pitch case (900 mm). The numerical modeling results revealed that a comparable bending stiffness ratio between the steel and Al cross sections is critical to the minimization of gap and distortion under paint baking condition. Digital image correlation technique was used to measure the overall part distortion and local strain distribution that were used to validate the model prediction. Weld bonding (adhesive bonding with friction bit joining) process was successfully employed to join Al to steel component without gap opening in adhesive after the paint baking and cooling.

36 MATERIALS SCIENCE↗

Establish the basis for Breadth-First Search on Frontier System: XBFS on AMD GPUs

Graphics Processing Units (GPUs) offer significant potential for accelerating various computational tasks, including Breadth-First Search (BFS). Numerous efforts have been made to deploy BFS on GPUs effectively. To address the dynamic nature of BFS, XBFS, the state-of-the-art work, employs an adaptive strategy that leverages different optimized frontier queue generation designs, accommodating the varying characteristics of levels in BFS. While XBFS demonstrates excellent performance on NVIDIA Quadro P6000 GPUs, it faces challenges when deployed on AMD GPUs. In this work, we present our efforts to implement XBFS’s adaptive approach on Frontier, the most powerful supercomputer system, by porting XBFS to AMD MI250X GPUs. Through targeted optimizations tailored to the unique features of AMD GPUs, our implementation achieves an average performance of 43 Giga-Traversed Edges Per Second (GTEPS) per Graphics Compute Dies (GCD). Based on these results, we observe potential for surpassing the performance of the official Frontier results from the Graph500 benchmark released in June 2024.

Yang, Haoshen↗

A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications

Our work on the DOE-sponsored project “A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications,” was an effort to address critical challenges in nu merical computing and its applications to optimization. The increasing demand for robust and scalable solutions to large-scale linear algebra problems has highlighted the limitations of traditional approaches, particularly in heterogeneous and extreme-scale computing environments. Randomized Numerical Linear Algebra (RandNLA) offers a promising framework to address these challenges, and this proposal builds on this foundation by introducing innovations in sensitivity analysis and computational adaptability.

97 MATHEMATICS AND COMPUTING↗