Search NASA⌕ Search

SEARCH · Search NASA

Results for “energy efficient computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Liquid-liquid phase transition of hydrogen and its critical point: Analysis from ab initio simulation and a machine-learned potential

We simulate high-pressure hydrogen in its liquid phase close to molecular dissociation using a machine-learned interatomic potential. The model is trained with density functional theory (DFT) forces and energies, with the Perdew-Burke-Ernzerhof (PBE) exchange-correlation functional. We show that an accurate NequIP model, an E(3)-equivariant neural network potential, accurately reproduces the phase transition present in PBE. Moreover, the computational efficiency of this model allows for substantially longer molecular dynamics trajectories, enabling us to perform a finite-size scaling (FSS) analysis to distinguish between a crossover and a true first-order phase transition. Here, we locate the critical point of this transition, the liquid-liquid phase transition (LLPT), at 1200-1300 K and 155-160 GPa, a temperature lower than most previous estimates and close to the melting transition.

08 HYDROGEN↗

Physics-Informed Machine Learning Model for Ceramic Matrix Composite Creep

A physics-informed recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear, time-dependent constitutive behavior of ceramic matrix composites (CMCs) driven by matrix damage and constituent creep at the microscale. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the high-fidelity generalized method of cells (HFGMC) approach which calls appropriate creep and damage models for each of the constituents. This coupling permits simulating the nonlinear behavior of CMCs based on constituent response at the microscale along with microstructural features such as fiber and porosity volume fraction and fiber radius. The microscale repeating unit cell is loaded under creep fatigue conditions to replicate the material loading experienced in a turbine engine. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input stress sequence, temperature, and microstructural features, the resulting strain history response while satisfying physical constraints related to creep rate, isochoric inelastic deformation, and strain energy density. The trained surrogate model is shown to effectively match the strain history over quantified distributions of microstructural features and relevant loading regimes and temperatures. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore, the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex microstructures, nonlinear time-dependent material response, and under non-monotonic loading conditions.

ceramic matrix composites↗

Spike-and-Slab Shrinkage Priors for Structurally Sparse Bayesian Neural Networks

Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.

97 MATHEMATICS AND COMPUTING↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Efficient calculation of self magnetic field, self-force, and self-inductance for electromagnetic coils with rectangular cross-section

Abstract For designing high-field electromagnets, the Lorentz force on coils needs to be computed in order to design suitable support structures, and the inductance should be computed to evaluate the stored energy and dynamics. Also, the magnetic field and its variation inside the conductor is of interest for computing stress and strain, and due to superconducting quench limits. For these force, inductance, energy, and internal field calculations, the coils cannot be naively approximated as infinitesimally thin filaments due to divergences when the source and evaluation points coincide, so more computationally demanding calculations are usually required, resolving the finite cross-section of the conductors. Here, we present a new alternative method that enables the internal magnetic field vector, self-force, and self-inductance to be computed rapidly and accurately within a 1D filament model. The method is applicable to coils for which the curve center-line can have general noncircular shape, as long as the conductor width is small compared to the radius of curvature. This paper extends a previous calculation for circular-cross-section conductors (Hurwitz et al 2024 IEEE Trans. Magn. ) to consider the case of rectangular cross-section. The reduced model is derived by rigorous analysis of the singularity, regularizing the filament integrals such that they match the true high-dimensional integrals at high coil aspect ratio. The new filament model exactly recovers analytic results for a circular coil, and is shown to accurately reproduce full finite-cross-section calculations for a non-planar coil of a stellarator magnetic fusion device. Due to the efficiency of the model here, it is well suited for use inside design optimization.

Landreman, Matt (ORCID:000000027233577X)↗

$\mathrm{SageNet}$: Fast Neural Network Emulation of the Stiff-amplified Gravitational Waves from Inflation

Accurate modeling of the inflationary gravitational waves (GWs) requires time-consuming, iterative numerical integrations of differential equations to take into account their backreaction on the expansion history. To improve computational efficiency while preserving accuracy, we present the Stiff-amplified Gravitational-wave Emulator Network (SageNet), a deep learning framework designed to replace conventional numerical solvers (code available at https://github.com/YifangLuo/SageNet). SageNet employs a long short-term memory architecture to emulate the present-day energy density spectrum of the inflationary GWs with possible stiff amplification, Ω GW (f). Trained on a data set of 25,689 numerically generated solutions, SageNet allows accurate reconstructions of Ω GW (f) and generalizes well to a wide range of cosmological parameters; 90.9% of the test emulations with randomly distributed parameters exhibit errors of under 4%. In addition, SageNet demonstrates its ability to learn and reproduce the artificial, adaptive sampling patterns in numerical calculations, which implement denser sampling of frequencies around changes in spectral indices in Ω GW (f). The dual capability of learning both physical and artificial features of the numerical GW spectra establishes SageNet as a robust alternative to exact numerical methods. Finally, our benchmark tests show that SageNet reduces the computation time from tens of seconds to milliseconds, achieving a speedup of ∼10 4 times over standard CPU-based numerical solvers with the potential for further acceleration on GPU hardware. These capabilities make SageNet a powerful tool for accelerating Bayesian inference procedures for extended cosmological models. In a broad sense, the SageNet framework offers a fast, accurate, and generalizable solution to modeling cosmological observables whose theoretical predictions demand costly differential equation solvers.

Astronomy data modeling↗

Entropy is an important design principle in the photosystem II supercomplex

Photosystem II (PSII) can achieve near-unity quantum efficiency of light harvesting in ideal conditions and can dissipate excess light energy as heat to prevent the formation of reactive oxygen species (ROS) under light stress. Understanding how this pigment–protein complex accomplishes these opposing goals is a topic of great interest that has so far been explored primarily through the lens of the system energetics. Despite PSII’s known flat energy landscape, a thorough consideration of the entropic effects on energy transfer in PSII is lacking. In this work, we aim to discern the free energetic design principles underlying the PSII energy transfer network. To accomplish this goal, we employ a structure-based rate matrix and compute the free energy terms in time following a specific initial excitation to discern how entropy and enthalpy drive ensemble system dynamics. We find that the interplay between the entropy and enthalpy components differ among each protein subunit, which allows each subunit to fulfill a unique role in the energy transfer network. This individuality ensures that PSII can accomplish efficient energy trapping in the reaction center (RC), effective nonphotochemical quenching (NPQ) in the periphery, and robust energy trapping in the other-monomer RC if the same-monomer RC is closed. We also show that entropy, in particular, is a dynamically tunable feature of the PSII free energy landscape accomplished through regulation of LHCII binding. These findings help rationalize natural photosynthesis and provide design principles for more efficient solar energy harvesting technologies.

59 BASIC BIOLOGICAL SCIENCES↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This report summarizes our contributions to the Department of Energy’s Tough Errors are no Match (TEAM) project (DE-SC0020377) under Thrust 2: Quantum Programming and Compilation. The central outcomes of this work included a novel efficient quantum compiling algorithm which works without requiring the quantum computer to exactly invert its operations, answering a longstanding open problem in quantum compiling. Additional results include the implementation of zero-noise extrapolation error mitigation in collaboration with the Unitary Fund, as well as novel quantum algorithms for entanglement detection and pseudorandomness.

Bouland, Adam [Stanford Univ., CA (United States)]↗

Artificial Intelligence for Data Center Operations (AIOps): Cooperative Research and Development (Final Report)

High performance computing data centers will increasingly need to rely on automation to keep pace with exascale growth in compute capability and to manage and optimize the data center environment and facility resources. Artificial intelligence and machine learning approaches provide the means to improve HPC data center operational efficiency, by learning historical trends and training models to operate on real-time data collected from both IT and facilities sources. NREL has developed methods of real-time collection, aggregation and streaming of these data in the ESIF HPC Data Center and has collected a significant dataset of relevant metrics across computer systems, racks, environmental, building and utility sources for research into various predictive analytics problems. HPE's Advanced Technology Group (ATG) is doing comprehensive research into exascale monitoring and management for High Performance Computing (HPC) systems (hereinafter HPE's Data Monitoring/ Management Technology). NREL and HPE will collaborate to add Artificial Intelligence (AI) to NREL's real-time data collection/ aggregation/ streaming system and HPE's Data Monitoring/ Management System, with the goal of improving the operational efficiency of NREL's Energy Systems Integration Facility (ESIF) HPC Data Center through data analytics on both historical and real-time data from IT systems and facilities operations. This collaboration will consist of efforts in Data Management, Data Analytics, and AI/ML Optimization for both manual and autonomous intervention in data center operations. This will be a multi-year, multi-staged effort with a goal towards building capabilities for an Advanced Smart Facility, and demonstration of these techniques in the NREL ESIF HPC Data Center.

97 MATHEMATICS AND COMPUTING↗

Wind Turbine Rotor Design Using High-Fidelity Aerostructural Optimization

Large wind turbines yield more energy but demand careful aeroelastic blade design. Coupled multiphysics design strategies can reduce wind energy costs by exploiting fluid-structure interactions. This work presents the first high-fidelity aerostructural optimization study of a large wind turbine rotor. We use blade-resolved fluid dynamics and structural solvers in a monolithic gradient-based optimization framework to explore steady-state torque and blade mass tradeoffs. The coupled-adjoint approach computes gradients efficiently, enabling the optimization of over 100 structural and geometric parameters simultaneously. Our optimization study modifies a DTU 10 MW benchmark with a simplified structure and isotropic material properties. The tightly coupled optimizations increase torque by 14% while reducing rotor mass by 9% or reduce blade mass by 27% while maintaining torque. Blade-resolved models provide greater design freedom, enabling 5% higher mass reductions than conventional parameterizations at equal torque. This framework paves the way for more detailed high-fidelity optimization studies to complement conventional design approaches.

17 WIND ENERGY↗

Towards a Deeper Fundamental Understanding of (Al,Sc)N Ferroelectric Nitrides

Density functional theory (DFT) calculations, within the virtual crystal alloy approximation, are performed, along with the development of a Landau-type model employing a symmetry-allowed analytical expression of the internal energy and having parameters determined from first principles, to investigate properties and energetics of Al1-xScxN ferroelectric nitrides in their hexagonal forms. These DFT computations and this model predict the existence of two different types of minima, namely, the fourfold-coordinated wurtzite (WZ) polar structure and a five-fold coordinated paraelectric hexagonal phase (denoted as H5), for any Sc composition up to 40%. The H5 minimum progressively becomes the lowest-energy state within hexagonal symmetry as the Sc concentration increases from 0 to 0.4. Furthermore, the model points to several key findings. Examples include the crucial role of the coupling between polarization and strains to create the WZ minimum, in addition to polar and elastic energies, and that the origin of the H5 state overcoming the WZ phase as the global minimum within hexagonal symmetry when increasing the Sc composition mostly lies in the compositional dependency of only two parameters-one linked to the polarization and another one being purely elastic in nature. Other examples are that forcing Al1-xScxN systems to have no or a weak change in lattice parameters when heating them allows us to reproduce their finite-temperature polar properties well and that a value of the axial ratio close to that of the ideal WZ structure implies a large polarization at low temperatures but not necessarily at high temperatures because of the ordered-disordered character of the temperature-induced formation of the WZ state. Such findings should allow for a better fundamental understanding of (Al,Sc)N ferroelectric nitrides, which may be used to design efficient devices having, e.g., low operating voltages.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A performant energy-conserving particle reweighting method for Particle-in-Cell simulations

A new particle-based reweighting method is developed and demonstrated in the Aleph Particle-in-Cell with Direct Simulation Monte Carlo (PIC-DSMC) program. Novel splitting and merging algorithms ensure that modified particles maintain physically consistent positions and velocities. This method allows a single reweighting simulation to efficiently model plasma evolution over orders of magnitude variation in density, while accurately preserving energy distribution functions (EDFs). Demonstrations on electrostatic sheath and collisional rate dynamics show that reweighting simulations achieve accuracy comparable to fixed weight simulations with substantial computational time savings. This highly performant reweighting method is recommended for modeling plasma applications that require accurate resolution of EDFs or exhibit significant density variations in time or space.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

High Performance, High Fidelity: A GPU‐Accelerated Doubly‐Periodic Configuration of the Simple Cloud‐Resolving E3SM Atmosphere Model Version 1 (DP‐SCREAMv1)

The development of the Simplified Cloud Resolving Energy Exascale Earth System Atmosphere Model (SCREAMv1) enables global storm-resolving simulations on modern GPU-based supercomputers. However, the high computational cost of SCREAMv1 limits its routine use for process-level studies, creating a need for efficient proxy configurations. This study addresses this gap by introducing DP-SCREAMv1, a doubly periodic cloud-resolving model designed to be fully consistent with SCREAMv1 while enabling high-resolution, long-duration simulations at significantly reduced computational expense by simulating a limited doubly periodic domain rather than the entire globe. Built on a C++/Kokkos architecture, DP-SCREAMv1 achieves exceptional performance scalability on GPU systems and includes a rich library of cases for validation and scientific exploration. In this work, we demonstrate short wall-clock times at SCREAMv1's default resolution and show that DP-SCREAMv1 supports routine execution of large-domain, high-resolution experiments that were previously challenging in practice. Furthermore, we show that DP-SCREAMv1 enables routine execution of “Giga-LES” style simulations and facilitates large-domain, high-resolution simulations that were recently considered burdensome to perform. These results document an efficient, fully consistent process-level configuration for SCREAMv1 (DP-SCREAMv1) and illustrate its use for long-duration and large-domain experiments at cloud-resolving to eddy-permitting resolution.

Environmental sciences↗

Plant Design for a Developing Bioeconomy Workshop Report: Frontier Science for the Bioeconomy Workshop Series

Recent advances in fundamental plant biology research, synthetic biology, and artificial intelligence (AI) are unlocking powerful new capabilities in plant biodesign, offering unprecedented potential to reimagine plants as programmable platforms for resource-efficient production of bioenergy, biomaterials, chemicals, and more. The U.S. Department of Energy (DOE) convened the Plant Design for a Developing Bioeconomy virtual workshop on March 12 through 14, 2025, to bring together leaders across plant science, engineering, and computation to assess the current landscape and define a bold vision for future research. Discussions during the workshop built upon findings included in DOE’s Biological and Environmental Research (BER) workshop report Overcoming Barriers in Plant Transformation: A Focus on Bioenergy Crops (U.S. DOE 2024; genomicscience. energy.gov/plant-transformation). Participants identified critical knowledge gaps, technical barriers, and emerging opportunities in the design and engineering of plant systems to support a robust, resilient domestic bioeconomy aligned with DOE’s mission.

09 BIOMASS FUELS↗

Transient uncertainty quantification and Global Sensitivity Analysis of the open-source Molten Chloride Reactor Experiment (MCRE) using GP-PCA surrogate models

Uncertainties in the thermophysical properties of molten salts impact both the steady-state and transient behavior of Molten Salt Reactors (MSRs). In this work, we aim to quantify the influence of such uncertainties on the transient operation of the Molten Chloride Reactor Experiment (MCRE), utilizing the open-source specifications provided for this reactor. Seven representative transient scenarios are considered. For each scenario, we evaluate the impact of thermophysical property uncertainties on four key multiphysics model output variables of interest (VoIs): maximum power density, maximum fuel temperature, maximum reflector temperature, and average fuel velocity magnitude. In addition, we perform a Global Sensitivity Analysis (GSA) by computing Sobol’ indices for the uncertain input parameters to determine their contribution to the variability of each VoI. Conducting GSA is computationally intensive due to the large number of required evaluations of the high-fidelity multiphysics model. To mitigate this cost, we develop a surrogate modeling framework that combines Gaussian Process (GP) regression with Principal Component Analysis (PCA), enabling efficient sample generation for the GSA. Our results show that for energy-related VoIs, thermal conductivity is the dominant contributor to uncertainty. In contrast, for flow-related VoIs, density and dynamic viscosity are the primary sources of uncertainty. The specific heat of the fuel salt was found to play a secondary role in the transient analyses.

42 - ENGINEERING↗

Gradient flow based phase-field modeling using separable neural networks

Allen–Cahn equation is a reaction–diffusion equation and is widely used for modeling phase separation. Machine learning methods for solving the Allen–Cahn equation in its strong form suffer from inaccuracies in collocation techniques, errors in computing higher-order spatial derivatives, and the large system size required by the space–time approach. To overcome these challenges, we propose solving the gradient flow of the Ginzburg–Landau free energy functional, which is equivalent to the Allen–Cahn equation, thereby avoiding the second-order spatial derivatives associated with the Allen–Cahn equation. A minimizing movement scheme is employed to solve the gradient flow problem, eliminating the complexities of a space–time approach. We utilize a separable neural network that efficiently represents the phase field through low-rank tensor decomposition. As we use the minimizing movement scheme to numerically solve the gradient flow problem, we thus, refer to the proposed method as the Separable Deep Minimizing Movement (SDMM) method. The evaluation of the functional in the minimizing movement scheme using the Gauss quadrature technique bypasses the inaccuracies associated with collocation techniques traditionally used to solve partial differential equations. A hyperbolic tangent transformation is introduced on the phase field prior to the evaluation of the functional to ensure that it remains strictly bounded within the values of the two phases. For this transformation, theoretical guarantee for energy stability of the minimizing movement scheme is established. Our results suggest that this transformation helps to improve the accuracy and efficiency significantly. The proposed method resolves the challenges faced by state-of-the-art machine learning techniques, outperforming them in both accuracy and efficiency. It is also the first machine learning method to achieve an order of magnitude speed improvement over the finite element method. In addition to its formulation and computational implementation, several case studies illustrate the applicability of the proposed method.

42 ENGINEERING↗

Application of a temporal multiscale method for efficient simulation of degradation in PEM Water Electrolysis under dynamic operating conditions

Hydrogen is emerging as a vital energy carrier, driven by the need to reduce carbon emissions. Proton Electrolyte Membrane Water Electrolysis (PEMWE) enables hydrogen production under fluctuating renewable power conditions but requires improved understanding and stability of the anode catalyst layer under dynamic operating conditions, especially with low noble metal loadings. Long-term degradation experiments are both time-consuming and costly; therefore, a systematic, model-aided approach is essential. In the present work, a temporal multiscale method is applied to reduce the computational effort of simulating long-term degradation processes in PEMWE, with an exemplary focus on catalyst dissolution. A mechanistic model incorporating the oxygen evolution reaction, catalyst dissolution, and hydrogen permeation from the cathode to the anode was hypothesized and implemented. In this way, the local periodicity of transport and reaction processes in dynamic PEMWE operation, which influence the gradual degradation of the catalyst layer, is captured. The temporal multiscale method significantly reduces the computational effort of simulation, decreasing processing time from hours to mere minutes. This efficiency gain is attributed to the limited evolution of Slow-Scale variables during each period of time P of the Fast-Scale variables. Consequently, simulation is required only until local periodicity is achieved within each Slow-Scale time step. Hence, the fully resolved dynamic problem is decoupled into these two scales, employing a heterogeneous multiscale technique. The developed approach effectively accelerates parameter estimation and predictive simulations, supporting systematic modeling of PEMWE degradation under dynamic conditions.

08 HYDROGEN↗

Thermo-hydraulic steam pipe models for district heating simulations: Simplifications to balance accuracy and simulation speed

Steam piping networks are essential for optimizing performance in industrial processes and district heating systems. However, dynamic models that balance thermo-hydraulic accuracy with computational efficiency remain limited. In response, this paper presents a new discretized steam pipe model based on the plug flow approach, capturing key thermo-hydraulic behaviors while simplifying steam phase change processes. Implemented in Modelica, the model accurately calculates temperature and pressure distributions along steam pipelines. To improve computational efficiency for district-scale simulations, five model simplifications are introduced: lumped thermo-hydraulic functions, empirical correlations, fluid state approximations, steady-state dynamics and inclusion of flow derivatives. These simplified models achieve 85%-98% accuracy in predicting pressure drop and condensation losses, including dynamic condensate behavior during pipe warm-up—a factor often overlooked in existing models. The models support diverse network configurations, scaling effectively to systems with multiple distribution pipes and connected building loads. Discrete models provide detailed insights but exhibit a cubic increase in simulation time as the network scales by N connected building O(N 2.42 ). In contrast, lumped models simulate 10–28 times faster than discrete, offering quadratic scaling of simulation time O(N 1.73 ). However, they still require 6 times more computation time than a lossless network, highlighting the inherent computational challenges of modeling compressible fluid flow. In conclusion, the steady-state lumped variant, with its near-linear scalability in computational time O(N 1.01 ), emerges as an efficient solution for preliminary design evaluations and extensive parametric studies.

15 GEOTHERMAL ENERGY↗