Search NASA⌕ Search

SEARCH · Search NASA

Results for “energy efficient computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Monte Carlo Event Generation with Continuous Normalizing Flows

We apply continuous normalizing flows trained with the flow matching method to the problem of phase-space sampling in Monte Carlo event generation for high-energy collider physics. Focusing on lepton-pair and top-quark pair production with multiple jets, the two computationally most expensive processes at the Large Hadron Collider, we train helicity-conditioned continuous normalizing flows to remap the random numbers used in matrix element evaluation. Compared to standard methods, we achieve unweighting efficiency improvements by factors of up to 184 and 25 for the two processes at their respective highest jet number, at the cost of an increased evaluation time. When combining the advantages of continuous normalizing flows with the fast evaluation times of coupling-layer-based flows, using the RegFlow approach, we find parton-level unweighted event generation walltime gains of about a factor of 10 at the highest jet numbers. These substantial gains highlight the promise of samplers based on machine learning for next-generation collider experiments.

Bothmann, Enrico [CERN; Gottingen U.] (ORCID:00000↗

A generalized and adaptable tensor-contraction-based cluster expansion formalism for multicomponent solids

Density functional theory (DFT)-based simulations of materials have first-principles accuracy, but are very computationally expensive. For simulating various properties of multi-component alloys, the cluster expansion (CE) technique has served as the standard workaround to improve computational efficiency. However, the standard CE technique is difficult to extend to exotic and/or low-symmetry lattices, often implemented via iteration over particular cluster types, which must be enumerated per lattice structure. In this work, we introduce the tensor cluster expansion (TCE), implemented in the open-source code tce-lib, which maps correlation functions to mixed tensor contractions, eliminating the need to iterate over cluster types and additionally making the calculation of correlation functions well-suited for massively parallel architectures like GPUs. We show that local interaction energies are an immediate consequence of the TCE formalism, yielding nearly $\mathcal{O}$(1) energy difference calculations. We then use this formalism to fit CE models for the TaW and CoNiCrFeMn systems, and use these models to respectively compute the enthalpy of mixing curve and Cowley short-range order parameters, showing excellent agreement with ground truth data.

Cluster expansion↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

Charge Transport in Solvated Donor–Acceptor Functionalized Peptoids: Molecular Dynamics and Rate Theory

Scalable solar-energy conversion requires photoactive materials that combine the efficiency of natural photosynthetic systems with the stability and processability needed for practical applications. Achieving reliable charge transport in soft, self-assembled organic materials remains challenging, as structural fluctuations and environmental effects strongly influence charge-transfer (CT) rates. Here, we present a broadly applicable computational framework for evaluating CT rates in the condensed phase, combining Fermi’s golden rule rate theory with inputs from all-atom molecular dynamics (MD) simulations and first-principles electronic-structure calculations. The approach does not rely on system-specific parametrization and is applicable to a wide range of soft and disordered materials. We demonstrate the applicability and usefulness of the framework on redox-active peptoids functionalized with iron–porphyrin (Fe–P) complexes, a bioinspired platform with programmable donor–acceptor units and tunable three-dimensional organization. The calculated CT rates exhibit strong sensitivity to molecular conformation, with variations spanning several orders of magnitude. This dependence is shown to arise from the pronounced variation in diabatic electronic coupling with the relative orientations and separations of the Fe–P complexes across the conformational ensemble. The framework provides a consistent route for connecting atomistic structure to CT kinetics in the condensed phase and enables analysis of structure–rate relations in organic semiconducting systems.

Charge transfer↗

eReaxFF force field development for BaZr 0.8 Y 0.2 O 3-δ solid oxide electrolysis cells applications

The use of solid-oxide materials in electrocatalysis applications, especially in hydrogen-evolution reactions, is promising. However, further improvements are warranted to overcome the fundamental bottlenecks to enhancing the performance of solid-oxide electrolysis cells (SOECs), which is directly linked to the more-refined fundamental understanding of complex physical and chemical phenomena and mass exchanges that take place at the surfaces and in the bulk of electrocatalysis materials. Here, we developed an eReaxFF force field for barium zirconate doped with 20 mol% of yttrium, BaZr 0.8 Y 0.2 O 3-δ (BZY20) to enable a systematic, large-length-scale, and longer-timescale atomistic simulation of solid-oxide electrocatalysis for hydrogen generation. All parameters for the eReaxFF were optimized to reproduce quantum-mechanical (QM) calculations on relevant condensed phase and cluster systems describing oxygen vacancies, vacancy migrations, electron localization, water adsorption, water splitting, and hydrogen generation on the surfaces of the BZY20 solid oxide. Using the developed force field, we performed both zero-voltage (excess electrons absent) and non-zero-voltage (excess electrons present) molecular dynamics simulations to observe water adsorption, water splitting, proton migration, oxygen-vacancy migrations, and eventual hydrogen-production reactions. Based on investigations offered in the present study, we conclude that the eReaxFF force field-based approach can enable computationally efficient simulations for electron conductivity, electron leakage, and other non-zero-voltage effects on the solid oxide materials using the explicit-electron concept. Moreover, we demonstrate how the eReaxFF force field-based atomistic-simulation approach can enhance our understanding of processes in SOEC applications and potentially other renewable-energy applications.

08 HYDROGEN↗

From Structured Solvents to Hybrid Materials (SS2HM) for Chemically Selective Capture and Electromagnetic Release of CO 2 : Mechanisms, Stability and Interfaces (Final Report)

The goal of this research program was to develop high capacity sorbents amenable for alternative regeneration approaches for direct air capture (DAC) of CO 2 . In particular, the research aimed to develop an understanding of CO 2 binding mechanism, thermal and oxidative stability, and regeneration energetics of functionalized ionic liquids (ILs), deep eutectic solvents (DESs), and porous materials. ILs and DESs are high-dielectric solvents with structural tunability that permits the rational-design for energy-efficient regeneration approaches based on electromagnetic (EM) field and moisture-swing. By further incorporating these solvents into polymeric capsules and other structural supports, multi-scale interfaces for targeted CO 2 and energy transfers were achieved. Aspects related to CO 2 capacity, selectivity, stability, dielectric properties, and binding energies were examined through experimental and computational design to identify molecular descriptors to inform future design of structured solvents and hybrid materials for DAC. Enclosed final report details the key findings, science advancements, and workforce development efforts from this project.

36 MATERIALS SCIENCE↗

3D printed optimized electrodes for electrochemical flow reactors

Recent advances in 3D printing have enabled the manufacture of porous electrodes which cannot be machined using traditional methods. With micron-scale precision, the pore structure of an electrode can now be designed for optimal energy efficiency, and a 3D printed electrode is not limited to a single uniform porosity. As these electrodes scale in size, however, the total number of possible pore designs can be intractable; choosing an appropriate pore distribution manually can be a complex task. To address this challenge, we adopt an inverse design approach. Using physics-based models, the electrode structure is optimized to minimize power losses in a flow reactor. The computer-generated structure is then printed and benchmarked against homogeneous porosity electrodes. We show how an optimized electrode decreases the power requirements by 16% compared to the best-case homogeneous porosity. Future work could apply this approach to flow batteries, electrolyzers, and fuel cells to accelerate their design and implementation.

25 ENERGY STORAGE↗

A Generation-Storage Coordination Dispatch Strategy for Power System Based on Causal Reinforcement Learning

In the backdrop of global energy transformation, power systems integrating high proportions of renewable energy sources are facing unprecedented challenges in operational stability and dispatch efficiency. To address these challenges, this study introduces a generation-storage coordination real-time dispatch strategy based on Causal Power System Dynamic Reinforcement Learning (CPSDRL). Diverging from traditional reinforcement learning approaches, CPSDRL innovatively incorporates causal inference within the state prediction model - the crux of model-based reinforcement learning - thereby establishing the Power Causal Dynamic Model (PCDM). Assisted by the prior knowledge of power systems, the model significantly enhances prediction accuracy and reliability through a two-stage training process. Utilizing PCDM, this study further applies a direct policy search algorithm to optimize the real-time dispatch strategy. Experimental results indicate that the proposed method improves the stability of generation-storage coordination real-time dispatch and exhibits competitive advantages in sample efficiency and computational speed, compared to traditional model-based and model-free reinforcement learning algorithms. This method is expected to enhance the practicality and adaptability of causal reinforcement learning techniques in power system scheduling and control.

causal reinforcement learning↗

Scattering wave packets of hadrons in gauge theories: Preparation on a quantum computer

Quantum simulation holds promise of enabling a complete description of high-energy scattering processes rooted in gauge theories of the Standard Model. A first step in such simulations is preparation of interacting hadronic wave packets. To create the wave packets, one typically resorts to adiabatic evolution to bridge between wave packets in the free theory and those in the interacting theory, rendering the simulation resource intensive. In this work, we construct a wave-packet creation operator directly in the interacting theory to circumvent adiabatic evolution, taking advantage of resource-efficient schemes for ground-state preparation, such as variational quantum eigensolvers. By means of an ansatz for bound mesonic excitations in confining gauge theories, which is subsequently optimized using classical or quantum methods, we show that interacting mesonic wave packets can be created efficiently and accurately using digital quantum algorithms that we develop. Specifically, we obtain high-fidelity mesonic wave packets in the Z 2 and U(1) lattice gauge theories coupled to fermionic matter in 1+1 dimensions. Our method is applicable to both perturbative and non-perturbative regimes of couplings. The wave-packet creation circuit for the case of the Z 2 lattice gauge theory is built and implemented on the Quantinuum H1-1 trapped-ion quantum computer using 13 qubits and up to 308 entangling gates. The fidelities agree well with classical benchmark calculations after employing a simple symmetry-based noise-mitigation technique. This work serves as a step toward quantum computing scattering processes in quantum chromodynamics.

97 MATHEMATICS AND COMPUTING↗

AutoSourceID-Classifier: Star-galaxy classification using a convolutional neural network with spatial information

Aims.Traditional star-galaxy classification techniques often rely on feature estimation from catalogs, a process susceptible to introducing inaccuracies, thereby potentially jeopardizing the classification’s reliability. Certain galaxies, especially those not manifesting as extended sources, can be misclassified when their shape parameters and flux solely drive the inference. We aim to create a robust and accurate classification network for identifying stars and galaxies directly from astronomical images. Methods.The AutoSourceID-Classifier (ASID-C) algorithm developed for this work uses 32x32 pixel single filter band source cutouts generated by the previously developed AutoSourceID-Light (ASID-L) code. By leveraging convolutional neural networks (CNN) and additional information about the source position within the full-field image, ASID-C aims to accurately classify all stars and galaxies within a survey. Subsequently, we employed a modified Platt scaling calibration for the output of the CNN, ensuring that the derived probabilities were effectively calibrated, delivering precise and reliable results. Results.We show that ASID-C, trained on MeerLICHT telescope images and using the Dark Energy Camera Legacy Survey (DECaLS) morphological classification, is a robust classifier and outperforms similar codes such as SourceExtractor. To facilitate a rigorous comparison, we also trained an eXtreme Gradient Boosting (XGBoost) model on tabular features extracted by SourceExtractor. While this XGBoost model approaches ASID-C in performance metrics, it does not offer the computational efficiency and reduced error propagation inherent in ASID-C’s direct image-based classification approach. ASID-C excels in low signal-to-noise ratio and crowded scenarios, potentially aiding in transient host identification and advancing deep-sky astronomy.

Astronomy & Astrophysics↗

Performance Improvements of the Griffin Solvers in FY24

The Griffin code is a MOOSE-based reactor physics application jointly developed by Idaho National Laboratory and Argonne National Laboratory under the Department of Energy Office of Nuclear Energy Nuclear Energy Advanced Modeling and Simulation Program. This fiscal year, we have made significant efforts to improve the performance of transport solver options and cross-section generation for the efficient use of Griffin in advanced reactor applications. For the HFEM-PN solver, the residual evaluations of HFEM kernels were optimized by utilizing the pre- computed averaged cross sections for individual elements. Numerical integration involving the evaluation of basis functions at quadrature points was bypassed by facilitating precomputed element mass matrices for response matrices. Red-black iterations were improved by introducing a new generalized minimum residual based solver. The memory usage of response matrix storage was significantly reduced by applying basis function rotations on interfaces and calculating volumetric odd-parity moments on the fly. Additionally, the adjoint flux and transient calculation capabilities of the HFEM-PN solver were successfully implemented and verified using the TWIGL benchmark problem. For the DFEM-SN solver, memory footprint and computation time were significantly reduced by not treating angular flux vectors as the MOOSE nonlinear system vectors. Specifically for IQS, scalar adjoint weighting was introduced to further eliminate angular adjoint flux storage in the MOOSE auxiliary system. It was demonstrated through the three-dimensional Advanced Burner Test Reactor core problem that the memory usage for transient calculations with the IQS method was reduced by over 7.5× compared to before the optimizations. For the self-shielding application programming interface, a new double-heterogeneity treatment method, named the Bell Function-Based Analytic Two-Region Slowing Down Method, was developed to efficiently flux-volume homogenize TRISO particles with the matrix. Additionally, optimizations were made to hyper- fine group (HFG) slowing down calculations by pretabulating collision probability coefficients and grouping isotopes, significantly reducing the computational time for calculating scattering sources per HFG. Lastly, the pin power reconstruction module was extended to account for temporal behavior in a microreactor analysis problem, specifically for a control drum transient. Verification tests for each of these improvements demonstrated significant performance enhancements and memory reduction.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensitivity analysis of design parameters in RowWise borehole layout for ground heat exchangers

Ground source heat pump systems offer a promising pathway toward energy-efficient building heating and cooling. The performance and cost-effectiveness of these systems heavily depend on the design of the ground heat exchanger (GHE), particularly the spatial placement of boreholes used in large commercial buildings. Among various borefield layout strategies, the RowWise approach generates and optimizes borehole configurations within irregular polygonal land boundaries, providing land-use efficiency and installation flexibility. While multiple geometric design parameters constrain the RowWise layout, their influence on the system thermal performance and total drilling requirements remains unclear. Thus, this study presents a sensitivity analysis of key design parameters influencing the RowWise layout of vertical borehole GHEs, including the perimeter spacing ratio, borehole spacing, borehole field rotation angle, and borehole length. GHEDesigner and Ray Tune are employed to generate and assess different RowWise configurations. A series of parametric simulations was conducted to quantify the impact of each parameter on geometric distribution, economic cost, and computational speed. The results provide critical insights into the sensitivity and relative importance of different design parameters of RowWise method, offering practical guidance for designers aiming to optimize GHE layouts by balancing thermal efficiency, land constraints, and economic feasibility.

Xu, Dikai [Purdue University]↗

Developing tools and process controls to manufacture energy-efficient powders for additive manufacturing feedstocks

Traditionally, metal powders have been produced through methods such as grinding, atomization, and electrolysis. In contrast to these techniques, Metal Powder Works, Inc. has pioneered a methodology based on metal cutting. This innovative approach utilizes a vibrating cutting tool to machine metal particles, in the form of chips, from a workpiece. This technique allows for control of powder particle size, morphology, and avoids any thermally induced material changes. This collaboration aims to elucidate metal cutting characteristics and assess performance on tough materials like Inconel alloys. Computational models, using FEA and SPH techniques, will be developed initially, focusing on aluminum alloy (Al 7075- T6) for studying mesh sensitivity, cutting forces, and chip morphology.

36 MATERIALS SCIENCE↗

District heating utilizing waste heat of a data center: High-temperature heat pumps

Data centers are energy-intensive facilities with substantial low-grade waste heat. High-temperature heat pumps can be critical in boosting the data center’s waste heat for district heating, improving the system-level energy efficiency of data centers, and reducing CO 2 emissions in district heating. This study built thermodynamic models to assess high-temperature heat pumps with six configurations using low global warming potential refrigerants to supply heat up to 120 °C. The heat pump configurations include single-stage or two-stage cycles with advanced components, such as internal heat exchanger, economizer, flash tank, or parallel compressor. The refrigerants include R1234ze(Z), R1233ed(E), R1224yd(Z), R600, and R600a, and R245fa is used as a reference. A case study was carried out to recover the waste heat from the Frontier high-performance computing data center and provide hot water for district heating at the US Department of Energy’s Oak Ridge National Laboratory campus. The optimized performance of high-temperature heat pumps is characterized with various effectiveness of internal heat exchangers, and the operating parameters of economizer or flash tank, as well as their combination. The results show that the configurations of two-stage cycles with internal heat exchanger + flash tank and internal heat exchanger + economizer/parallel-compressor provide the highest coefficient of performance under scenarios of the maximum allowable value and a fixed value (0.3) of the internal heat exchangers’ effectiveness, respectively. R1234ze(Z) and R600a are the most promising refrigerants, considering trade-offs between the coefficient of performance and the volumetric heating capacity. The single-stage cycle with internal heat exchanger + economizer/parallel-compressor using R1234ze(Z) is recommended for utilizing Fronter’s waste heat in district heating. A one mega-watt high-temperature heat pump will reduce 33,100–33,200 metric tons of CO2 emission annually, corresponding to 85.4 %–85.6 % of equivalent CO2 emissions from natural gas boilers. Here, this study provides good guidelines for designing and deploying high-temperature heat pumps to support sustainable data centers and decarbonize district heating in the US.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SymbolNet: neural symbolic regression with adaptive dynamic pruning for compression

Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).

Tsoi, Ho Fung (ORCID:0000000225502184)↗

A CHIL Validation of Machine Learning-Assisted Methods for Real-Time Controls of Solar PV for Grid Services

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been proposed; however, these technologies lack comprehensive validation under real-world application scenarios. This paper addresses this gap by designing and developing a controller-hardware-in-the-loop framework to evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. Simulation results indicate the superior performance of an ML-based approach compared to the conventional reference-control grouping-based approach, showcasing its potential to support grid stability and operational efficiency.

closed-loop validation↗

Ab Initio Many Body Quantum Embedding and Local Correlation in Crystalline Materials using Interpolative Separable Density Fitting

We present an efficient implementation of ab initio many-body quantum embedding and local correlation methods for infinite periodic systems through translational symmetry adapted interpolative separable density fitting, an approach which reduces the scaling of the calculations to only linear with the number of k-points. Employing this methodology, we compute correlated ground-state coupled cluster energies within density matrix embedding and local natural orbital correlation frameworks for both weakly and strongly correlated solids, using up to 1000 k-points. By extrapolating the local correlation domains and k-point sampling we further obtain estimates of the full coupled cluster with singles, doubles, and perturbative triples ground-state energies in the thermodynamic limit.

Chemical Physics (physics.chem-ph)↗