Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer systems performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Insights into Native Single-Atom Electrocatalyst Site Structures

Single-atom electrocatalysts consisting of metal atoms embedded in a carbon matrix are promising next-generation catalysts for green hydrogen production and utilization, CO2 reduction, low-temperature CO oxidation, ammonia production, plastic decomposition, and electrochemical energy storage. The origins of activity and stability for the single-atom sites are still debatable, however, because of constrained insights into their local structure resulting from idealized models and experiments derived from a large number of individual sites. Insights into structural variations around single atomic sites are therefore critical for the continued development of these next-generation catalysts. While electron microscopy commonly provides atomic-scale information about these materials, the beam sensitivity of individual sites makes structural determination by conventional low-voltage (60 keV) techniques challenging. Here, we introduce ultralow-voltage electron ptychography, performed at 30 keV, that enables determination of the lattice structure around individual metal sites in a well-defined single-atom electrocatalyst system while essentially eliminating knock-on structural modifications. Pairing these atomic-scale, site-specific measurements with computational methods will broaden our understanding of the activity and stability of these materials, which will accelerate the development of the next generation of catalysts.

Zachman, Michael [ORNL] (ORCID:0000000319101357)↗

Data-driven analysis to understand GPU hardware resource usage of optimizations

With heterogeneous systems, the number of GPUs per chip increases to provide computational capabilities for solving science at a nanoscopic scale. However, low utilization for single GPUs defies the need to invest more money in expensive accelerators. Although related work develops optimizations to improve application performance, none studies how these optimizations impact hardware resource usage or average GPU utilization. Here, this paper takes a data-driven analysis approach in addressing this gap by (1) characterizing how hardware resource usage affects device utilization, execution time, or both, (2) presenting a multiobjective metric to identify important application-device interactions that can be optimized to improve device utilization and application performance jointly, (3) studying hardware resource usage behaviors of several optimizations for a benchmark application, and finally (4) identifying optimization opportunities for several scientific proxy applications based on their hardware resource usage behaviors. Furthermore, we demonstrate the applicability of our methodology by applying the identified optimizations to a proxy application, which improves the execution time, device utilization, and power consumption by up to 29.6%, 5.3% and 26.5% respectively.

Computer science↗

Quantum for Energy Systems and Technologies

Quantum Information Science (QIS) is expected to profoundly change the practice of science and engineering in the coming decades. QIS technology exploits quantum phenomena for performing tasks that are impossible to do today and is a rapidly progressing field, fueled by large investments from the private sector and governments. Its importance to the U.S. economy and national security is underscored by the National Quantum Initiative Act (NQIA) passed in December 2018, which creates a coordinated multiagency program to support research and training in QIS. After the NIQA signed into law, NETL has launched an initiative to apply QIS to problems encountered in energy technology development. In Jan. 2020, an Quantum for Energy Systems & Technologies (QUEST) working group was formed to establish a workforce capable of developing, reviewing, managing, and advising on QIS-related technologies for NETL and FECM. Since then, the QUEST team have been working on developing quantum sensing technology and performing quantum computing to solve energy-related problems. This poster summarized the QUEST team’s activities & accomplishments on QIS targeting energy-related applications.

Paudel, Hari P.↗

Identifying Decoherence Mechanisms in Superconducting Qubits through Advanced Materials Characterization

Although superconducting qubits have emerged as a leading technology platform for quantum computing through large improvements in device coherence times and gate fidelity in recent years, the presence of defects and impurities at the interfaces and surfaces in the constituent materials continue to limit performance and serve as a critical barrier in achieving scalable quantum systems. Understanding and eliminating these sources of quantum decoherence in superconducting qubit devices requires dedicated studies aimed at establishing robust structure-property relationships that will enable researchers to target and eliminate defects strategically. As part of the Superconducting Materials and Systems (SQMS) center, we have extensively employed state-of-the-art materials characterization techniques, including scanning/transmission electron microscopy, secondary ion mass spectrometry, atom probe tomography, x-ray diffraction, and x-ray photoelectron spectroscopy in conjunction with device measurements to elucidate such relationships. In this talk, I will discuss some of our recent findings, including linking atomic defects to microwave loss in surface oxides, linking impurities in the Josephson Junction to qubit parameters, and linking low temperature precipitates to device performance. By applying these insights, we have been able to strategically develop and implement mitigation strategies for reliable fabrication of high coherence superconducting qubits.

Murthy, A. [Fermilab] (ORCID:0000000176776866)↗

Identifying Decoherence Mechanisms in Superconducting Qubits through Advanced Materials Characterization

Although superconducting qubits have emerged as a leading technology platform for quantum computing through large improvements in device coherence times and gate fidelity in recent years, the presence of defects and impurities at the interfaces and surfaces in the constituent materials continue to limit performance and serve as a critical barrier in achieving scalable quantum systems. Understanding and eliminating these sources of quantum decoherence in superconducting qubit devices requires dedicated studies aimed at establishing robust structure-property relationships that will enable researchers to target and eliminate defects strategically. As part of the Superconducting Materials and Systems (SQMS) center, we have extensively employed state-of-the-art materials characterization techniques, including scanning/transmission electron microscopy, secondary ion mass spectrometry, atom probe tomography, x-ray diffraction, and x-ray photoelectron spectroscopy in conjunction with device measurements to elucidate such relationships. In this talk, I will discuss some of our recent findings, including linking atomic defects to microwave loss in surface oxides, linking impurities in the Josephson Junction to qubit parameters, and linking low temperature precipitates to device performance. By applying these insights, we have been able to strategically develop and implement mitigation strategies for reliable fabrication of high coherence superconducting qubits.

Murthy, A. [Fermilab] (ORCID:0000000176776866)↗

Rapid Evaluation Framework for the CMIP7 Assessment Fast Track

As Earth system models (ESMs) grow in complexity and in volume of output data, there is an increasing need for rapid, comprehensive evaluation of their scientific performance. The upcoming Assessment Fast Track for the Seventh Phase of the Coupled Model Intercomparison Project (CMIP7) will require expeditious response for model analyses designed to inform and drive integrated Earth system assessments. To meet this challenge, the Rapid Evaluation Framework (REF), a community-driven platform for benchmarking and performance assessment of ESMs, was designed and developed. The initial implementation of the REF, constructed to meet the near-term needs of the CMIP7 Assessment Fast Track, builds upon four disparate community evaluation and benchmarking tools that are coupled together using the Coordinated Model Evaluation Capabilities (CMEC) framework. The REF runs within a containerized workflow for portability and reproducibility and is aimed at generating and organizing diagnostics covering a variety of model variables. The REF leverages well documented observational datasets to provide assessments of model fidelity across a collection of diagnostics. All diagnostics were identified and selected with community involvement and consultation. Operational integration with the Earth System Grid Federation (ESGF) will permit automated execution of the REF for selected diagnostics as soon as model output data are published on ESGF by the originating modeling centers. The REF is designed to be portable across a range of current computational platforms to facilitate use by modeling centers for assessing the evolution of model versions or gauging the relative performance of CMIP simulations before being published on ESGF. When integrated into production simulation workflows, results from the REF provide immediate quantitative feedback that allows model developers and scientists to quickly identify model biases and performance issues. After the REF is released to the community, its subsequent development and support will be prioritized by an international consortium of scientists and engineers, enabling a broader impact across Earth science disciplines. For instance, the REF will facilitate improvements to models and will enhance confidence in model projections through process-based selection of models based on their performance with respect to observations. Production of reproducible diagnostics and community-based assessments are key features of the REF. Furthermore, providing interoperability with existing evaluation packages assures that contributions from previous community efforts will be available for use in future model intercomparison projects.

Hoffman, Forrest [ORNL] (ORCID:0000000158024134)↗

Open-source library for performance-portable neutrino reaction rates: Application to neutron star mergers

A realistic and detailed description of neutrinos in binary neutron star (BNS) mergers is essential to build reliable models of such systems. To this end, we present bns_nurates, a novel open-source numerical library designed for the efficient on-the-fly computation of neutrino interactions, with particular focus on regimes relevant to BNS mergers. bns_nurates targets a higher level of accuracy and realism in the implementation of commonly employed reactions by accounting for relevant microphysics effects on the interactions, such as weak magnetism and mean field effects. It also includes the contributions of inelastic neutrino scattering off electrons and positrons and (inverse) nucleon decays. Finally, it offers a way to reconstruct the neutrino distribution function in the framework of moment-based transport schemes. As a first application, we compute both energy-dependent and energy-integrated neutrino emissivities and opacities for conditions extracted from a BNS merger simulation with m1 transport scheme. We find some qualitative differences in the results when considering the impact of the additional relevant reactions and of microphysics effects. For example, neutrino-electron/positron scattering reactions are important for the energy exchange of heavy-type neutrinos as they do not undergo semileptonic charged-current processes, when μ± are not accounted for. Moreover, weak magnetism and mean field effects can significantly modify the contribution of β processes for electron-type (anti)neutrinos, increasing at the same time the importance of (inverse) neutron decays. Here, the improved treatment for the reaction rates also modifies the conditions at which neutrinos decouple from matter in the system, potentially affecting their emission spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

Unification of finite symmetries in the simulation of many-body systems on quantum computers

Symmetry is fundamental in the description and simulation of quantum systems. Leveraging symmetries in classical simulations of many-body quantum systems can result in significant overhead due to the exponentially growing size of some symmetry groups as the number of particles increases. Quantum computers hold the promise of achieving exponential speedup in simulating quantum many-body systems; however, a general method for utilizing symmetries in quantum simulations has not yet been established. In this work, we present a unified framework for incorporating symmetry group transforms on quantum computers to simulate many-body systems. The core of our approach lies in the development of efficient quantum circuits for symmetry-adapted projection onto irreducible representations of a group or pairs of commuting groups. We provide resource estimations for common groups, including the cyclic and permutation groups. Our algorithms demonstrate the capability to prepare coherent superpositions of symmetry-adapted states and to perform quantum evolution across a wide range of models in condensed-matter physics and ab initio electronic structure in quantum chemistry. Specifically, we execute a symmetry-adapted quantum subroutine for small molecules in first-quantization on noisy hardware and demonstrate the emulation of symmetry-adapted quantum phase estimation for preparing coherent superpositions of quantum states in various irreducible representations of a symmetry group. In addition, we present a discussion of open problems regarding treating symmetries in digital quantum simulations of many-body systems, paving the way for future systematic investigations into leveraging symmetries quantumly for practical quantum advantage. The broad applicability and rigorous resource estimation for symmetry transformations make our framework appealing for achieving provable quantum advantage on fault-tolerant quantum computers, especially for symmetry-related properties.

quantum algorithms↗

Differentiable hybrid neural network approach for enhancing reactor dynamics simulations

Reactor dynamics simulations provide essential insights into the time-dependent behavior of nuclear reactors under various operating conditions. However, high-fidelity simulations can be computationally intensive, requiring significant computational resources. Here, to address this challenge, this study employs a differentiable hybrid model that utilizes neural networks as a corrector to enhance the performance of a low-fidelity simulation, aligning its predictions with those of a high-fidelity simulation. Low-fidelity and high-fidelity simulations were obtained by adjusting the mesh size in the System Dynamics Analysis Tool. The differentiable hybrid model was trained in two approaches: time-step-wise and sequence-wise. It was then applied to simulate various transients in a molten salt reactor. Its performance was evaluated by comparing its responses to transients against those of the high-fidelity simulation. An additional approach was performed using a data-driven model to correct the low-fidelity simulation. In comparison, the differentiable hybrid model showed significant improvements in transient prediction, effectively addressing the limitations of the low-fidelity simulations. The results highlighted the robustness of the differentiable hybrid model in both training approaches. It delivered simulations that were at least 3.8 times faster than high-fidelity models. In the time-step-wise approach, it achieved at least a 39% improvement in accuracy. In the sequence-wise approach, it showed at least an 81% accuracy improvement over the full transient. This approach offers a promising path for improving computational efficiency without compromising accuracy in nuclear reactor simulations, making it suitable for real-time digital twin applications.

42 - ENGINEERING↗

Designing Cellular Metal Structures for Thermal Insulation

This project focused on developing topology optimization software to design advanced metal thermal insulators. Initially, solid designs were created that matched the thermal performance of current baseline designs but were significantly heavier. To address this, cellular materials were incorporated, specifically the octet structure which is known for its high strength-to-weight ratio and thermal properties. By leveraging these cellular designs at various densities, superior thermal and mechanical performance was achieved without added weight. This novel approach enhances thermal management and structural integrity under extreme conditions, offering promising advancements for thermal protection systems.

36 MATERIALS SCIENCE↗

Offline Maximizing Minimally Invasive Proper Orthogonal Decomposition for Reduced-Order Modeling of S n Radiation Transport

Deterministic solutions to the Sn radiation transport equation can be computationally expensive to calculate. Reduced-order modeling enables efficient approximation of the full-order model (FOM) solution. We propose a novel method for constructing reduced-order models (ROMs) of the S n radiation transport equation, offline maximizing minimally invasive (OMMI) proper orthogonal decomposition (POD). POD uses the method of snapshots to create a reduced-order basis for constructing an ROM. Minimally invasive POD leverages the sweep infrastructure existing in deterministic transport codes to create a POD-based ROM, even when infeasible by traditional methods. Offline maximizing minimally invasive proper orthogonal decomposition (OMMI-POD) extends minimally invasive POD by performing sweeps offline, therefore maximizing the potential speedup. OMMI-POD does so by creating a library of reduced systems from a training set. This library of reduced systems is then interpolated to provide a rapid approximate solution of the S n radiation transport equation. The model is evaluated on a set of test problems, achieving a low error with a 466 times speedup over the FOM. Also presented is a study of the effect of sampling method on the performance of OMMI-POD, specifically comparing naive uniform sampling to the more accurate and computationally expensive greedy sampling.

97 MATHEMATICS AND COMPUTING↗

2023 Operational Assessment Report: Argonne Leadership Computing Facility

In 2004, the U.S. Department of Energy’s (DOE’s) Advanced Scientific Computing Research (ASCR) program established the Leadership Computing Facility (LCF) with a mission to provide the world’s most advanced computational resources to the open science community. The LCF is a huge investment in the nation’s scientific and technological future, inspired by a growing demand for large-scale computing and its impact on science and engineering. The LCF operates two world-class centers in support of open science at Argonne National Laboratory (Argonne) and at Oak Ridge National Laboratory (Oak Ridge) and deploys diverse machines that are among the most powerful systems in the world today. The LCF ranks among the top U.S. scientific facilities delivering impactful science. The work performed at these centers informs policy decisions and advances innovations in far-reaching areas such as energy assurance, ecological sustainability, and global security. The leadership-class systems at Argonne and Oak Ridge operate around the clock every day of the year. The high level of services these centers provide and the exceptional science they produce justify their existence to the DOE Office of Science and the U.S. Congress

97 MATHEMATICS AND COMPUTING↗

Divide and conquer: Learning chaotic dynamical systems with multistep penalty neural ordinary differential equations

Forecasting high-dimensional dynamical systems is a fundamental challenge in various fields, such as geosciences and engineering. Neural Ordinary Differential Equations (NODEs), which combine the power of neural networks and numerical solvers, have emerged as a promising algorithm for forecasting complex nonlinear dynamical systems. However, classical techniques used for NODE training are ineffective for learning chaotic dynamical systems. In this work, we propose a novel NODE-training approach that allows for robust learning of chaotic dynamical systems. Here, our method addresses the challenges of non-convexity and exploding gradients associated with underlying chaotic dynamics. Training data trajectories from such systems are split into multiple, non-overlapping time windows. In addition to the deviation from the training data, the optimization loss term further penalizes the discontinuities of the predicted trajectory between the time windows. The window size is selected based on the fastest Lyapunov time scale of the system. Multi-step penalty(MP) method is first demonstrated on Lorenz equation, to illustrate how it improves the loss landscape and thereby accelerates the optimization convergence. MP method can optimize chaotic systems in a manner similar to least-squares shadowing with significantly lower computational costs. Our proposed algorithm, denoted the Multistep Penalty NODE, is applied to chaotic systems such as the Kuramoto-Sivashinsky equation, the two-dimensional Kolmogorov flow, and ERA5 reanalysis data for the atmosphere. It is observed that MP-NODE provide viable performance for such chaotic systems, not only for short-term trajectory predictions but also for invariant statistics that are hallmarks of the chaotic nature of these dynamics.

Chaotic dynamical systems↗

Identification of Solid-Electrolyte Interphase Species by Joint Characterization of Li-Ion Battery Chemistry by Mass Spectrometry and Electrochemical Reaction Networks

The formation and stability of the solid-electrolyte interphase (SEI) play central roles in determining the long-term performance and safety of modern electrochemical energy storage systems. Despite decades of research, the SEI’s heterogeneous, dynamic, and multiphase nature has defied comprehensive molecular-level characterization, creating a critical knowledge gap that limits rational battery design. In this work, we introduce a computational−experimental framework that integrates high-throughput quantum chemistry calculations, data-driven electrochemical reaction networks (eCRNs), stochastic algorithms, and laser desorption/ionization Fourier transform ion cyclotron resonance mass spectrometry (LDI-FTICR-MS) to unravel SEI formation in carbonatebased electrolytes without imposing predefined mechanisms. We constructed the most comprehensive eCRN to date, spanning over 10,000 species and 209 million reactions. Through stochastic network analysis, we successfully recovered 27 species that were previously reported in the literature and predicted 28 novel SEI species nearly doubling our scientific knowledge in this area. Each new species was rigorously confirmed through advanced mass spectral analysis of its distinct molecular and isotopic signatures. We kinetically refined the formation pathways for a select set of both previously reported and novel SEI products, revealing kinetically feasible elementary reaction mechanisms with activation barriers below 1 eV. This computational−experimental approach deepens our molecular-level understanding of SEI chemistry by resolving which species form and through which decomposition mechanisms they emerge. Such knowledge provides the foundation necessary to connect electrolyte composition to the resulting SEI components, a critical step toward a more informed electrolyte development in next-generation lithium-based batteries.

25 ENERGY STORAGE↗

Accurate and uncertainty-aware multi-task prediction of HEA properties using prior-guided deep Gaussian processes

Surrogate modeling techniques have become indispensable in accelerating the discovery and optimization of high-entropy alloys (HEAs), especially when integrating computational predictions with sparse experimental observations. This study systematically evaluates the training and testing performance of four prominent surrogate models—conventional Gaussian processes (cGP), Deep Gaussian processes (DGP), encoder-decoder neural networks for multi-output regression and eXtreme Gradient Boosting (XGBoost)—applied to a hybrid dataset of experimental and computational properties of the 8-component HEA system Al-Co-Cr-Cu-Fe-Mn-Ni-V. We specifically assess their capabilities in predicting correlated material properties, including yield strength, hardness, modulus, ultimate tensile strength, elongation, and average hardness under dynamic/quasi-static conditions, alongside auxiliary computational properties. The comparison highlights the strengths of hierarchical deep modeling approaches in handling heteroscedastic, heterotopic, and incomplete data commonly encountered in materials science. Our findings illustrate that combined surrogate models such as DGPs infused with machine-learned priors outperform other surrogates by effectively capturing inter-property correlations and by assimilating prior knowledge. This enhanced predictive accuracy positions the combined surrogate models as powerful tools for robust and data-efficient materials design.

36 MATERIALS SCIENCE↗

Simulate Wind Loading on RUTE SunTracker: Cooperative Research and Development (Final Report)

RUTE produces and delivers efficient, sustainable structural foundation systems for the wind and solar industries, instrumental in the effort to lower the cost of clean electricity and reduce CO2. RUTE's SunTracker is a high-clearance agrivoltaic solar array. The system uses raised cables, providing support with less materials and cost. These cables can provide tracking for increased energy generation and greater revenue, allowing the land beneath to be used for other means simultaneously. RUTE provided data files that NREL will use to perform computational fluid dynamics (CFD) simulations to characterize and predict wind loading on the 6x6 RUTE SunTracker photovoltaic (PV) array. Varying both the panel orientations and wind speeds provides information on how the array can manage a variety of weather conditions. These results, in turn, inform RUTE's design and construction of these arrays.

14 SOLAR ENERGY↗

Hydraulic Conductivity Measurements, Utqiagvik (Barrow), Alaska, 2014

Six individual ice cores were collected from the Barrow Environmental Observatory in Barrow, Alaska, in May of 2013 as part of the Next Generation Ecosystem Experiment (NGEE). Each core was drilled at a different location to varying depths. After drilling, the cores were stored in coolers packed with dry ice and flown to Lawrence Berkeley National Laboratory (LBNL) in Berkeley, CA. 3-dimensional images of the cores were constructed using medical X-ray computed tomography (CT) scanner at 120kV. Hydraulic conductivity samples were extracted from these cores at LBNL Richmond Field Station in Richmond, CA, in February 2014 by cutting 5 to 8 inch segments using a chop saw. Samples were packed individually and stored at -20C freezing temperatures to minimize any changes in structure or loss of ice content prior to analysis. Hydraulic conductivity was determined through falling head tests using a permeameter [ELE International, Model #: K-770B] (Appendix A). Samples were placed in a latex membrane via a membrane stretcher while frozen. Use of a membrane stretcher made the membranes easier to secure and minimized contact with the sample. A clear polycarbonate sleeve, fabricated with a stainless steel ring at the bottom to keep the sleeve from floating, was placed around the sample inside the permeameter to minimize deformation during analysis. The permeameter was filled with water and 1.0 PSI of air was applied for confining pressure during sample defrost. Outflow valves were left open to allow for incremental thawing and samples were left to thaw for approximately 12 hours. After approximately 12 hours of thaw, initial falling head tests were performed. When the flow was significantly too fast or too slow, the analysis was stopped and the burette size was adjusted accordingly (i.e. a larger diameter burette was used for flows that were faster than desired or a smaller diameter burette was used for flows that were slower than desired). Two to four measurements were collected on each sample and collection stopped when the applied head load exceeded 25% change from the original load. Analyses were performed between 2 to 3 times for each sample. The final hydraulic conductivity calculations were computed using methodology of Das et al., 1985.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗