Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

Manufacturability and performance response of developed Ni-based alloys for reactor environments

Under the Advanced Materials and Manufacturing Technologies program, two Ni-based alloys fabricated by laser powder bed fusion (LPBF) at Oak Ridge National Laboratory and by laser powder directed energy deposition (LP-DED) at Idaho National Laboratory have been evaluated: γ′-strengthened Haynes 282 and solution-strengthened Inconel 625. Printing defects were observed in the LPBF 282 alloy fabricated using a Renishaw AM250 machine, likely due to particle spattering during printing. For both the LPBF and LP-DED 282 alloys, annealing at 1,180°C for 1 h resulted in partial recrystallization and a bimodal grain distribution, and subsequent aging for 4 h at 800°C led to the formation of a high density of nano-size γ′-strengthening precipitates. Creep testing performed at 750°C revealed lower creep life and ductility for the LPBF and LP-DED 282 compared with wrought 282. X-ray computed tomography combined with optical and scanning electron microscopy microstructural characterization revealed crack formation in the LPBF 282 alloy during creep testing, initiated either from printing defects or from creep cavitation at grain boundaries. Creep cavitation and cracking at grain boundaries were also observed for the LP-DED 282 alloy, and optimization of the printing strategies and postheat treatments will be needed to fabricate high-performance 282 components for high-temperature nuclear applications.

36 MATERIALS SCIENCE↗

Sub-microsecond Transformers for Jet Tagging on FPGAs

We present the first sub-microsecond transformer implementation on an FPGA achieving competitive performance for state-of-the-art high-energy physics benchmarks. Transformers have shown exceptional performance on multiple tasks in modern machine learning applications, including jet tagging at the CERN Large Hadron Collider (LHC). However, their computational complexity prohibits use in real-time applications, such as the hardware trigger system of the collider experiments up until now. In this work, we demonstrate the first application of transformers for jet tagging on FPGAs, achieving $\mathcal{O}(100)$ nanosecond latency with superior performance compared to alternative baseline models. We leverage high-granularity quantization and distributed arithmetic optimization to fit the entire transformer model on a single FPGA, achieving the required throughput and latency. Furthermore, we add multi-head attention and linear attention support to hls4ml, making our work accessible to the broader fast machine learning community. This work advances the next-generation trigger systems for the High Luminosity LHC, enabling the use of transformers for real-time applications in high-energy physics and beyond.

Laatu, Lauri [Imperial Coll., London]↗

Computing Sparse Tensor Decompositions via Chapel and C++/MPI Interoperability without Intermediate I/O

We extend an existing approach for efficient use of shared mapped memory across Chapel and C++ for graph data stored as 1-D arrays to sparse tensor data stored using a combination of 2-D and 1-D arrays. We describe the specific extensions that provide use of shared mapped memory tensor data for a particular C++ tensor decomposition tool called GentenMPI. We then demonstrate our approach on several real-world datasets, providing timing results that illustrate minimal overhead incurred using this approach. Finally, we extend our work to improve memory usage and provide convenient random access to sparse shared mapped memory tensor elements in Chapel, while still being capable of leveraging high performance implementations of tensor algorithms in C++.

97 MATHEMATICS AND COMPUTING↗

Ultra-thick three-dimensional interpenetrating graphene electrode architectures for high volumetric density energy storage

For electrochemical energy storage, increasing the electrode thickness is an effective approach to achieving higher energy density from a given material. However, this often compromises ion transport, leading to diminished performance. Here, in this study, we present a novel platform for fabricating complex 3D interpenetrating electrode structures via photo-polymerization 3D printing, integrated with computational structural optimization for energy storage. The platform employs an acrylate resin system infused with graphene oxide (GO), enabling high-fidelity printing of optimized porous structures and facilitating efficient electron and ion transport in ultra-thick electrodes. The optimized 3D layouts substantially enhance energy and power densities compared to conventional configurations, ensuring superior material utilization and minimal ohmic losses. Supercapacitors fabricated using this approach achieved an exceptional energy density of 4.7 Wh L−1 at a power density of 1689.0 W L−1, surpassing traditional designs. This work underscores the transformative role of structural optimization in advancing electrochemical performance and establishes a versatile pathway for developing next-generation energy storage systems with exceptional efficiency and functionality.

Wang, Zhen [University of California, Berkeley, CA↗

Efficient and flexible multirate temporal adaptivity

In this work we present two new families of multirate time step adaptivity controllers, that are designed to work with embedded multirate infinitesimal (MRI) time integration methods for adapting time steps when solving problems with multiple time scales. We compare these controllers against competing approaches on two benchmark problems, showing that the proposed methods offer dramatically improved performance and flexibility. The combination of embedded MRI methods and the proposed controllers enable adaptive simulations of problems with a potentially arbitrary number of time scales, achieving high accuracy while maintaining low computational cost. Additionally, we introduce a new set of embeddings for the family of explicit multirate exponential Runge–Kutta (MERK) methods of orders 2 through 5, resulting in the first-ever fifth-order embedded MRI method. Finally, we compare the performance of a wide range of embedded MRI methods on our benchmark problems to provide guidance on how to select an appropriate MRI method and multirate controller.

97 MATHEMATICS AND COMPUTING↗

Three Birds with One Stone: Improving Performance, Convergence, and System Throughput with NEST

Variational quantum algorithms (VQAs) have the potential to demonstrate quantum utility on near-term quantum computers. However, these algorithms often get executed on the highest-fidelity qubits and computers to achieve the best performance, causing low system throughput. Recent efforts have shown that VQAs can be run on low-fidelity qubits initially and high-fidelity qubits later on to still achieve good performance. We take this effort forward and show that carefully varying the qubit fidelity map of the VQA over its execution using our technique, Nest, does not just (1) improve performance (i.e., help achieve close to optimal results), but also (2) lead to faster convergence. We also use Nest to co-locate multiple VQAs concurrently on the same computer, thus (3) increasing the system throughput, and therefore, balancing and optimizing three conflicting metrics simultaneously.

qaoa↗

Differentiable multiphase flow model for physics-informed machine learning in reservoir pressure management

Accurate subsurface reservoir pressure control is extremely challenging due to geological heterogeneity and multiphase fluid-flow dynamics. Predicting behavior in this setting relies on high-fidelity physics-based simulations that are computationally expensive. Yet, the uncertain, heterogeneous properties that control these flows make it necessary to perform many of these expensive simulations, which is often prohibitive. To address these challenges, we introduce a physics-informed machine learning workflow that couples a fully differentiable multiphase flow simulator, which is implemented in the DPFEHM framework with a convolutional neural network (CNN). The CNN learns to predict fluid extraction rates from heterogeneous permeability fields to enforce pressure limits at critical reservoir locations. By incorporating transient multiphase flow physics into the training process, our method enables more practical and accurate predictions for realistic injection-extraction scenarios compared to previous works. To speed up training, we pretrain the model on single-phase, steady-state simulations and then finetune it on full multiphase scenarios, which dramatically reduces the computational cost. We demonstrate that high-accuracy training can be achieved with fewer than three thousand full-physics multiphase flow simulations – compared to previous estimates requiring up to ten million. This drastic reduction in the number of simulations is achieved by leveraging transfer learning from much less expensive single phase simulations.

25 ENERGY STORAGE↗

PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU

This paper introduces Pandora, a parallel algorithm for computing dendrograms, the hierarchical cluster trees for single linkage clustering (SLC). Current parallel approaches construct dendrograms by partitioning a minimum spanning tree and removing edges. However, they struggle with skewed, hard-to-parallelize real-world dendrograms. Consequently, computing dendrograms is the sequential bottleneck in HDBSCAN*[21], a popular SLC variant. Pandora uses recursive tree contraction to address this limitation. Pandora contracts nodes to construct progressively smaller trees. It computes the smallest contracted dendrogram and expands it by inserting contracted edges. This recursive strategy is highly parallel, skew-independent, work-optimal, and well-suited for GPUs and multicores. We develop a performance portable implementation of Pandora in Kokkos[31] and evaluate its performance on multicore CPUs and multi-vendor GPUs (e.g., Nvidia, AMD) for dendrogram construction in HDBSCAN*. Multithreaded Pandora is 2.2x faster than the current best-multithreaded implementation. Our GPU version achieves 6-20x speedup on AMD GPUs and 10-37x on NVIDIA GPUs over multithreaded Pandora. Pandora removes HDBSCAN*’s sequential bottleneck, greatly boosting efficiency, particularly with GPUs.

Sao, Piyush↗

Coefficient-to-Basis Network: a fine-tunable operator learning framework for inverse problems with adaptive discretizations and theoretical guarantees

We propose a Coefficient-to-Basis Network (C2BNet), a novel framework for solving inverse problems within the operator learning paradigm. C2BNet efficiently adapts to different discretizations through fine-tuning, using a pre-trained model to significantly reduce computational cost while maintaining high accuracy. Unlike traditional approaches that require retraining from scratch for new discretizations, our method enables seamless adaptation without sacrificing predictive performance. Furthermore, we establish theoretical approximation and generalization error bounds for C2BNet by exploiting low-dimensional structures in the underlying datasets. Our analysis demonstrates that C2BNet adapts to low-dimensional structures without relying on explicit encoding mechanisms, highlighting its robustness and efficiency. To validate our theoretical findings, we conducted extensive numerical experiments that showcase the superior performance of C2BNet on several inverse problems. The results confirm that C2BNet effectively balances computational efficiency and accuracy, making it a promising tool to solve inverse problems in scientific computing and engineering applications.

97 MATHEMATICS AND COMPUTING↗

Unlocking the potential: machine learning applications in electrocatalyst design for electrochemical hydrogen energy transformation

Machine learning (ML) is rapidly emerging as a pivotal tool in the hydrogen energy industry for the creation and optimization of electrocatalysts, which enhance key electrochemical reactions like the hydrogen evolution reaction (HER), the oxygen evolution reaction (OER), the hydrogen oxidation reaction (HOR), and the oxygen reduction reaction (ORR). This comprehensive review demonstrates how cutting-edge ML techniques are being leveraged in electrocatalyst design to overcome the time-consuming limitations of traditional approaches. ML methods, using experimental data from high-throughput experiments and computational data from simulations such as density functional theory (DFT), readily identify complex correlations between electrocatalyst performance and key material descriptors. Leveraging its unparalleled speed and accuracy, ML has facilitated the discovery of novel candidates and the improvement of known products through its pattern recognition capabilities. This review aims to provide a tailored breakdown of ML applications in a format that is readily accessible to materials scientists. Hence, we comprehensively organize ML-driven research by commonly studied material types for different electrochemical reactions to illustrate how ML adeptly navigates the complex landscape of descriptors for these scenarios. We further highlight ML's critical role in the future discovery and development of electrocatalysts for hydrogen energy transformation. Potential challenges and gaps to fill within this focused domain are also discussed. As a practical guide, we hope this work will bridge the gap between communities and encourage novel paradigms in electrocatalysis research, aiming for more effective and sustainable energy solutions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fabrication, Modeling, and Testing of a Prototype Thermal Energy Storage Containment

Increasing penetration of variable renewable energy resources requires the deployment of energy storage at a range of durations. Long-duration energy storage (LDES) technologies will fulfill the need to firm variable renewable energy resource output year round; lithium-ion batteries are uneconomical at these durations. Thermal energy storage (TES) is one promising technology for LDES applications because of its siting flexibility and ease of scaling. Particle-based TES systems use low-cost solid particles that have higher temperature limits than the molten salts used in traditional concentrated solar power systems. A key component in particle-based TES systems is the containment silo for the high-temperature (>1100 degrees C) particles. This study combined experimental testing and computational modeling methods to design and characterize the performance of a particle containment silo for LDES applications. A laboratory-scale silo prototype was built and validated the congruent transient finite element analysis (FEA) model. The performance of a commercial-scale silo was then characterized using the validated model. The commercial-scale model predicted a storage efficiency above 95% after 5 days of storage with a design storage temperature of 1200 degrees C. Insulation material and concrete temperature limits were considered as well. The validation of the methodology means the FEA model can simulate a range of scenarios for future applications. This work supports the development of a promising LDES technology with implications for grid-scale electrical energy storage, but also for thermal energy storage for industrial process heating applications.

clean energy↗

Q-OPT:Quantum Optimization Toolkit

SF-2025-003 Quantum computing has the potential to solve classical optimization problems. To bring these algorithms into practical use, a comprehensive, high-performance and user-friendly toolkit is essential. The Q-OPT: Quantum Optimization Toolkit is a collection of software tools designed to support complete end-to-end framework for quantum optimization.

Hovland, Paul [Argonne National Laboratory (ANL), ↗

NA-22 Quarterly Report: IST for Extended Deterrence (Q4FY24)

Department of Energy (DOE) scientists have long enjoyed technical collaboration with Japanese colleagues – valuing their technical and scientific prowess in materials science and engineering, large-scale, scientific computing, and many other areas. In particular, NNSA’s laboratories are attractive partners for this effort because they have a heritage of performing high quality, basic research; they work with an awareness of the dual use nature of emerging science and technology; and, importantly, they work with deep national security sensibilities through engagement with NNSA and other national security mission responsibilities. Other DOE laboratories with significant national security bona fides are also significant collaborators. In FY23 a Trilateral Leaders Summit was convened at Camp David. The DOE was given responsibility for “Trilateral National Laboratories Cooperation: The United States, Japan, and the ROK will drive new trilateral cooperation between the U.S. Department of Energy’s National Laboratories and counterpart laboratories—supported by a budget of at least $6 million—to advance knowledge, strengthen scientific collaboration, and spearhead innovation in support of the three countries’ shared interests. Scientists and innovators from the three countries will advance collaborative projects on priority critical and emerging technology areas; potential areas of cooperation include advanced computing, artificial intelligence, materials research, and climate and earthquake modeling among other technology areas.”

36 MATERIALS SCIENCE↗

Deployment of BISON models of fuel restructuring at high burnup and related fission gas behavior in UO 2

This milestone report details the advancements made in fiscal year 2024 under the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program to improve the modeling of fission gas behavior in high burnup UO 2 nuclear fuel in the BISON fuel performance code. As nuclear fuel is pushed to higher burnups, significant microstructural changes occur within the fuel, including the formation of a high burnup structure (HBS) on the pellet rim and a dark zone deeper within the pellet. These regions, characterized by subgrain formation and increased pore densities, have critical implications for fission gas behavior and release, which are not well understood. The modeling capabilities in BISON did not adequately predict these phenomena, leading to an underestimation of fuel restructuring and - potentially - of fission gas release. To address these gaps, this milestone focused on three key objectives: (1) reviewing and assessing Sifgrs's capabilities for low burnup fuel, on which high burnup capabilities rely, (2) validating and expanding HBS fission gas modeling capabilities, including investigating mechanisms for fission gas release from HBS, and (3) expanding Sifgrs to enable modeling of dark zone formation and its effects on fission gas behavior. These objectives were achieved and are described herein. The achievements of this NEAMS milestone are significant for the industry's goal of burnup extension. The improved predictive modeling capabilities for both low- and high-burnup conditions enhance our understanding of fuel performance under both normal operations and transient scenarios. Although goals were reached, future work is necessary to validate these models against experimental data and quantify their accuracy in different conditions. In parallel, mechanistic modeling efforts should continue to extend and refine these capabilities to increase accuracy while reducing reliance on empirical models. This will ensure robust performance across a broader range of conditions.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

Reactive Processing of Furan‐Based Monomers via Frontal Ring‐Opening Metathesis Polymerization for High Performance Materials

Frontal ring-opening metathesis polymerization (FROMP) presents an energy-efficient approach to produce high-performance polymers, typically utilizing norbornene derivatives from Diels–Alder reactions. This study broadens the monomer repertoire for FROMP, incorporating the cycloaddition product of biosourced furan compounds and benzyne, namely 1,4-dihydro-1,4-epoxynaphthalene (HEN) derivatives. A computational screening of Diels–Alder products is conducted, selecting products with resistance to retro-Diels–Alder but also sufficient ring strain to facilitate FROMP. The experiments reveal that varying substituents both modulate the FROMP kinetics and enable the creation of thermoplastic materials characterized by different thermomechanical properties. Moreover, HEN-based crosslinkers are designed to enhance the resulting thermomechanical properties at high temperatures (>200 °C). The versatility of such materials is demonstrated through direct ink writing (DIW) to rapidly produce 3D structures without the need for printed supports. This research significantly extends the range of monomers suitable for FROMP, furthering efficient production of high-performance polymeric materials.

36 MATERIALS SCIENCE↗

ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors

With Moore’s law approaching its end, traditional von Neumann architectures are struggling to keep up with the exceeding performance and memory requirements of artificial intelligence and machine learning algorithms. Unconventional computing approaches such as neuromorphic computing that leverage spiking neural networks (SNNs) to perform computation are gaining traction and seek the paradigm shift necessary to sustain the increasing demands of modern applications. Novel memory technologies, such as resistive RAM (ReRAM), employ a crossbar architecture that possesses the inherent capability of efficiently computing vector-matrix multiplication—a dominant operation in SNNs. The prospect of naturally mapping SNNs to the crossbar structures provides a unique opportunity for achieving a high-performance, power-efficient neuromorphic system. In this work, we present ReSpike, which is a new framework, behavioral simulator, and architectural design based on ReRAM crossbar architectures, enabling modeling and co-design to achieve efficient execution of SNNs. We drive this co-design forward by quantifying the impact that ReRAM cell nonidealities have on the corresponding accuracy of an SNN application.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Improving the modeling of near-wall interphase heat transfer in porous media models of Pebble Bed Reactors

Here, this work aims to improve capabilities for modeling localized effects in porous media models of Pebble Bed Reactors. The wall-channeling effect is the primary local phenomenon of interest in a PBR, where the presence of the reflector wall disrupts the pebble packing, causing the pebbles near the wall to pack less efficiently and creating large void regions. Accurate modeling of the near-wall region is important as it will affect core bypass flow and temperature predictions. Porous media models are commonly used for design scoping and plant-level simulations of PBRs. Although these models have some capabilities to model the near-wall region, the correlations that are available in porous media codes are often inaccurate when a multi-region model is used to discretize the near-wall region. This work employs a high-to-low analysis to study the accuracy of available interphase heat transfer closures. NekRS, a spectral element computational fluid dynamics code, is used to perform Large Eddy Simulations. These LES simulation results are compared to porous media model results from the Pronghorn porous media code. The friction term of the KTA drag closure is first improved, reducing the error in the prediction of the near-wall velocity from over 50% to less than 5%. This is combined with improvements to the form term from previous works to produce a drag closure that is capable of accurately modeling the wall-channeling effect across a variety of flow conditions. The Nusselt number predictions of several heat transfer correlations are compared to the high-fidelity results where it is found that the KTA heat transfer correlation is capable of accurately predicting the local Nusselt numbers that were determined in the high-fidelity simulation. Comparison of the radial solid temperature profiles, however, reveal discrepancies between NekRS and Pronghorn. It is discovered that the implementation of the interphase heat transfer coefficient that exists in many current porous media codes is not valid when local porosities are modeled. Instead, it is suggested that the interphase heat transfer coefficient should be dependent on the local porosity, the Nusselt number, and the local solid surface-to-volume ratio. Implementation of this change produces improvement in the agreement between the results obtained by NekRS and Pronghorn while using the KTA heat transfer correlation.

interphase heat transfer↗