Search NASA⌕ Search

SEARCH · Search NASA

Results for “High Performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Accelerating detector simulations with Celeritas: Profiling and performance optimizations

Celeritas is a GPU-optimized Monte Carlo (MC) particle transport code designed to meet the growing computational demands of next-generation high energy physics (HEP) experiments. It provides efficient simulation of electromagnetic (EM) physics processes in complex geometries with magnetic fields, detector hit scoring, and seamless integration into Geant4-driven applications to offload EM physics to GPUs. Recent efforts have focused on performance optimizations and expanding profiling capabilities. This paper presents some key advancements, including the integration of the Perfetto system profiling tool for detailed performance analysis and the development of track-sorting methods to improve computational efficiency.

Lund, Amanda [Argonne National Laboratory (ANL)]↗

Gaia: segmented germanium detector for high-energy X-ray fluorescence and spectroscopic imaging

We present Gaia, a monolithic array of 96 high-purity germanium pixel detectors integrated with a custom low-noise application-specific integrated circuit (ASIC) and a field-programmable gate array (FPGA)-based data acquisition system. The sensor operates at ∼100 K using a commercial closed-cycle cryocooler, with the in-vacuum electronics thermally isolated from the cold finger to ensure thermal stability. The system demonstrates an average energy resolution of 711 eV at 122 keV, measured using a 57 Co source, and 253 eV at 5.89 keV, measured with 55 Fe across all channels. The readout architecture incorporates a high-performance FPGA paired with a dual-core ARM processor, forming a complete embedded Linux-based computing platform. Communication between the processor and FPGA is handled via memory-mapped I/O, and data are streamed over high-speed gigabit Ethernet. A full-scale 384-pixel Gaia detector, based on this 96-element module, is currently under fabrication.

36 MATERIALS SCIENCE↗

SPIKANs: separable physics-informed Kolmogorov–Arnold networks

Physics-Informed Neural Networks (PINNs) have emerged as a promising method for solving partial differential equations (PDEs) in scientific computing. While PINNs typically use multilayer perceptrons (MLPs) as their underlying architecture, recent advancements have explored alternative neural network structures. One such innovation is the Kolmogorov–Arnold Network (KAN), which has demonstrated benefits over traditional MLPs, including faster neural scaling and better interpretability. The application of KANs to physics-informed learning has led to the development of Physics-Informed KANs (PIKANs), enabling the use of KANs to solve PDEs. However, despite their advantages, KANs often suffer from slower training speeds, particularly in higher-dimensional problems where the number of collocation points grows exponentially with the dimensionality of the system. To address this challenge, we introduce Separable Physics-Informed Kolmogorov–Arnold Networks (SPIKANs). This novel architecture applies the principle of separation of variables to PIKANs, decomposing the problem such that each dimension is handled by an individual KAN. This approach drastically reduces the computational complexity of training without sacrificing accuracy, facilitating their application to higher-dimensional PDEs. Through a series of benchmark problems, we demonstrate the effectiveness of SPIKANs, showcasing their superior scalability and performance compared to PIKANs and highlighting their potential for solving complex, high-dimensional PDEs in scientific computing.

Kolmogorov-Arnold networks↗

Scaling and performance portability of the particle-in-cell scheme for plasma physics applications through mini-apps targeting exascale architectures

We perform a scaling and performance portability study of the particle-in-cell scheme for plasma physics applications through a set of mini-apps we name "Alpine", which can make use of exascale computing capabilities. The mini-apps are based on Independent Parallel Particle Layer, a framework that is designed around performance portable and dimension independent particles and fields. We benchmark the simulations with varying parameters such as grid resolutions (5123 to 20483) and number of simulation particles (109 to 1011) with the following mini-apps: weak and strong Landau damping, bump-on-tail and two-stream instabilities, and the dynamics of an electron bunch in a charge-neutral Penning trap. We show strong and weak scaling and analyze the performance of different components on several pre-exascale architectures such as Piz-Daint, Cori, Summit and Perlmutter. While the scaling and portability study helps identify the performance critical components of the particle-in-cell scheme in the current state-of-the-art computing architectures, the mini-apps by themselves can be used to develop new algorithms and optimize their high performance implementations targeting exascale architectures.

Muralikrishnan, Sriramkrishnan↗

Real-time High-resolution X-Ray Computed Tomography

Computed Tomography (CT) serves as a key imaging technology that relies on computationally intensive filtering and back-projection algorithms for 3D image reconstruction. While conventional high-resolution image reconstruction (> 2K3) solutions provide quick results, they typically treat reconstruction as an offline workload to be performed remotely on large-scale HPC systems. The growing demand for post-construction AI-driven analytics and the need for real-time adjustments call for high-resolution reconstruction solutions that are feasible on local computing resources, i.e. a multi-GPU server at most. In this paper, we propose a novel approach that utilizes Tensor Cores to optimize image reconstruction without sacrificing precision. We also introduce a framework designed to enable real-time execution of end-to-end distributed image reconstruction in a multi-GPU environment. Evaluations conducted on a single Nvidia A100 and H100 GPU show performance improvements of 1.91 × and 2.15 × compared to highly optimized production libraries. Furthermore, our framework, when deployed on 8-card Nvidia A100 GPU system, demonstrates the ability to reconstruct real-world datasets into 20483 volumes (32 GB) in slightly more than one minute and 40963 volumes (256 GB) in 7 minutes.

Wu, Du↗

Machine Learning‐Guided Discovery of High‐Entropy Perovskite Oxide Electrocatalysts via Oxygen Vacancy Engineering

Abstract High‐entropy perovskite oxides (HEPOs) have recently emerged as multifunctional catalysts. However, the HEPOs’ structural and compositional complexity hinders the easy and accurate extrapolation of activity indicators, which are essential for establishing structure‐property correlations. Here, OxiGraphX, is introduced as a novel graph neural network (GNN) model designed to capture the complex relationships among structure, composition, and atomic chemical environments for accurate prediction of oxygen vacancy formation energies (OVFEs) in HEPOs. By integrating machine learning (ML), density functional theory (DFT), and experimental validation, this work demonstrates an efficient framework for rapidly and accurately screening HEPO electrocatalysts for oxygen evolution reaction (OER). The OxiGraphX predicts OVFEs with a precision exceeding existing data, enabling the identification of compositions of higher oxygen vacancy content (OVC) and, thus, higher catalytic activity. Furthermore, the model explores latent spaces that translate effectively into experimental domains, bridging computational predictions with real‐world applications. This approach accelerates the discovery of high‐performance HEPO catalysts while providing deeper insights into their catalytic mechanisms.

Chemistry↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Topology Optimization of Curved Electrodes with Suggested Simplified Corrugated Designs

The demand for space-efficient energy storage, from grid to device scale, has soared with the rapid electrification of society. To increase energy density while maintaining high rate performance, shaping electrodes has been shown to be an effective strategy. Computational optimization can suggest designs that improve energy storage performance; however, generated designs are often complex, leading to manufacturing challenges. Additionally, most optimization is typically conducted in standard rectangular domains, whereas some applications, such as in the aerospace industry or in consumer electronics, may benefit from optimized electrodes that fit in a nonstandard form factor. Here, in this work, we use topology optimization to design capacitive full-cell electrodes in a curved domain. We propose a simplified corrugated electrode design based on the optimization results and show that the energy stored in the simplified design is only reduced by about 5–27% when compared to that in the fully optimized design, with more deviation occurring when ionic transport in the porous electrode is particularly impeded. Overall, we find that these simplified designs are promising, particularly given the more conventional manufacturing techniques required, and have a 42–369% increase in energy stored compared to monolithic electrodes.

Energy - Storage↗

Quantum Transfer Learning to Boost Dementia Detection

Dementia is a devastating condition with profound implications for individuals, families, and healthcare systems. Early and accurate detection of dementia is critical for timely intervention and improved patient outcomes. While classical machine learning and deep learning approaches have been explored extensively for dementia prediction, these solutions often struggle with high-dimensional biomedical data and large-scale datasets, quickly reaching computational and performance limitations. To address this challenge, quantum machine learning (QML) has emerged as a promising paradigm, offering faster training and advanced pattern recognition capabilities. This work aims to demonstrate the potential of quantum transfer learning (QTL) to enhance the performance of a weak classical deep learning model applied to a binary classification task for dementia detection. Besides, we show the effect of noise on the QTL-based approach, investigating the reliability and robustness of this method. Using the OASIS 2 dataset, we show how quantum techniques can transform a suboptimal classical model into a more effective solution for biomedical image classification, highlighting their potential impact on advancing healthcare technology.

Bhowmik, Sounak [University of Tennessee, Knoxvill↗

The Radical Atom: Mechanosynthetic 3D Printing of an Atomically Precise SPM Tip

This research effort sought to overcome current limitations in scanning probe-based atomic manipulation to enable atomically precise manufacturing (APM). Previous theoretical and experimental works on atom by atom and molecule by molecule fabrication of precise structures are limited to essentially to two-dimensions. APM will enable a paradigm shift in 21st century manufacturing practices in which every single atom in a electronic chip, device or machine can be placed in an exact and predefined position in three-dimensions. By providing a general method for generating reproducible SPM tip structure, this project will drive forward the entire field of atomically precise scanning probe microscopy, opening the door to positional control of nearly arbitrary covalent chemistry. Such control could, for example, be used in applications such as novel 2.5 or 3D microchip fabrication. The creation of a unique manufacturing method through APM has the potential to impact technologies at the theoretical limits of performance, weight, and utility including: solid-state quantum and spintronic computing systems, high efficiency optical antenna, solar power systems, defect engineered materials and extremely efficient catalysts. Although this experiment focused on pick-and-place non-scalable APM, the better understanding of the chemistry is crucial to the eventual goal of scalable APM. To place individual atoms into a specified location is a seminal aspiration of researchers and engineers in the many fields and may have early premium applications in medical devices and microelectronics.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Self-driving thin film laboratory: autonomous epitaxial atomic-layer synthesis via real-time computer vision analysis of electron diffraction

Emerging materials science platforms with the ability to make autonomous decisions on the fly are fundamentally changing the outlook and protocols for materials optimization and discovery. Because AI-driven self-navigating schemes can effectively reduce the total number of iterations needed to arrive at the "answer" (i.e. the best stochiometric composition for a desired physical property, optimum materials processing parameters, etc.) by significant margins, they have the potential to revolutionize materials and chemical manufacturing processes at large in research laboratory settings as well as in industrial plants. Here, we demonstrate a successful implementation of real-time closed-loop autonomous navigation of a multi-dimensional materials synthesis parameter space for fabricating phase-pure epitaxial films of a metastable phase of a functional oxide in a combinatorial pulsed laser deposition chamber. Sequential epitaxial growth iterations in search of the optimized recipe to stabilize the desired crystal phase were performed using frame-by-frame quantitative computer vision analysis of reflection high-energy electron diffraction (RHEED) images of the unit-cell level film being deposited. The autonomous scheme regularly resulted in > 30-fold reduction in the number of required experiments compared to a comprehensive mapping of the parameter space. The real-time workflow developed here can be readily extended to a variety of thin film synthesis platforms opening the door for self-driving atomic-level materials design as well as autonomous optimization of semiconductor manufacturing.

36 MATERIALS SCIENCE↗

Fast and Accurate Intersections on a Sphere

We introduce a fast, high-precision algorithm for calculating intersections between great circle arcs and lines of constant latitude on the unit sphere. We first propose a simplified intersection point formula with improved speed and numerical robustness over the ones traditionally implemented in geoscience software. We then show how algorithms based on the concept of error-free transformations (EFT) can be applied to evaluate this formula within a relative error bound that is on the order of machine precision. Here, we demonstrate that, with a vectorized and parallelized implementation, this enhanced accuracy is achieved with no compute time overhead compared to a direct calculation in hardware floating point, making our algorithm suitable for performance-sensitive applications like regridding of high-resolution climate data. In contrast, evaluating our formula using high-precision data types like quadruple precision and arbitrary precision, or using the robust intersection computation routines from the Computational Geometry Algorithms Library, leads to significant computational overhead, especially since these alternatives inhibit vectorization. More generally, our work demonstrates how EFT techniques can be combined and extended to implement nontrivial geometric calculations with high accuracy and speed.

Environmental sciences↗

Unrolled Video Super-Resolution Network with Autoregressive Prior for the Case of Known Motion

Real-time detection and classification of distant objects is necessary for many national security applications. However, when objects are far from the sensor, they occupy only a small number of pixels in the captured video, limiting the amount of visual detail available for recognition. State-of-the-art classification methods typically rely on high-resolution (HR) video streams to capture characteristic object features, but obtaining such detail is challenging for distant objects that occupy only a few pixels. This motivates the development of video super-resolution (VSR) methods that enhance object classification by recovering fine details from low-pixel representations. Current VSR methods rely either on model-based optimization, which is interpretable but computationally expensive, or on learning-based approaches, which are efficient and high-performing but often lack flexibility and interpretability. In this report, we propose an end-to-end trainable unrolled VSR network, UVSRNet, which super-resolves each frame in a video by exploiting sub-pixel motion between neighboring low-resolution (LR) frames as well as incorporating high-frequency detail from previously super-resolved frames. In particular, by unrolling a plug-and-play (PnP) half-quadratic splitting (HQS) algorithm, we leverage a model-based data-fitting module alongside a learning-based autoregressive prior module. This combination yields a method that maintains the flexibility and interpretability of model-based methods while achieving the performance advantages of learning-based methods.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

A Synthesis Methodology for Intelligent Memory Interfaces in Accelerator Systems

Domain-specific systems improve the performance of a specific set of applications compared to general-purpose processing systems by deploying custom hardware accelerators. These hardware accelerators are generated using high-level synthesis (HLS) tools. The HLS tools enable a comprehensive design space exploration to optimize the compute performance of the generated accelerators. However, they often ignore the challenges of implementing the accelerators in a system-on-chip, particularly how the accelerators access memory. Our work introduces a buffering system design that improves accelerators' memory accesses by intelligently employing burst transactions to prefetch useful data from external memory to on-chip local buffers. Our design is dynamic, parametric, and transparent to the accelerators generated by HLS tools. We derive the buffering system parameters using appropriate compiler-based analysis passes and memory channel latency constraints. The proposed buffering system design results in, on average, 8.8x performance improvements while lowering memory channel utilization on average by 53.2% for a set of PolyBench kernels.

Limaye, Ankur M. (ORCID:0000000194062584)↗

Brittle failure analysis and modeling of high-burnup PWR fuel cladding alloys

The aim of this research is the development of methods for predicting mechanical behavior and identification of limiting conditions to prevent brittle failure of high-burnup (HBU) pressure water reactor (PWR) fuel cladding alloys. A finite element (FE) model of the ring compression test (RCT) was created to analyze the failure behavior of zirconium-based alloys with radial hydrides during the RCT. An elastic-plastic material model describes the zirconium alloy. The stress-strain curve needed for the elastic-plastic material model was derived by inverse finite element analyses. Cohesive zone modeling is used to reproduce sudden load drops during RCT loading. Based on the failure mechanism in non-irradiated ZIRLO (R) claddings, a micro-mechanical model was developed that distinguishes between brittle failure along hydrides and ductile failure of the zirconium matrix. Two different cohesive laws representing these types of failure are present in the same cohesive interface. The key differences between these constitutive laws are the cohesive strength, the stress at which damage initiates, and the cohesive energy, which is the damage energy dissipated by the cohesive zone. Statistically generated matrix-hydride distributions were mapped onto the cohesive elements and simulations with focus on the first load drop were performed. Computational results are in good agreement with the RCT results conducted on high-burnup M5 (R) samples. It could be shown that crack initiation and propagation strongly depend on the specific configuration of hydrides and matrix material in the fracture area.

Simbruner, Kai↗

Leveraging High-throughput Computation and Machine Learning to Discover and Understand Low-Temperature Fast Oxygen Conductors (Final Technical Report)

The major goals of this work are twofold: (1) to enable transformative basic understanding of structure-property-performance relationships governing oxygen transport in oxygen-active materials and (2) facilitate the discovery and rational design of new oxygen-active materials which transport oxygen efficiently at low temperature. Transformative understanding and materials design will be accomplished by synergistically combining materials data mining, machine learning, high-throughput computation and targeted experiments.

36 MATERIALS SCIENCE↗

The Remarkable I 2 O 3 Molecule: A New View from Theory

Atmospheric iodine chemistry has garnered increasing attention as a result of increased iodine emissions. A key subset of this chemistry involves iodine oxides (I 2 O 2–5 ), which serve as precursors to particle formation. Among these, I 2 O 3 is the simplest iodine oxide involved in particle formation, but it has remained undetected in the atmosphere. Previous theoretical studies have characterized this peculiar molecule, primarily using energies to refine geometries obtained at low levels of theory. Due to the reemerging interest in I 2 O 3 , this study presents geometries optimized at the CCSD(T)/aug-cc-pwCVTZ-PP level of theory─marking the first instance, to the best of our knowledge, where this system has been studied exclusively with CCSD(T). Harmonic vibrational frequencies were computed at the same level of theory. Final energetics were obtained using the very high level CCSDT(Q) method with basis sets up to quintuple-zeta cardinality (aug-cc-pwCV5Z-PP) and extrapolated to the CBS limit to yield CCSDT(Q)/CBS//CCSD(T)/aug-cc-pwCVTZ-PP energies. These energies include harmonic zero-point vibrational energy corrections and scalar relativistic energy corrections. Additionally, this study discovers new isomers along the I 2 O 3 potential energy surface, a novel contribution to the field. The performance of different computational methods and DFT functionals commonly used in atmospheric chemistry is also assessed relative to high-level theoretical methods.

basis sets↗

Design-to-Deployment Continuum Platform for Microscopes and Computing Ecosystems

Science ecosystems with networked computing systems and physical instruments are increasingly being deployed with a goal to achieve the productivity promised by AI-supported remote automation. In support of these efforts, the virtual infrastructure twins (VITs) have been successfully utilized to develop the orchestration codes for these ecosystems without requiring physical access to expensive instruments, such as electron microscopes. Currently, the utility of such a VIT is severely limited by the computing capacity and capability of the computing system used as its host. Furthermore, codes developed on the VIT typically need to be transferred and refactored for production use, particularly, on high-performance systems with accelerators. In response, we develop a design-to-deployment continuum platform wherein a VIT runs natively on the ecosystem's own computing system, and thereby facilitates the continual in-situ testing and transition of codes for production use. Here, we describe the development and testing of software for remote microscope steering and GPU-based image reconstruction using this platform on a multi-GPU computing system networked to Nion microscopes. We demonstrate a continual transition of steering and reconstruction codes developed under VIT platform to production ecosystem deployment.

Al-Najjar, Anees [Oak Ridge National Laboratory (O↗