Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49

LHC EFT WG note: SMEFT predictions, event reweighting, and simulation

This note provides a comprehensive overview of tools for predicting observables in the Standard Model effective field theory (SMEFT) at both tree level and one loop using event generators. We evaluate three primary methodologies–event reweighting, separate simulation of squared matrix elements, and full SMEFT process simulation–focusing on their statistical performance, computational efficiency, and potential biases. Each approach is assessed in terms of its accuracy, highlighting trade-offs between precision and resource demands. Practical insights into their applicability for high-energy physics analyses are offered, with particular attention to processes where SMEFT effects are significant. Additionally, we discuss the role of helicity in reweighting strategies and its impact on the quality of predictions. By comparing the methods across various LHC processes, this note provides guidance for selecting the most effective strategy for various SMEFT studies, ensuring robust predictions while optimizing computational resources.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Preliminary Plan to Inform Testing of a Heat Exchanger Test Article

This report presents a preliminary plan to guide the qualification testing of advanced heat exchanger (HX) components for nuclear-to-industrial heat transfer applications. The objective is to establish a defensible, physics-based methodology that integrates computational modeling, targeted experimentation, and in-service inspection considerations to demonstrate component performance and reliability under representative reactor conditions. The analysis identifies Sodium-cooled Fast Reactor (SFR) and High-Temperature Gas-cooled Reactor (HTGR) systems as reference configurations in terms of temperature, pressure, and chemical environment. Within these operating envelopes, dominant degradation mechanisms— including creep–fatigue interaction, flow-induced vibration, corrosion, and diffusion-bond deterioration—were evaluated to define test requirements. A comprehensive computationalexperimental framework is proposed to support life prediction and qualification activities. The framework couples high-fidelity structural-mechanics, thermal-hydraulic, and fluid-structure interaction models with accelerated degradation testing to produce a traceable linkage between microstructural evolution, mechanical performance, and remaining useful life (RUL). The approach adheres to established Verification, Validation, and Uncertainty Quantification (VVUQ) standards (ASME V&V 10/20; NUREG-2152) and incorporates a digital-twin architecture for continuous model refinement through data assimilation. The plan further outlines testing methodologies, including pre-test analyses, test-loop design parameters, and sensor placement strategies that maximize information yield while maintaining mechanistic fidelity. Complementary sections describe in-service inspection (ISI), on-line monitoring (OLM), and structural-health-monitoring (SHM) techniques applicable to compact HX geometries typical of advanced reactors. Collectively, these activities establish the technical foundation for demonstrating 40-60-year equivalent service life of advanced heat exchangers in support of the U.S. Department of Energy’s Advanced Reactor and Integrated Energy Systems programs. The forthcoming phase will execute the defined pre-test analyses, initiate hardware fabrication, and implement the integrated testing campaign to validate the proposed qualification methodology.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

High tensile alloy of copper to mitigate current collector deformation in silicon electrodes for lithium-ion batteries

The volumetric changes of silicon electrodes, along with the strong adhesive properties of certain binders, can lead to plastic deformation of the current collector and create damage in the electrode coating. Here, in this study, we report a detailed study of silicon coatings on a high-tensile alloy (HTA) foil of copper with strength over twice that of conventional copper foils. The HTA current collectors with high mechanical strength can mitigate plastic deformation upon continuous cycling. At moderate areal capacities (2.5–3 mAh cm −2 ), conventional copper foils show significant wrinkling after only a few electrochemical cycles, whereas the HTA foils remain intact. We demonstrate viability of the HTA foils in large format xx6395 pouch cells, in which the HTA current collectors remain intact even at an areal capacity of 4.5 mAh cm −2 ; in contrast, wrinkles form in conventional copper current collectors increasing the likelihood of lithium plating. Computational studies show that stresses generated during cycling of silicon electrodes are very high in the current collector and at the current collector-coating interface, explaining the wrinkling of conventional Cu foils. Our studies highlight importance of current collector to solve the electrochemical and chemo-mechanical performance challenges associated with high-loading silicon electrodes.

Chemo-mechanical degradation↗

High-performance data management for whole slide image analysis in digital pathology

When dealing with giga-pixel digital pathology in whole-slide imaging, a notable proportion of data records holds relevance during each analysis operation. For instance, when deploying an image analysis algorithm on whole-slide images (WSI), the computational bottleneck often lies in the input-output (I/O) system. This is particularly notable as patch-level processing introduces a considerable I/O load onto the computer system. However, this data management process could be further paralleled, given the typical independence of patch-level image processes across different patches. This paper details our endeavors in tackling this data access challenge by implementing the Adaptable IO System version 2 (ADIOS2). Our focus has been constructing and releasing a digital pathology-centric pipeline using ADIOS2, which facilitates streamlined data management across WSIs. Additionally, we’ve developed strategies aimed at curtailing data retrieval times. The performance evaluation encompasses two key scenarios: (1) a pure CPU-based image analysis scenario (“CPU scenario”), and (2) a GPU-based deep learning framework scenario (“GPU scenario”). Our findings reveal noteworthy outcomes. Under the CPU scenario, ADIOS2 showcases an impressive two-fold speed-up compared to the brute-force approach. In the GPU scenario, its performance stands on par with the cutting-edge GPU I/O acceleration framework, NVIDIA Magnum IO GPU Direct Storage (GDS). From what we know, this appears to be among the initial instances, if any, of utilizing ADIOS2 within the field of digital pathology. The source code has been made publicly available at https://github.com/hrlblab/adios.

Wang, Xiao↗

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

ObstacleSense: Low-Power Neuromorphic Vision for Corridor Obstacle Awareness in Low-Level ADAS

The automotive industry’s pursuit of Level 5 autonomy is constrained by substantial perception-compute power requirements, often reaching 1, 000 + watts in full autonomy stacks. Reducing this energy burden requires rethinking perception not only at the high-end autonomy level, but also at the foundational Advanced Driver Assistance Systems (ADAS) level where low-power, safety-critical sensing can have broad impact. Neuromorphic vision provides a promising starting point: HD Dynamic Vision Sensors (DVS) can operate below 100 mW at the sensor level by reporting only asynchronous brightness changes. However, low-power sensing alone is insufficient if downstream perception reintroduces dense, energy-intensive computation. In particular, many event-driven object-detection pipelines still rely on CNN backbones, while purely spiking alternatives often trade away accuracy or ignore deployment constraints. We introduce ObstacleSense, a highly compact, CNN-free hybrid ANN–SNN framework for Level 0–1 forward-corridor obstacle awareness. Instead of performing full-scene object detection with a convolutional feature backbone, ObstacleSense targets the safety-critical question of whether the ego corridor is occupied and how far the nearest obstacle is. The architecture combines polarity-conditioned event encoding, lightweight temporal spiking dynamics, axial spatial mixing, and coarse-to-fine range estimation within a regular fixed-grid compute pattern. This design avoids the dense CNN backbone commonly used in event-based detection while maintaining a small state footprint suitable for eventual small-FPGA deployment. Before hardware mapping, we evaluate the software implementation using a model-side power proxy derived from MACs, weight and activation traffic, and spiking state updates under shared FP16 assumptions. On simulated CARLA event corpora, the deployment-oriented model achieves 0.9464 objectness F1, 0.9978 grid-level mAP, and 0.8987 m distance Mean Absolute Error at an estimated 1.92 mW proxy cost, while maintaining performance on unseen generalization test sequences.

Johnson-Scott, Zac [ORNL]↗

PHASE: Personalized Head-based Automatic Simulation for Electromagnetic properties in 7T MRI

Accurate and individualized human head models are becoming increasingly important for electromagnetic (EM) simulations. These simulations depend on precise anatomical representations to realistically model electric and magnetic field distributions, particularly when evaluating Specific Absorption Rate (SAR) within safety guidelines. State of the art simulations use the Virtual Population due to limited public resources and the impracticality of manually annotating patient data at scale. Here, this paper introduces Personalized Head-based Automatic Simulation for EM properties (PHASE), an automated open-source toolbox that generates high-resolution, patient-specific head models for EM simulations using paired T1-weighted (T1w) magnetic resonance imaging (MRI) and computed tomography (CT) scans with 14 tissue labels. To evaluate the performance of PHASE models, we conduct semi-automated segmentation and EM simulations on 15 real human patients, serving as the gold standard reference. The PHASE model achieved comparable global SAR and localized SAR averaged over 10 grams of tissue (SAR-10g), demonstrating its potential as a promising tool for generating large-scale human model datasets in the future. The code and models of PHASE toolbox have been made publicly available: https://github.com/hrlblab/PHASE.

Deep learning↗

Hierarchical Testing of a Hybrid Machine Learning‐Physics Global Atmosphere Model

Machine learning (ML)-based models have demonstrated high skill and computational efficiency, often outperforming conventional physics-based models in weather and subseasonal predictions. While prior studies have assessed their fidelity in capturing synoptic-scale atmospheric dynamics, their performance across timescales and under out-of-distribution forcing, such as +3K or +4K uniform-warming forcings, and the sources of biases remain elusive, to establish the model's reliability for Earth science. Here, we design three sets of experiments targeting synoptic-scale phenomena, interannual variability, and out-of-distribution uniform-warming forcings. We evaluate the Neural General Circulation Model (NeuralGCM), a hybrid model integrating a dynamical core with ML-based component, against observations and physics-based Earth system models (ESMs). At the synoptic scale, NeuralGCM captures the evolution and propagation of extratropical cyclones with performance comparable to ESMs. At the interannual scale, when forced by El Niño-Southern Oscillation sea surface temperature (SST) anomalies, NeuralGCM successfully reproduces associated teleconnection patterns but exhibits deficiencies in capturing nonlinear response. Under out-of-distribution uniform-warming forcings, NeuralGCM simulates similar responses in global-average temperature and precipitation and reproduces large-scale tropospheric circulation features similar to those in ESMs. Notable weaknesses include overestimating the tracks and spatial extent of extratropical cyclones, biases in the teleconnected wave train triggered by tropical SST anomalies, and differences in upper-level warming and stratospheric circulation responses to SST warming compared to physics-based ESMs. The causes of these weaknesses were explored. Despite the noted weaknesses, NeuralGCM reproduces responses across experiments reasonably and performs comparably to ESMs. By integrating a dynamical core with ML, NeuralGCM shows potential for developing ML-based ESMs.

global warming↗

Deep learning‑based metal artefact reduction in X-ray computed tomography of TRISO fuel compacts

TRISO (TRi-structural ISOtropic) – compact-type micro-particle fuels – are next generation nuclear fuel compacts designed with safety in mind. Structural integrity and characterisation before and after irradiation are important to determine the performance of the fuel compacts under real reactor conditions. X-ray computed tomography can be an important tool for non-destructive evaluation of these fuel compacts. The fuel particles are highly attenuating for X-rays which creates metal artefact, rendering the images unusable. Our artefact correction can mitigate these artefacts significantly. The proposed method works by first segmenting highly-attenuating structures (metal) and forward projecting to localise the source of artefacts in the projection domain, then, using a traditional or U-net-based deep learning architecture, contextually interpolating those regions to remove the artefacts. Finally the reconstruction from the modified projection is fused with the segmented metal. The proposed method shows a significant improvement in the image quality while significantly reducing the reconstruction time compared to the standard technique.

Rahman, Obaid [ORNL] (ORCID:0000000277810840)↗

JACC: Leveraging HPC Meta-Programming and Performance Portability with the Just-in-Time and LLVM-based Julia Language

We present JACC (Julia for Accelerators), the first high-level, and performance-portable model for the just-in-time and LLVM-based Julia language. JACC provides a unified and lightweight front end across different back ends available in Julia, enabling the same Julia code to run efficiently on many HPC CPU and GPU targets. We evaluated the performance of JACC for common HPC kernels as well as for the most computationally demanding kernels used in applications, HPCCG, a supercomputing benchmark test for sparse domains, and HARVEY, a blood flow simulator to assist in the diagnosis and treatment of patients suffering from vascular diseases. We carried out the performance analysis on the most advanced US DOE supercomputers: Aurora, Frontier, and Perlmutter. Overall, we show that JACC has a negligible overhead versus vendor-specific solutions, reporting GPU speedups with no extra cost to programmability.

Valero-Lara, Pedro↗

Non-Electricity Based Renewable Fuels: Theory and Computation for Solar Thermochemical Hydrogen

Dominated by photovoltaics and wind, current renewable energy sources generate mostly electricity, but 80% of the global final energy consumption occurs in form of fuels. Therefore, direct solar fuel generation would be a major breakthrough for the energy transition. Solar thermochemical hydrogen (STCH) is one of the very few potential routes towards scalable renewable fuels, but currently suffers from lack of an oxide working material that could optimally perform energy conversion within the thermodynamic boundary conditions. Theory and computation can contribute in two distinct ways, through materials search and discovery, but also by providing detailed mechanistic models for specific systems so to advance our understanding of possible design strategies. To enable high-throughput materials screening, we developed a defect graph neural network (dGNN) machine learning approach,[1] which accelerates the prediction of defect formation energies by replacing the tedious density functional theory (DFT) supercell calculations for all possible defect sites. This approach enables high-throughput database screening of oxides, which was integrated with thermodynamic modeling to extract the reduction entropies as additional selection criterion for STCH. Once potential candidate materials are identified, detailed models can guide materials design by predicting performance characteristics. One challenge is to quantitatively predict thermochemical equilibria at high concentrations when the redox active defects start to interact with each other, thereby impeding the formation of additional defects. Introducing a model for the free energy of defect interaction, parametrized on the basis of DFT data, we simulated the complete STCH redox cycle for (Sr,Ce)MnO3 alloys, achieving near-quantitative agreement with experimental data.[2] The analysis of these simulations reveals how defect interactions diminish the reduction entropy and H2 yield, suggesting to include these interactions in design considerations. Finally, we revisit the popular van't Hoff method for analyzing reduction enthalpies and entropies. This method is not ideal, as it involves a temperature-dependent convolution of gas-phase and solid-state entropies, causing uncertainties in the same order of magnitude as the physical quantities of interest. To avoid this problem, we suggest a simple alternative approach which can be applied to experimental and simulated data alike.

first-principles calculations↗

Cost-efficient finite-volume high-order schemes for compressible magnetohydrodynamics

We present an efficient dimension-by-dimension finite-volume method which solves the adiabatic magnetohydrodynamics equations at high discretization order, using the constrained-transport approach on Cartesian grids. Results are presented up to tenth order of accuracy. The algorithmic architecture of this method is very close to that of commonly employed second-order schemes: it requires only one reconstructed value per face for each computational cell, independently of the scheme's order. This property is highly beneficial for the numerical efficiency. It results from reusing the required values already available in neighboring grid cells, in contrast to standard algorithms that require a number of reconstructions and evaluations which increases with the scheme's order of accuracy. At a given resolution, these high-order schemes present significantly less numerical dissipation than commonly employed lower-order approaches. Thus, results of comparable accuracy are achievable at a substantially coarser resolution, yielding overall performance gains. We also present a way to include physical dissipative terms: viscosity, magnetic diffusivity and cooling functions, respecting the finite-volume and constrained-transport frameworks. Benefits of this method are shown through applications in turbulent flows.

97 MATHEMATICS AND COMPUTING↗

Monocrystalline CdSeTe/MgCdTe Double‐Heterostructure Solar Cells

This article reports monocrystalline CdSeTe/MgCdTe double‐heterostructure (DH) solar cells with varying Se compositions in the absorber layers that are grown on InSb substrates by using molecular beam epitaxy. The Se composition in the samples studied is determined to be 4%–11% through high‐resolution X‐ray diffraction (XRD) and photoluminescence measurements. Increased Se incorporation induces higher defect density attributed to increased lattice mismatch, and mixed‐phase formation due to the small difference in formation energies between the wurtzite and zinc blende phases of CdSe, which causes stacking faults and grain boundaries. Devices are fabricated by directly depositing an n‐type indium tin oxide (ITO) layer on the CdSeTe/MgCdTe DHs followed by Ag metal contacts. Reduced bandgap is observed in solar cells with increased Se composition. The devices with an absorber containing 4% Se exhibit an average open‐circuit voltage ( V OC ) of 0.925 V, a short‐circuit current density ( J SC ) of 23.4 mA/cm 2 , a fill factor ( FF ) of 0.650 and an efficiency of 14.1% without anti‐reflection coating. Absorbers with higher Se compositions result in poor device performance, mainly attributed to high defect density in the materials.

Ju, Zheng [Center for Photonics Innovation Arizona↗

Optimal 3D chemical imaging with multimodal electron tomography

Accurate mapping of nanoscale chemistry in three dimensions (3D) has been a longstanding challenge. Modern electron microscopy provides chemical images by electron energy loss spectroscopy (EELS) and energy dispersive x-ray spectrometry (EDX) but requires high fluences that damage specimens. In 3D, the requirements are worse; electron tomography demands many high-fluence chemical maps for reconstruction, creating a tradeoff between resolution, accuracy, and sample survival. Fused multimodal electron tomography (MM-ET) alleviates this requirement by leveraging lower-fluence high-angle annular dark-field (HAADF) images alongside a few chemical maps to dramatically improve chemical resolution. Here, experimental and computational parameter space is systematically explored to determine when MM-ET performs best. Ideal imaging conditions balance sample survival with resolution and chemical specificity; we recommend a tilt range of at least ± 70°, acquiring 40 equally spaced HAADF projections (signal-to-noise > 10), and 7 EELS/EDX maps of each chemistry (signal-to-noise > 4).

36 MATERIALS SCIENCE↗

Efficient lattice QCD computation of radiative-leptonic-decay form factors at multiple positive and negative photon virtualities

In previous work [D. Giusti, Methods for high-precision determinations of radiative-leptonic decay form factors using lattice QCD, Phys. Rev. D 107, 074507 (2023)], we showed that form factors for radiative leptonic decays of pseudoscalar mesons can be determined efficiently and with high precision from lattice QCD using the “three-dimensional (3D) method,” in which three-point functions are computed for all values of the current insertion time and the time integral is performed at the data-analysis stage. Here, we demonstrate another benefit of the 3D method: the form factors can be extracted for any number of nonzero photon virtualities from the same three-point functions at no extra cost. We present results for the $D_s → ℓνγ*$ vector form factor as a function of photon energy and photon virtuality, for both positive and negative virtuality, for a single ensemble with 340 MeV pion mass and 0.11 fm lattice spacing. In our analysis, we separately consider the two different time orderings and the different quark flavors in the electromagnetic current. We discuss in detail the behavior of the unwanted exponentials contributing to the three-point functions, as well as the choice of fit models and fit ranges used to remove them for various values of the virtuality. While positive photon virtuality is relevant for decays to multiple charged leptons, negative photon virtuality suppresses soft contributions and is of interest in QCD-factorization studies of the form factors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Asymptotic-state prediction for fast flavor transformation in neutron star mergers

Neutrino flavor instabilities appear to be omnipresent in dense astrophysical environments, thus presenting a challenge to large-scale simulations of core-collapse supernovae and neutron star mergers (NSMs). Subgrid models offer a path forward, but require an accurate determination of the local outcome of such conversion phenomena. Focusing on “fast” instabilities, related to the existence of a crossing between neutrino and antineutrino angular distributions, we consider a range of analytical mixing schemes, including a new, fully three-dimensional one, and also introduce a new machine learning (ML) model. We compare the accuracy of these models with the results of several thousands of local dynamical calculations of neutrino evolution from the conditions extracted from classical NSM simulations. Our ML model shows good overall performance, but struggles to generalize to conditions from a NSM simulation not used for training. The multidimensional analytic model performs and generalizes even better, while other analytic models (which assume axisymmetric neutrino distributions) do not have reliably high performances, as they notably fail as expected to account for effects resulting from strong anisotropies. As a result, the ML and analytic subgrid models extensively tested here are both promising, with different computational requirements and sources of systematic errors.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward a 2D Local Implementation of Quantum Low-Density Parity-Check Codes

Geometric locality is an important theoretical and practical factor for quantum low-density parity-check (qLDPC) codes that affects code performance and ease of physical realization. For device architectures restricted to two-dimensional (2D) local gates, naively implementing the high-rate codes suitable for low-overhead fault-tolerant quantum computing incurs prohibitive overhead. In this work, we present an error-correction protocol built on a bilayer architecture that aims to reduce operational overheads when restricted to 2D local gates by measuring some generators less frequently than others. We investigate the family of bivariate-bicycle qLDPC codes and show that they are well suited for a parallel syndrome-measurement scheme using fast routing with local operations and classical communication (LOCC). Through circuit-level simulations, we find that in some parameter regimes, bivariate-bicycle codes implemented with this protocol have logical error rates comparable to the surface code while using fewer physical qubits. Published by the American Physical Society 2025

Berthusen, Noah (ORCID:0000000275862786)↗