Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51

Image processing tools for petabyte-scale light sheet microscopy data

Light sheet microscopy is a powerful technique for high-speed three-dimensional imaging of subcellular dynamics and large biological specimens. However, it often generates datasets ranging from hundreds of gigabytes to petabytes in size for a single experiment. Conventional computational tools process such images far slower than the time to acquire them and often fail outright due to memory limitations. To address these challenges, we present PetaKit5D, a scalable software solution for efficient petabyte-scale light sheet image processing. This software incorporates a suite of commonly used processing tools that are optimized for memory and performance. Notable advancements include rapid image readers and writers, fast and memory-efficient geometric transformations, high-performance Richardson–Lucy deconvolution and scalable Zarr-based stitching. These features outperform state-of-the-art methods by over one order of magnitude, enabling the processing of petabyte-scale image data at the full teravoxel rates of modern imaging cameras. The software opens new avenues for biological discoveries through large-scale imaging experiments.

97 MATHEMATICS AND COMPUTING↗

ObstacleSense: Low-Power Neuromorphic Vision for Corridor Obstacle Awareness in Low-Level ADAS

The automotive industry’s pursuit of Level 5 autonomy is constrained by substantial perception-compute power requirements, often reaching 1, 000 + watts in full autonomy stacks. Reducing this energy burden requires rethinking perception not only at the high-end autonomy level, but also at the foundational Advanced Driver Assistance Systems (ADAS) level where low-power, safety-critical sensing can have broad impact. Neuromorphic vision provides a promising starting point: HD Dynamic Vision Sensors (DVS) can operate below 100 mW at the sensor level by reporting only asynchronous brightness changes. However, low-power sensing alone is insufficient if downstream perception reintroduces dense, energy-intensive computation. In particular, many event-driven object-detection pipelines still rely on CNN backbones, while purely spiking alternatives often trade away accuracy or ignore deployment constraints. We introduce ObstacleSense, a highly compact, CNN-free hybrid ANN–SNN framework for Level 0–1 forward-corridor obstacle awareness. Instead of performing full-scene object detection with a convolutional feature backbone, ObstacleSense targets the safety-critical question of whether the ego corridor is occupied and how far the nearest obstacle is. The architecture combines polarity-conditioned event encoding, lightweight temporal spiking dynamics, axial spatial mixing, and coarse-to-fine range estimation within a regular fixed-grid compute pattern. This design avoids the dense CNN backbone commonly used in event-based detection while maintaining a small state footprint suitable for eventual small-FPGA deployment. Before hardware mapping, we evaluate the software implementation using a model-side power proxy derived from MACs, weight and activation traffic, and spiking state updates under shared FP16 assumptions. On simulated CARLA event corpora, the deployment-oriented model achieves 0.9464 objectness F1, 0.9978 grid-level mAP, and 0.8987 m distance Mean Absolute Error at an estimated 1.92 mW proxy cost, while maintaining performance on unseen generalization test sequences.

Johnson-Scott, Zac [ORNL]↗

PHASE: Personalized Head-based Automatic Simulation for Electromagnetic properties in 7T MRI

Accurate and individualized human head models are becoming increasingly important for electromagnetic (EM) simulations. These simulations depend on precise anatomical representations to realistically model electric and magnetic field distributions, particularly when evaluating Specific Absorption Rate (SAR) within safety guidelines. State of the art simulations use the Virtual Population due to limited public resources and the impracticality of manually annotating patient data at scale. Here, this paper introduces Personalized Head-based Automatic Simulation for EM properties (PHASE), an automated open-source toolbox that generates high-resolution, patient-specific head models for EM simulations using paired T1-weighted (T1w) magnetic resonance imaging (MRI) and computed tomography (CT) scans with 14 tissue labels. To evaluate the performance of PHASE models, we conduct semi-automated segmentation and EM simulations on 15 real human patients, serving as the gold standard reference. The PHASE model achieved comparable global SAR and localized SAR averaged over 10 grams of tissue (SAR-10g), demonstrating its potential as a promising tool for generating large-scale human model datasets in the future. The code and models of PHASE toolbox have been made publicly available: https://github.com/hrlblab/PHASE.

Deep learning↗

Hierarchical Testing of a Hybrid Machine Learning‐Physics Global Atmosphere Model

Machine learning (ML)-based models have demonstrated high skill and computational efficiency, often outperforming conventional physics-based models in weather and subseasonal predictions. While prior studies have assessed their fidelity in capturing synoptic-scale atmospheric dynamics, their performance across timescales and under out-of-distribution forcing, such as +3K or +4K uniform-warming forcings, and the sources of biases remain elusive, to establish the model's reliability for Earth science. Here, we design three sets of experiments targeting synoptic-scale phenomena, interannual variability, and out-of-distribution uniform-warming forcings. We evaluate the Neural General Circulation Model (NeuralGCM), a hybrid model integrating a dynamical core with ML-based component, against observations and physics-based Earth system models (ESMs). At the synoptic scale, NeuralGCM captures the evolution and propagation of extratropical cyclones with performance comparable to ESMs. At the interannual scale, when forced by El Niño-Southern Oscillation sea surface temperature (SST) anomalies, NeuralGCM successfully reproduces associated teleconnection patterns but exhibits deficiencies in capturing nonlinear response. Under out-of-distribution uniform-warming forcings, NeuralGCM simulates similar responses in global-average temperature and precipitation and reproduces large-scale tropospheric circulation features similar to those in ESMs. Notable weaknesses include overestimating the tracks and spatial extent of extratropical cyclones, biases in the teleconnected wave train triggered by tropical SST anomalies, and differences in upper-level warming and stratospheric circulation responses to SST warming compared to physics-based ESMs. The causes of these weaknesses were explored. Despite the noted weaknesses, NeuralGCM reproduces responses across experiments reasonably and performs comparably to ESMs. By integrating a dynamical core with ML, NeuralGCM shows potential for developing ML-based ESMs.

global warming↗

Deep learning‑based metal artefact reduction in X-ray computed tomography of TRISO fuel compacts

TRISO (TRi-structural ISOtropic) – compact-type micro-particle fuels – are next generation nuclear fuel compacts designed with safety in mind. Structural integrity and characterisation before and after irradiation are important to determine the performance of the fuel compacts under real reactor conditions. X-ray computed tomography can be an important tool for non-destructive evaluation of these fuel compacts. The fuel particles are highly attenuating for X-rays which creates metal artefact, rendering the images unusable. Our artefact correction can mitigate these artefacts significantly. The proposed method works by first segmenting highly-attenuating structures (metal) and forward projecting to localise the source of artefacts in the projection domain, then, using a traditional or U-net-based deep learning architecture, contextually interpolating those regions to remove the artefacts. Finally the reconstruction from the modified projection is fused with the segmented metal. The proposed method shows a significant improvement in the image quality while significantly reducing the reconstruction time compared to the standard technique.

Rahman, Obaid [ORNL] (ORCID:0000000277810840)↗

JACC: Leveraging HPC Meta-Programming and Performance Portability with the Just-in-Time and LLVM-based Julia Language

We present JACC (Julia for Accelerators), the first high-level, and performance-portable model for the just-in-time and LLVM-based Julia language. JACC provides a unified and lightweight front end across different back ends available in Julia, enabling the same Julia code to run efficiently on many HPC CPU and GPU targets. We evaluated the performance of JACC for common HPC kernels as well as for the most computationally demanding kernels used in applications, HPCCG, a supercomputing benchmark test for sparse domains, and HARVEY, a blood flow simulator to assist in the diagnosis and treatment of patients suffering from vascular diseases. We carried out the performance analysis on the most advanced US DOE supercomputers: Aurora, Frontier, and Perlmutter. Overall, we show that JACC has a negligible overhead versus vendor-specific solutions, reporting GPU speedups with no extra cost to programmability.

Valero-Lara, Pedro↗

Non-Electricity Based Renewable Fuels: Theory and Computation for Solar Thermochemical Hydrogen

Dominated by photovoltaics and wind, current renewable energy sources generate mostly electricity, but 80% of the global final energy consumption occurs in form of fuels. Therefore, direct solar fuel generation would be a major breakthrough for the energy transition. Solar thermochemical hydrogen (STCH) is one of the very few potential routes towards scalable renewable fuels, but currently suffers from lack of an oxide working material that could optimally perform energy conversion within the thermodynamic boundary conditions. Theory and computation can contribute in two distinct ways, through materials search and discovery, but also by providing detailed mechanistic models for specific systems so to advance our understanding of possible design strategies. To enable high-throughput materials screening, we developed a defect graph neural network (dGNN) machine learning approach,[1] which accelerates the prediction of defect formation energies by replacing the tedious density functional theory (DFT) supercell calculations for all possible defect sites. This approach enables high-throughput database screening of oxides, which was integrated with thermodynamic modeling to extract the reduction entropies as additional selection criterion for STCH. Once potential candidate materials are identified, detailed models can guide materials design by predicting performance characteristics. One challenge is to quantitatively predict thermochemical equilibria at high concentrations when the redox active defects start to interact with each other, thereby impeding the formation of additional defects. Introducing a model for the free energy of defect interaction, parametrized on the basis of DFT data, we simulated the complete STCH redox cycle for (Sr,Ce)MnO3 alloys, achieving near-quantitative agreement with experimental data.[2] The analysis of these simulations reveals how defect interactions diminish the reduction entropy and H2 yield, suggesting to include these interactions in design considerations. Finally, we revisit the popular van't Hoff method for analyzing reduction enthalpies and entropies. This method is not ideal, as it involves a temperature-dependent convolution of gas-phase and solid-state entropies, causing uncertainties in the same order of magnitude as the physical quantities of interest. To avoid this problem, we suggest a simple alternative approach which can be applied to experimental and simulated data alike.

first-principles calculations↗

Cost-efficient finite-volume high-order schemes for compressible magnetohydrodynamics

We present an efficient dimension-by-dimension finite-volume method which solves the adiabatic magnetohydrodynamics equations at high discretization order, using the constrained-transport approach on Cartesian grids. Results are presented up to tenth order of accuracy. The algorithmic architecture of this method is very close to that of commonly employed second-order schemes: it requires only one reconstructed value per face for each computational cell, independently of the scheme's order. This property is highly beneficial for the numerical efficiency. It results from reusing the required values already available in neighboring grid cells, in contrast to standard algorithms that require a number of reconstructions and evaluations which increases with the scheme's order of accuracy. At a given resolution, these high-order schemes present significantly less numerical dissipation than commonly employed lower-order approaches. Thus, results of comparable accuracy are achievable at a substantially coarser resolution, yielding overall performance gains. We also present a way to include physical dissipative terms: viscosity, magnetic diffusivity and cooling functions, respecting the finite-volume and constrained-transport frameworks. Benefits of this method are shown through applications in turbulent flows.

97 MATHEMATICS AND COMPUTING↗

Monocrystalline CdSeTe/MgCdTe Double‐Heterostructure Solar Cells

This article reports monocrystalline CdSeTe/MgCdTe double‐heterostructure (DH) solar cells with varying Se compositions in the absorber layers that are grown on InSb substrates by using molecular beam epitaxy. The Se composition in the samples studied is determined to be 4%–11% through high‐resolution X‐ray diffraction (XRD) and photoluminescence measurements. Increased Se incorporation induces higher defect density attributed to increased lattice mismatch, and mixed‐phase formation due to the small difference in formation energies between the wurtzite and zinc blende phases of CdSe, which causes stacking faults and grain boundaries. Devices are fabricated by directly depositing an n‐type indium tin oxide (ITO) layer on the CdSeTe/MgCdTe DHs followed by Ag metal contacts. Reduced bandgap is observed in solar cells with increased Se composition. The devices with an absorber containing 4% Se exhibit an average open‐circuit voltage ( V OC ) of 0.925 V, a short‐circuit current density ( J SC ) of 23.4 mA/cm 2 , a fill factor ( FF ) of 0.650 and an efficiency of 14.1% without anti‐reflection coating. Absorbers with higher Se compositions result in poor device performance, mainly attributed to high defect density in the materials.

Ju, Zheng [Center for Photonics Innovation Arizona↗

Optimal 3D chemical imaging with multimodal electron tomography

Accurate mapping of nanoscale chemistry in three dimensions (3D) has been a longstanding challenge. Modern electron microscopy provides chemical images by electron energy loss spectroscopy (EELS) and energy dispersive x-ray spectrometry (EDX) but requires high fluences that damage specimens. In 3D, the requirements are worse; electron tomography demands many high-fluence chemical maps for reconstruction, creating a tradeoff between resolution, accuracy, and sample survival. Fused multimodal electron tomography (MM-ET) alleviates this requirement by leveraging lower-fluence high-angle annular dark-field (HAADF) images alongside a few chemical maps to dramatically improve chemical resolution. Here, experimental and computational parameter space is systematically explored to determine when MM-ET performs best. Ideal imaging conditions balance sample survival with resolution and chemical specificity; we recommend a tilt range of at least ± 70°, acquiring 40 equally spaced HAADF projections (signal-to-noise > 10), and 7 EELS/EDX maps of each chemistry (signal-to-noise > 4).

36 MATERIALS SCIENCE↗

Efficient lattice QCD computation of radiative-leptonic-decay form factors at multiple positive and negative photon virtualities

In previous work [D. Giusti, Methods for high-precision determinations of radiative-leptonic decay form factors using lattice QCD, Phys. Rev. D 107, 074507 (2023)], we showed that form factors for radiative leptonic decays of pseudoscalar mesons can be determined efficiently and with high precision from lattice QCD using the “three-dimensional (3D) method,” in which three-point functions are computed for all values of the current insertion time and the time integral is performed at the data-analysis stage. Here, we demonstrate another benefit of the 3D method: the form factors can be extracted for any number of nonzero photon virtualities from the same three-point functions at no extra cost. We present results for the $D_s → ℓνγ*$ vector form factor as a function of photon energy and photon virtuality, for both positive and negative virtuality, for a single ensemble with 340 MeV pion mass and 0.11 fm lattice spacing. In our analysis, we separately consider the two different time orderings and the different quark flavors in the electromagnetic current. We discuss in detail the behavior of the unwanted exponentials contributing to the three-point functions, as well as the choice of fit models and fit ranges used to remove them for various values of the virtuality. While positive photon virtuality is relevant for decays to multiple charged leptons, negative photon virtuality suppresses soft contributions and is of interest in QCD-factorization studies of the form factors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Asymptotic-state prediction for fast flavor transformation in neutron star mergers

Neutrino flavor instabilities appear to be omnipresent in dense astrophysical environments, thus presenting a challenge to large-scale simulations of core-collapse supernovae and neutron star mergers (NSMs). Subgrid models offer a path forward, but require an accurate determination of the local outcome of such conversion phenomena. Focusing on “fast” instabilities, related to the existence of a crossing between neutrino and antineutrino angular distributions, we consider a range of analytical mixing schemes, including a new, fully three-dimensional one, and also introduce a new machine learning (ML) model. We compare the accuracy of these models with the results of several thousands of local dynamical calculations of neutrino evolution from the conditions extracted from classical NSM simulations. Our ML model shows good overall performance, but struggles to generalize to conditions from a NSM simulation not used for training. The multidimensional analytic model performs and generalizes even better, while other analytic models (which assume axisymmetric neutrino distributions) do not have reliably high performances, as they notably fail as expected to account for effects resulting from strong anisotropies. As a result, the ML and analytic subgrid models extensively tested here are both promising, with different computational requirements and sources of systematic errors.

79 ASTRONOMY AND ASTROPHYSICS↗

Toward a 2D Local Implementation of Quantum Low-Density Parity-Check Codes

Geometric locality is an important theoretical and practical factor for quantum low-density parity-check (qLDPC) codes that affects code performance and ease of physical realization. For device architectures restricted to two-dimensional (2D) local gates, naively implementing the high-rate codes suitable for low-overhead fault-tolerant quantum computing incurs prohibitive overhead. In this work, we present an error-correction protocol built on a bilayer architecture that aims to reduce operational overheads when restricted to 2D local gates by measuring some generators less frequently than others. We investigate the family of bivariate-bicycle qLDPC codes and show that they are well suited for a parallel syndrome-measurement scheme using fast routing with local operations and classical communication (LOCC). Through circuit-level simulations, we find that in some parameter regimes, bivariate-bicycle codes implemented with this protocol have logical error rates comparable to the surface code while using fewer physical qubits. Published by the American Physical Society 2025

Berthusen, Noah (ORCID:0000000275862786)↗

Hybrid Simulations of FRC Merging and Compression

An improved understanding of field-reversed configuration (FRC) merging and stability in high acceleration and compression magnetic fields is needed to speed up the development of the pulsed fusion concept developed at Helion Energy. All previous theoretical and simulation work on FRC merging and compression was performed using two-dimensional (2D) magnetohydrodynamic (MHD) models. The results of novel 2D hybrid simulations (fluid electrons and full-orbit kinetic ions) of FRC merging and compression are presented. Results of kinetic and MHD simulations, computed using the HYM code, are compared and analyzed. In cases without axial magnetic compression, both the MHD and hybrid simulations show a high sensitivity to the initial parameters (i.e. FRC separation, velocity, normalized separatrix radius, and plasma viscosity), showing that FRCs with large elongation and separatrix radius either do not merge or merge partially, forming a doublet FRC. In conclusion, application of a mirror coil field at the FRC ends with increasing strength is shown to lead to fast and complete merging of the FRCs in MHD and kinetic simulations.

FRC↗

Assessing methods in fusion and fitting for time series construction in remote sensing-based earth observations

This study evaluates the comparative performance of spatiotemporal fusion and time-series fitting methods for constructing high-spatiotemporal-resolution remote sensing time-series data. Due to in-class similarity of fusion methods and fitting methods, we employ the Fit-FC (Fitting, spatial Filtering, and residual Compensation) model as a representative fusion method and the linear harmonic fitting model as a representative fitting method. Both Fit-FC and the linear harmonic fitting are widely used for high-spatiotemporal-resolution time-series data construction, and we modify the original Fit-FC model to enable automatic time-series fusion. To ensure data representativeness, we use 3 years (2019–2021) of Harmonized Landsat and Sentinel-2 surface reflectance datasets and Terra MCD43A4 products. Eight experimental regions are selected worldwide to guarantee generalization of the comparative performance between fusion and fitting methods, covering diverse land-use types (cropland, developed land, forest, and grassland) and varying climatological conditions. Time-series of NDVI and surface reflectance are analyzed under both actual observations and simulated data-missing scenarios. The constructed time-series data reveals that (1) the modified Fit-FC and linear harmonic fitting model achieve excellent performance in constructing high-resolution time-series images; (2) the fusion method outperforms the fitting method in constructing time-series of NDVI and surface reflectance images in cropland-, forest-, and grassland-dominated regions; (3) both methods achieve comparable performance in developed-dominated regions; (4) the fusion method is more robust to missing data, and better captures abrupt phenological transitions under conditions of continuous missing data; (5) the fitting method is computationally more efficient, making it suitable for large-scale time-series image reconstruction. This study provides valuable insights for selecting optimal strategies to generate high-resolution time-series images across diverse application scenarios and lays a foundation for extensions to other vegetation indices or land surface variables.

54 ENVIRONMENTAL SCIENCES↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

Assessing Resilience in Lane Detection Methods: Infrastructure-Based Sensors and Traditional Approaches for Autonomous Vehicles

Traditional autonomous vehicle perception subsystems that use onboard sensors have the drawbacks of high computational load and data duplication. Infrastructure-based sensors, which can provide high quality information without the computational burden and data duplication, are an alternative to traditional autonomous vehicle perception subsystems. However, these technologies are still in the early stages of development and have not been extensively evaluated for lane detection system performance. Therefore, there is a lack of quantitative data on their performance relative to traditional perception methods, especially during hazardous scenarios, such as lane line occlusion, sensor failure, and environmental obstructions. We address this need by evaluating the influence of hazards on the resilience of three different lane detection methods in simulation: (1) traditional camera detection using a U-Net algorithm, (2) radar detections using infrastructure-based radar retro-reflectors (RRs), and (3) direct communication of lane line information using chip-enabled raised pavement markers (CERPMs). The performance of each of these methods is assessed using resilience engineering metrics by simulating the individual methods for each sensor technology’s response to related hazards in the CARLA simulator. Using simulation techniques to replicate these methods and hazards acquires extensive datasets without lengthy time investments. Specifically, the resilience triangle was used to quantitatively measure the resilience of the lane detection system to obtain unique insights into each of the three lane detection methods; notably the infrastructure-based CERPMs and RRs had high resistance to hazards and were not as easily affected as the vision-based U-Net. However, while U-Net was able to recover the fastest from the disruption as compared to the other two methods, it also had the most performance loss. Overall, this study demonstrates that while infrastructure-based lane keeping technologies are still in early development, they have great potential as alternatives to traditional ones.

Patil, Pritesh↗