Search NASA⌕ Search

SEARCH · Search NASA

Results for “supercomputing technologies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

86 records · Page 5

Enabling Low-Temperature (LTP) Ignition Technologies for Multi-Mode Engines through the Development of a Validated High-Fidelity LTP Model for Predicative Simulations Tools

The goal of multi-mode engine architectures is to extend current lean-burn dilution limits with renewable fuels, which requires spark plugs to deposit high energies (hundreds of mJ) in order to initiate ignition and complete combustion. At elevated energy deposition rates, spark plugs experience increased electrode erosion and thermal losses, which ultimately shortens the spark-plug lifetime and lowers ignition efficiency. As such, in order to safeguard the efficiency gains of multi-mode concepts, new and improved ignition technologies are required. Recently, non-equilibrium low-temperature plasmas (LTP) have been shown to promote energy-efficient ignition via quenching and transport of electronically excited atoms and molecules, selective radical production and fast heating of hydrocarbon/air mixtures [1-2]. Thus, LTP is seen as a technology that can potentially improve the energy extraction efficiency of fuels, while enabling kinetically controlled combustion modes towards fuel leaner conditions to realize current DOE VTO goals of improving the sustainability of future mobility [3]. Although many previous studies have demonstrated the efficacy of plasma-assisted ignition to enhance combustion, the detailed enhancement mechanisms remain largely unknown, especially for oxygenated fuels and at elevated pressures that are most relevant to practical engine conditions. These barriers hinder the development of accurate and comprehensive numerical models that seek to describe LTP-based ignition in existing engine design software tools and methods. Current state-of-the-art simulation capabilities for LTP ignition systems are in need of improvements since they deliver qualitative results only due to important limitations of existing approaches. Firstly, validated kinetic models with elementary steps for plasma discharges in oxygenated fuel/air mixtures of relevance to the transportation sector are required. Such kinetic models do not exist at present and will be developed and validated within this project. Secondly, plasma discharges and reactive mixture ignition are multi-scale, unsteady processes requiring high-performance numerical methods and software that execute efficiently on DOE supercomputers. Such software does not exist at present and will be developed and applied to practical LTP ignition scenarios as part of this project. Thirdly, experimental databases that are tailored to serve as benchmark in support of the development of predictive computational models of LTP ignition do not exist and will be part of this project.

33 ADVANCED PROPULSION SYSTEMS↗

Modeling Large Dust Aerosols in the Community Earth System Model Version 2 (CESM2)

Dust aerosols have a wide size distribution from less than 0.1 to over 100 μm and dominate Earth's atmospheric aerosol mass. However, most Earth system models (ESMs) inadequately represent dust aerosols larger than 10 μm in diameter, limiting the accuracy of the simulated dust cycle and climate impacts. Here, we introduce a new modeling framework that captures the full observed size distribution of dust aerosols, incorporating recent advances into a mineral-resolved version of the Community ESM, while addressing known issues in previous versions. Comprehensive evaluation against diverse observations of bulk dust and component minerals demonstrates that the model reproduces the observed dust cycle across particle sizes. Incorporating the previously unrepresented large-dust fractions substantially alters dust budget estimates, highlighting potential changes in simulated climate impacts and underscoring the importance of comprehensive size-resolved dust modeling. Despite these advancements, uncertainties persist. Our results indicate that a size-dependent reduction in settling velocity is required to reproduce the observed dust size distribution downwind of source regions. Specifically, in the new model, the gravitational settling velocity of dust particles larger than 10 μm in diameter must be reduced by as much as 85% to achieve agreement with observations. This empirical reduction serves as a constraint on physics-based models of dust settling. Future developments should address misrepresented physical processes that hinder accurate modeling of the large dust aerosol transport. Expanding observational data sets covering the full-size distribution is also essential to better constrain the dust cycle and improve the representation of dust optical properties and climate effects.

Li, Longlei [Cornell Univ., Ithaca, NY (United Sta↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

HydraGNN v4.0

The new version of HydraGNN v4.0 provides additional core capabilities, such as: Inclusion of multi-body atomistic cluster expansion MACE, polarizable atom interaction neural network PAINN, and equivariant principal neighborhood aggregation (PNAEq) among the message passing layers supported -Inclusion of graph transformers to directly model long-range interactions between nodes that are distant in the graph topology Integration of graph transformers with message passing layers by combining the graph embedding generated by the two mechanisms, which allows for an improved expressivity of the HydraGNN architecture Improved re-implementation of multi-task learning (MTL) to allow its use for stabilized training across imbalanced, multi-source, multi-fidelity data Introduction of multi-task parallelism, a newly proposed type of model parallelism specifically for MTL architectures, which allows to dispatch different output decoding heads to different GPU devices Integration of multi-task parallelism with pre-existing distributed data parallelism to enable a 2D parallelization for distributed training Improved portability of the distributed training across Intel GPUs, which has been testes on ALCF exascale supercomputer Aurora Inclusion of 2-level fine-grained energy profilers portable across NVIDIA, AMD, and Intel GPUs to monitor the power and energy consumption associated with different functions executed by the HydraGNN code during data pre-load and training Restructuring of previous examples and inclusion of new sets of examples to illustrate the download, preprocess, and training of HydraGNN models on new large-scale open-source datasets for atomistic materials modeling (e.g., Alexandria, Transition1x, OMat24, OMol25)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

DAmodel: hierarchical Bayesian modelling of DA white dwarfs for spectrophotometric calibration

We use hierarchical Bayesian modelling to calibrate a network of 32 all-sky faint DA white dwarf (DA WD) spectrophotometric standards (⁠16.5 < V , 19.5⁠) alongside three CALSPEC standards, from 912 Å to 32 μm. The framework is the first of its kind to jointly infer photometric zero points and WD parameters (surface gravity log g⁠, effective temperature T eff ⁠, extinction A V ⁠, dust relation parameter R V ) by simultaneously modelling both photometric and spectroscopic data. We model panchromatic Hubble Space Telescope Wide Field Camera 3 (HST/WFC3) UVIS and IR photometry, HST/STIS UV spectroscopy, and ground-based optical spectroscopy to sub-per cent precision. Photometric residuals for the sample are the lowest yet yielding < 0.004 mag RMS on average from the UV to the NIR, achieved by jointly inferring time-dependent changes in system sensitivity and WFC3/IR count-rate nonlinearity. Our GPU-accelerated implementation enables efficient sampling via Hamiltonian Monte Carlo, critical for exploring the high-dimensional posterior space. The hierarchical nature of the model enables population analysis of intrinsic WD and dust parameters. Inferred spectral energy distributions from this model will be essential for calibrating the James Webb Space Telescope as well as next-generation surveys, including Vera Rubin Observatory’s Legacy Survey of Space and Time and the Nancy Grace Roman Space Telescope.

methods: statistical↗

Small-scale properties from exascale computations of turbulence on a $\mathbf{32\,768^3}$ periodic cube

To study the physics of small-scale properties of homogeneous isotropic turbulence at increasingly high Reynolds numbers, direct numerical simulation results have been obtained for forced isotropic turbulence at Taylor-scale Reynolds number R λ = 2500 on a 32 768 3 three-dimensional periodic domain using a GPU pseudo-spectral code on a 1.1 exaflop GPU supercomputer (Frontier). These simulations employ the multi-resolution independent simulation (MRIS) technique (Yeung & Ravikumar 2020, Phys. Rev. Fluids, vol. 5, 110517) where ensemble averaging is performed over multiple short segments initiated from velocity fields at modest resolution, and subsequently taken to higher resolution in both space and time. Reynolds numbers are increased by reducing the viscosity with the large-scale forcing parameters unchanged. Although MRIS segments at the highest resolution for each Reynolds number last for only a few Kolmogorov time scales, small-scale physics in the dissipation range is well captured – for instance, in the probability density functions and higher moments of the dissipation rate and enstrophy density, which appear to show monotonic trends persisting well beyond the Reynolds number range in prior works in the literature. Attainment of range of length and time scales consistent with classical scaling also reinforces the potential utility of the present high-resolution data for studies of short-time-scale turbulence physics at high Reynolds numbers where full-length simulations spanning many large-eddy time scales are still not accessible. A single snapshot of the 32 768 3 data is publicly available for further analyses via the Johns Hopkins Turbulence Database.

intermittency↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

Acceleration of the particle-in-cell code Osiris with graphics processing units

Fully relativistic particle-in-cell (PIC) simulations are crucial for advancing our knowledge of plasma physics. Modern supercomputers based on graphics processing units (GPUs) offer the potential to perform PIC simulations of unprecedented scale, but require robust and feature-rich codes that can fully leverage their computational resources. In this work, this demand is addressed by adding GPU acceleration to the PIC code Osiris. An overview of the algorithm, which features a CUDA extension to the underlying Fortran architecture, is given. Detailed performance benchmarks for thermal plasmas are presented, which demonstrate excellent weak scaling on NERSC's Perlmutter supercomputer and high levels of absolute performance. The robustness of the code to model a variety of physical systems is demonstrated via simulations of Weibel filamentation and laser-wakefield acceleration run with dynamic load balancing. Finally, measurements and analysis of energy consumption are provided that indicate that the GPU algorithm is up to ~14 times faster and ~7 times more energy efficient than the optimized CPU algorithm on a node-to-node basis. The described development addresses the PIC simulation community's computational demands both by contributing a robust and performant GPU-accelerated PIC code and by providing insight into efficient use of GPU hardware.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The DELVE Quadruple Quasar Search. I. A Lensed Low-luminosity Active Galactic Nucleus

A quadruply lensed source, J125856.3–031944, has been discovered using the DELVE survey and Wide-field Infrared Survey Explorer W1–W2 colors. Follow-up direct imaging carried out with the Magellan Baade 6.5 m telescope is analyzed, as is spectroscopy from the 2.5 m Nordic Optical Telescope. The lensed image configuration is kite-like, with the major axis of the lensing galaxy along the symmetry axis of the kite, and with the faintest image at its tail. Redward of 6000 Å, the tail image is strongly blended with the lensing galaxy. The Sloan g direct imaging carried out with Magellan permits deblending. As the lensed image configuration is nearly circular, simple models give high predicted magnifications for all four images. The source’s narrow emission lines at redshift z = 2.225 and low intrinsic luminosity qualify it as a type 2 active galactic nucleus. The Magellan image shows a substantial residual that suggests a second lensing galaxy.

79 ASTRONOMY AND ASTROPHYSICS↗

What to Support When You’re Compressing

Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.

Error-Bounded Lossy Compression↗

Dark Energy Survey Year 6 results: cell-based coadds and METADETECTION weak lensing shape catalogue

We present the metadetection weak lensing galaxy shape catalogue from the 6-yr Dark Energy Survey (DES Y6) imaging data. This data set is the final release from DES, spanning 4422 deg 2 of the southern sky. We describe how the catalogue was constructed, including the two new major processing steps, cell-based image coaddition, and shear measurements with metadetection. The DES Y6 M etadetection weak lensing shape catalogue consists of 151 922 791 galaxies detected over riz bands, with an effective number density of n eff = 8.22 galaxies per arcmin 2 and shape noise of σ e = 0.29. We carry out a suite of validation tests on the catalogue, including testing for point spread function (PSF) leakage, testing for the impact of PSF modelling errors, and testing the correlation of the shear measurements with galaxy, PSF, and survey properties. In addition to demonstrating that our catalogue is robust for weak lensing science, we use the DES Y6 image simulation suite to estimate the overall multiplicative shear bias of our shear measurement pipeline. We find no detectable multiplicative bias at the roughly half-per cent level, with m = (3.4 ± 6.1) x 10 –3 , at 3σ uncertainty. This is the first time both cell-based coaddition and Metadetection algorithms are applied to observational data, paving the way to the Stage-IV weak lensing surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Milestone in predicting core plasma turbulence: successful multi-channel validation of the gyrokinetic code GENE

On the basis of several recent breakthroughs in fusion research, many activities have been launched around the world to develop fusion power plants on the fastest possible time scale. In this context, high-fidelity simulations of the plasma behavior on large supercomputers provide one of the main pathways to accelerating progress by guiding crucial design decisions. When it comes to determining the energy confinement time of a magnetic confinement fusion device, which is a key quantity of interest, gyrokinetic turbulence simulations are considered the approach of choice – but the question, whether they are really able to reliably predict the plasma behavior is still open. The present study addresses this important issue by means of careful comparisons between state-of-the-art gyrokinetic turbulence simulations with the GENE code and experimental observations in the ASDEX Upgrade tokamak for an unprecedented number of simultaneous plasma observables.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗