Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

NEML2: An efficient and modular multiphysics constitutive modeling library for hybrid computing environments

This paper presents NEML2, an open-source, high-performance library developed for constitutive material modeling, designed to support the flexible and modular development of models for complex material behavior. Building on the foundational structure of its predecessor, NEML, the NEML2 library introduces significant improvements, including enhanced vectorization, automatic differentiation, and seamless integration with PyTorch, facilitating the application of machine learning techniques in material simulations. NEML2 provides a C++ backend with Python bindings, enabling users to create custom material models that can be executed efficiently on both CPU and GPU platforms. The library also supports coupling with Multiphysics simulation frameworks like MOOSE, making it suitable for realistic simulations involving coupled physical processes. Rigorous quality assurance through unit and regression testing ensures the reliability of results, while the extensible, user-friendly design encourages collaboration and reproducibility across the scientific community. This paper provides an overview of NEML2’s architecture, core features, and applications, highlighting its impact on accelerating material qualification and advancing computational methods in materials science.

GPU↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Operational Analytics Studies for ATLAS Distributed Computing: Data Popularity Forecast and Utilization of the WLCG Centers

Operational analytics is the direction of research related to the analysis of the current state of computing processes and the prediction of future states in order to anticipate imbalances and take timely measures to stabilize a complex system. There are two relevant areas in ATLAS Distributed Computing that are currently the focus of studies: user physics analysis including the forecast of popularity of data samples among users, and evaluating WLCG centers for their readiness to process user analysis payloads. Studying these areas is challenging due to the complexity involved, as it requires a comprehensive understanding of numerous boundary conditions typically found in large-scale distributed computing infrastructures. Forecasts of data popularity are problematic without the categorization of user tasks by their types (data transformation or physics analysis), which do not always appear on the surface but may induce noise, which introduces significant distortions for predictive analysis. Evaluating the WLCG resources by their analysis workloads is also a challenging task as it is necessary to find a balance between the workload of the resource, its performance, the waiting time for jobs on it, as well as the volume of jobs that it processes. This is especially difficult in a heterogeneous computing environment, where legacy resources are used along with modern high-performance machines. We will look at these areas of research in detail and discuss what tools and methods are used in our work, demonstrating results already obtained.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

How much is surface dopant enough to maximize CO2-to-liquid chemicals conversion at industrially relevant current density?

This invited talk will be presented at the symposium "Waste Feedstock to Fuels and Petrochemicals”, Fuels and Petrochemical Division at the 2024 AIChE Annual Meeting. In this study, we will discuss how minute amount of surface heteroatoms would maximize CO2 reduction to formate/formic acid at industrially relevant current densities and high selectivity which was validated by both experimental and computational studies. The outstanding performance of the best-in-class catalysts in both H-cell and full-cell electrolyzer cell is also demonstrated. Our findings would provide additional design concepts of high performance CO2R electrocatalysts.

Nguyen Phan, Thuy Duong↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

An enrichment wall modeling framework for spectral element methods

In the present work, a first-of-its-kind enrichment wall-model is developed within the spectral element method (SEM) framework for large-eddy simulations (LES) of wall-bounded turbulent flows. The method augments the polynomial solution in the wall-adjacent elements with an analytical law-of-the-wall enrichment function representing the mean velocity near the wall. In the solution representation, this enrichment function captures the large gradients in the boundary layer, which allows the polynomial modes to represent the turbulent fluctuations. The enriched solution is able to resolve the shear stress at the wall without any modification to the no-slip wall boundary conditions, which allows for greater accuracy in the near-wall region compared to traditional methods. The enrichment wall modeling approach is implemented in a high-order SEM computational fluid dynamics solver, Nek5000, and its performance is assessed in turbulent channel flow wall-modeled LES for a range of Reynolds numbers. It is demonstrated that the enrichment wall-model improves solution accuracy on under-resolved near-wall grids as compared to traditional shear stress wall-models.

42 ENGINEERING↗

Solving key challenges in collider physics with foundation models

Foundation models are neural networks that are capable of simultaneously solving many problems. Large language foundation models like ChatGPT have revolutionized many aspects of daily life, but their impact for science is not yet clear. In this paper, we use a new foundation model for hadronic jets to solve three key challenges in collider physics. In particular, we show how experiments can (1) save significant computing power when developing reconstruction algorithms, (2) perform a complete uncertainty quantification for high-dimensional measurements, and (3) search for new physics with model agnostic methods using low-level inputs. In each case, there are significant computational or methodological challenges with current methods that limit the science potential of deep learning algorithms. By solving each problem, we take jet foundation models beyond proof-of-principle studies and into the toolkit of practitioners.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Space Charge Simulations of High Intensity Proton Beams in the AGS Booster

Computer simulation studies have been performed to understand the beam behavior and to explore intensity limitations of proton beams in the AGS Booster at higher beam intensities. During the 100 GeV polarized proton operations of RHIC Run 2024, sPHENIX operated in modes with a crossing angle at collisions in order to mitigate beam-beam effects. Three different running modes were employed: (a) sPHENIX operated with a negative (-2 mrad) crossing angle, and STAR operated with 0 mrad. Both experiments were brought into collisions at the start of the store. (b) sPHENIX was brought into collisions with 0 mrad first. Then STAR was brought into collisions after the beam-beam parameter from sPHENIX reduced to below $10 \times 10^{-3}$. (c) sPHENIX operated with a positive (+1.5 mrad) crossing angle, and STAR operated with 0 mrad. Both experiments were brought into collisions at the start of the store. The collisions with a crossing angle of up to $\pm 2$ mrad, as in running modes (a) and (c), lead to large Piwinski angle in the new sPHENIX detector, which reduces luminosity if other parameters are unchanged. There are two ways to compensate the reduction in luminosity: squeeze $\beta^{*}$ if there is sufficient dynamic aperture, or increase the injected beam intensity. The first part of polarized proton operation during RHIC Run 2024 was dedicated to increasing the intensity. Different configurations were tested with crossing angle and lattice adjustments on RHIC. At the same time, new injector configurations were developed and tested in an effort to push for both higher intensity and better quality of the beam injected into RHIC. When the beam intensity is increased, space charge is a concern particularly in the lower energy stages of acceleration, such as during the injection and the early part of the Booster cycle, which could become a dominant effect in limiting the intensity of the beam that can be delivered to RHIC.

43 PARTICLE ACCELERATORS↗

Density functional theory-based surrogate kinetic models for heterogeneous reactions of hydrocarbon intermediates on silicon carbide

The increasing demand for high-performance materials in advanced technologies highlights the importance of achieving a fundamental understanding and potential control of silicon carbide (SiC) deposition processes. However, existing models often lack sufficient theoretical detail, relying heavily on empirical data and offering limited predictive capability. In particular, the complex surface chemistry governing SiC growth remains poorly understood. This study addresses these challenges by employing density functional theory (DFT) to investigate key heterogeneous reactions involving hydrocarbon intermediates on SiC surfaces, including dehydrogenation, hydrogenation, and carbon deposition. Transition state searches were conducted to identify reaction pathways and energy barriers. While first-principles calculations offer high accuracy, they are computationally intensive. To extend the utility of these first-principles results, vibrational analyses were performed using phonon-based statistical thermochemistry to compute temperature-dependent reaction rates which were used to develop Arrhenius-type surrogate kinetic models. Furthermore, the resulting framework provides a more rigorous, physically grounded basis for integrating atomistic insights into continuum-scale modeling, ultimately enabling improved prediction and optimization of SiC film growth in high-performance material systems.

Density Functional Theory↗

AGS-GNN: Attribute-guided Sampling for Graph Neural Networks

We propose AGS-GNN, a novel attribute-guided sampling algorithm for Graph Neural Networks (GNNs) that exploits node features and connectivity structure of a graph while simultaneously adapting for both homophily and heterophily in graphs. (In homophilic graphs vertices of the same class are more likely to be connected, and vertices of different classes tend to be linked in heterophilic graphs.) While GNNs have been successfully applied to homophilic graphs, their application to heterophilic graphs remains challenging. The best-performing GNNs for heterophilic graphs do not fit the sampling paradigm, suffer high computational costs, and are not inductive. We employ samplers based on feature-similarity and feature-diversity to select subsets of neighbors for a node, and adaptively capture information from homophilic and heterophilic neighborhoods using dual channels. Currently, AGS-GNN is the only algorithm that we know of that explicitly controls homophily in the sampled subgraph through similar and diverse neighborhood samples. For diverse neighborhood sampling, we employ submodularity, which was not used in this context prior to our work. The sampling distribution is pre-computed and highly parallel, achieving the desired scalability. Using an extensive dataset consisting of 35 small (<=100K nodes) and large (>100K nodes) homophilic and heterophilic graphs, we demonstrate the superiority of AGS-GNN compare to the current approaches in the literature. AGS-GNN achieves comparable test accuracy to the best-performing heterophilic GNNs, even outperforming methods using the entire graph for node classification. AGS-GNN also converges faster compared to methods that sample neighborhoods randomly, and can be incorporated into existing GNN models that employ node or graph sampling.

artificial intelligence↗

International fuel performance study of fresh fuel experiments for PCMI effects during RIA experiments

This paper presents the results of High-burnup Experiments for Reactivity-initiated Accident (HERA) Modeling & Simulation (M&S) exercise. The HERA project under the Nuclear Energy Agency (NEA) Second Framework for Irradiation Experiments (FIDES-II) program is focused on studying Light Water Reactor (LWR) fuel behavior during Reactivity-Initiated Accident (RIA) conditions. The Part I M&S cases are based on a series of tests in the Transient Reactor Test (TREAT) facility in the United States and the Nuclear Safety Research Reactor (NSRR) in Japan. The purpose of this work is to evaluate the test design to accomplish its goals in establishing clearer understanding of the effects of power pulse width during RIA conditions. Further, the blind predictions using various computational tools have been performed and compared amongst to interpret the behaviors of high burnup fuels during RIA. While many international participants evaluate the thermal–mechanical behavior of fuel rod under different conditions, a considerable scatter of outputs comes out for the cases due to the disparity between codes in predicting mechanical behaviors. In general, however, the results of thermal–mechanical analysis elaborate that nominal design conditions the shorter pulse width tests in NSRR should cause cladding failures while the TREAT tests appear to have more split prediction of failure or not. Furthermore, the sensitivity analysis varying key testing parameters reveals the considerable effect of power pulse width and total energy deposition on prediction of fuel rod failure.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Remark on Algorithm 1012: Computing Projections with Large Datasets

In ACM TOMS Algorithm 1012, the DELAUNAYSPARSE software is given for performing Delaunay interpolation in medium to high dimensions. When extrapolating outside the convex hull of the training set, DELAUNAYSPARSE calls the nonnegative least squares solver DWNNLS to compute projections onto the convex hull. However, DWNNLS and many other available sum-of-squares optimization solvers were not intended for usage with many variable problems, which result from the large training sets that are typical in machine learning applications. Thus, a new PROJECT subroutine is given, based on the highly customizable quadratic program solver BQPD. This solution is shown to be as robust as DELAUNAYSPARSE for projection onto both synthetic and real-world datasets, where other available solvers frequently fail. Although it is intended as an update for DELAUNAYSPARSE, due to the difficulty and prevalence of the problem, this solution is likely to be of external interest as well.

97 MATHEMATICS AND COMPUTING↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Learning nonlinear operators in latent spaces for real-time predictions of complex dynamics in physical systems

Abstract Predicting complex dynamics in physical applications governed by partial differential equations in real-time is nearly impossible with traditional numerical simulations due to high computational cost. Neural operators offer a solution by approximating mappings between infinite-dimensional Banach spaces, yet their performance degrades with system size and complexity. We propose an approach for learning neural operators in latent spaces, facilitating real-time predictions for highly nonlinear and multiscale systems on high-dimensional domains. Our method utilizes the deep operator network architecture on a low-dimensional latent space to efficiently approximate underlying operators. Demonstrations on material fracture, fluid flow prediction, and climate modeling highlight superior prediction accuracy and computational efficiency compared to existing methods. Notably, our approach enables approximating large-scale atmospheric flows with millions of degrees, enhancing weather and climate forecasts. Here we show that the proposed approach enables real-time predictions that can facilitate decision-making for a wide range of applications in science and engineering.

97 MATHEMATICS AND COMPUTING↗