Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

MC Formula Protocol for H35HF Fueling (CRADA Final Report)

The National Renewable Energy Lab (NREL), Frontier Energy, and the industry partners worked together to help SAE J2601-5 develop an H35 high-flow (HF) medium-duty (MD) and heavy-duty (HD) fueling protocol. The team upgraded NREL's hydrogen filling simulations (H2FillS) model to accommodate an MC Formula fueling (t-final) table generation capability by leveraging NREL's high-performance computing system. Based on protocol boundary conditions (e.g., allowable maximum flow rate, range of storage system size) set by SAE J2601-5, the team generated the fueling tables and then validated the reliability of those tables by installing them on NREL's HD dispenser and ZBT's H35HF dispenser and then performing H35HF fueling experiments. Through the validation process, this team certified that the fueling tables generated were reliable to install in commercial H35HF dispensers and then performed H35HF fueling of commercial MD/HD vehicles.

08 HYDROGEN↗

FY25 MOOSE Usability Improvements: 3D Meshing Capabilities, Initiation of Geometry Support for Monte Carlo Tools, and Enhancement of MOOSE/Workbench User Input Interactions

Usability improvements have been made to MOOSE and Workbench in FY25 to enhance usability and user workflows. Assorted enhancement have been made to MOOSE’s intrinsic meshing capabilities in order to enable more flexible and complex meshing of nuclear reactor systems, in particular for 3D applications. Mesh generators have been added to perform operations such as batch mesh generation, surface mesh generation, and creation of 3D transition layers. These mesh generation capabilities make it much easier to generate high quality non-extruded 3D meshes. Additionally, work to integrate Monte Carlo reactor physics simulations into MOOSE-based multi-physics workflows has reached another milestone with the implementation of the Constructive Solid Geometry (CSG) base framework. This framework lays the foundation for mesh generators to offer the user a generic CSG output option (as opposed to a finite element mesh). To support users, workshop on the MOOSE Reactor Module was delivered which featured hands-on examples using the NEAMS Workbench on INL’s High Performance Computing system. Recent updates to the NEAMS Workbench, WASP, and the MOOSE language server have introduced several improvements aimed at making MOOSE-based simulation setup and input management faster, more accurate, and easier to use. Key capabilities that have been added include multi-tab-stop autocompletion, visual input diagnostics, developer-directed data visualizations, upgraded ParaView integration, and Workspace-level file tracking. Together, these changes make it easier for users to build, validate, and manage complex MOOSE-based simulation models — especially those involving reusable components, included files, and datasets. The improvements are designed to save time, reduce input errors, and help users get to a successful simulation run faster, with more confidence in the results.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

2025 Atmospheric Radiation Measurement (ARM) Annual Report

ARM is a multi-laboratory, U.S. Department of Energy (DOE) Office of Science user facility and a key contributor to atmospheric research efforts. For more than 30 years, ARM has supported the DOE and Office of Science missions by providing atmospheric observations to enable scientific discovery, transform our understanding of the atmosphere, and evaluate and improve the accuracy of atmospheric models. ARM’s cutting-edge capabilities for scientists include continuously operating ground-based observatories, aerial observation platforms, and high performance computing.

54 ENVIRONMENTAL SCIENCES↗

Summer Internship Report: ARA2 Benchmarking

Over the past decade, the RISC-V Instruction Set Architecture (ISA) has emerged as a significant player in both academic and industrial processor design due to its open-source nature, modular extension system, and versatility across domains ranging from microcontrollers to high-performance computing (HPC). One of its most important recent advancements is the RISC-V Vector Extension (RVV), which enables explicit data-level parallelism through vector registers and vectorized instructions. Unlike traditional SIMD (Single Instruction, Multiple Data) architectures that fix vector lengths at design time, RVV uses the concept of VLEN (vector register length) as a hardware-independent parameter and allows software to adapt dynamically to the available vector width. This flexible approach ensures portability across implementations while enabling scalable performance. The ARA2 core is a parameterizable RISC-V vector processor developed at the Integrated Systems Lab at ETH Zürich and the University of Bologna. Designed as a tightly-coupled accelerator to a scalar RISC-V core, ARA2 implements the RVV 1.0 specification and offers tunable architectural parameters such as the number of vector lanes, VLEN, and cache sizes.

97 MATHEMATICS AND COMPUTING↗

HPC Resource Allocation Under Energy Constraints

We discuss the new problem faced by High-Performance Computing (HPC) facilities in allocating resources to users of their facilities: while facilities once allocated a single finite resource—node-hours—now facilities must also concurrently allocate a second scarce resource: electrical energy, which is bounded within each facility's annual operations budget. Current application optimization practices encourage conservation of the first resource, but can be potentially unaffordably wasteful of the second. We describe a framework for reasoning about such allocations that can be utilized by facilities to articulate policy, while encouraging scientific application developers to write code mindfully of both constraints. We outline the requirements on facilities, on developers, and on hardware vendors and integrators that are necessary to enable the implementation of this framework.

97 MATHEMATICS AND COMPUTING↗

Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science

High Performance Computing (HPC) centers, such as the Oak Ridge Leadership Computing Facility (OLCF), provide advanced infrastructure that enables scientific research at extreme scale. These centers operate with unique hardware configurations, specialized software environments, and elevated security re quirements that differ substantially from what most users encounter on their local systems. As a result, users often develop customized digital artifacts that are tightly coupled to the specific configuration of a given HPC center. Although necessary, this practice can lead to significant duplication of effort as multiple users independently create similar solutions to common problems.

97 MATHEMATICS AND COMPUTING↗

Benchmarking a Defeatured Geant4 NIF Model with HTOAD and HNED Neutron Data

This study benchmarks a simplified Geant4 NIF Target Chamber (TC) model against foil measurements in two instruments: the HTOAD and HNED Snout. The number of product atoms per source neutron per gram ( N 0 /n/g ) is compared across multiple locations and material configurations. The model reproduces high energy threshold reactions dominated by 14 MeV neutrons to within a few percent, depending on location. Discrepancies occur where scattering and moderation are significant for low energy threshold reactions. The defeatured Geant4 model achieves comparable statistical precision with ∼400 times less CPU time than a full fidelity TC model in MCNP. The results can be produced locally on a laptop without the need for high performance computing. Low energy biasing can be improved by preserving room return pathways and implementing IRDFF cross section evaluations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

BISON: A Finite Element-Based Nuclear Fuel Performance Code

BISON is a finite element-based nuclear fuel performance code applicable to a variety of fuel forms including light water reactor fuel rods, TRISO particle fuel, and metallic rod and plate fuel. It is a multiphysics fuel analysis tool that solves fully-coupled thermomechanical problems. BISON is based on MOOSE and can efficiently solve problems using standard workstations or very large high-performance computers in a variety of different dimensions, including full 3D, 2D-RZ axisymmetric, layered axisymmetric 1D, and spherically symmetric 1D systems. It is developed by a team of scientists and engineers at Idaho National Laboratory and by collaborators. The development of BISON is supported by various funding agencies, principally the United States Department of Energy.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Additive Manufactured Composite Phase-Change Material for Thermal Energy Storage Applications

Phase-change materials play a critical role in industrial energy storage applications to drive efficiency improvements, thermal energy management, and carbon emissions reductions. Recently, it has been shown that rapid solidification of alloys with metastable immiscibility in the liquid phase has the potential to form unique microstructures in which a low-melting phase is uniformly distributed in a high-melting matrix. This feature can be exploited using additive manufacturing to produce components with complex geometries containing such unique phase-change microstructures. Phase-field simulations utilizing high-performance computing were used to provide a detailed description of the evolution of the active phase during service in terms of their morphology and composition in different polycrystalline matrix grain morphologies that are typically produced during additive manufacturing. Phase field simulations were performed using, MEUMAPPS-SL (Microstructure Evolution Using Massively Parallel Phase-field Simulations – Solid Liquid) code that was developed in-house by the Oak Ridge National Laboratory. The simulations utilized the capabilities of the Kestrel supercomputer at the National Renewable Energy Laboratory. The simulation results were compared with experimental results generated at Siemens Energy, Inc. The results indicate that the kinetics of liquid spreading along grain boundaries is largely determined by the mobility of the triple line along the intersection of the grain boundary liquid and the grain boundary plane.

25 ENERGY STORAGE↗

Providing Thermal Stability for an Exascale Supercomputer: A Case Study of Frontier's Cooling System

High performance computing (HPC) systems frequently produce large dynamic power swings, even under typical operating conditions, that can present a significant challenge for their direct-liquid cooling systems. Further, the primary cooling loops that must remove this waste heat have response times measured in minutes while the underlying HPC component thermal stress is measured in seconds. The per-socket power demand for both compute processing units (CPUs) and graphic processing units ( GPUs) continues to increase with each successive generation while case temperatures are declining. New HPC systems are expected to exacerbate the challenge of these dynamic power swings and the impact on effective and timely cooling systems. This paper describes the cooling and controls system for Oak Ridge National Laboratory’s Frontier Supercomputer, the first sustained exascale system, as a case study for this situation. The cooling and control system for Frontier demonstrates specific success, but with a number of trade-offs and decisions that suggest further design and operating optimizations for the community at large to consider.

42 ENGINEERING↗

Aviation Fuel Characterization at Operationally Relevant Conditions

To accelerate approval and potentially expand the allowable property range for aviation fuels, we are using high performance computing simulations to reveal fuel property effects on aviation combustor performance. These simulations are supported by fuel property measurements over temperatures and pressures that the fuel experiences in an aircraft engine and by validated chemical kinetics models for SAF combustion. Here we report density, viscosity, and surface tension results for conventional jet fuel and multiple synthetic fuels from -30 degrees Celsius to 200 degrees Celsius (-40 degrees Celsius for viscosity) and 1 atm to 70 atm including an assessment of method repeatability. Properties of surrogate mixtures are also investigated. Distillation, ICN, LHV, flashpoint, and Cp are also reported, and data are being used to develop models to predict fuel properties from composition (GCxGC).

33 ADVANCED PROPULSION SYSTEMS↗

Flow Reactor Study and Kinetic Model Development of HEFA-SPK and its Surrogate

In this work, we formulate a two-component surrogate for HEFA-SPK, incorporating aromatic or cycloalkane components, and develop reduced kinetic models for the surrogates to be used in high-performance computing simulations. The HEFA-SPK surrogate was selected and optimized based on the fuel's physical and combustion properties, including ignition delay times and flame speeds. The resulting surrogate consists of 40% n-undecane and 60% 2-methylnonane. The surrogate was confirmed by flow reactor experiments for both the HEFA-SPK fuel and the suggested two-component surrogates, where excellent agreement was observed. To meet aromatic requirements, 1,2,4-trimethylbenzene (8%) was selected, and we determined that incorporating 30% propylcyclohexane into the HEFA-SPK will achieve a volume swell equivalent to 8% aromatics. The properties of the surrogates were measured, and a new reduced kinetic model was developed based on the semi-decoupling methodology, using a reduced CH4 chemistry from NUIG 1.0 as base chemistry. Kinetic models will be employed in combustor simulations to enable a comprehensive understanding of the effect of SAF fuel properties on aviation combustor performance.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

OptiBench: An Optimization Benchmark Tool for Renewable Energy Problems

We propose a benchmark framework and visualization tool, OptiBench, for analyzing the performance of state-of-the-art optimization solvers across a variety of optimization problems in renewable energy research. Our framework is designed from the ground up in the Julia programming language and enables analysis at scale on high performance computing (HPC) systems. Our visualization tool allows effortless evaluation of optimization solver performance, robustness, and accuracy through intuitive plots, e.g., performance profiles, heat maps, and distribution plots. We have tested three benchmark suites relevant to the modeling of renewable energy systems, viz., CUTEst, PGLib-OPF, and WaterTAP water treatment optimization problems. We illustrate benchmarking of CUTEst using OptiBench on the National Laboratory of the Rockies's (NLR) HPC Kestrel. Our findings indicate that MA57 HSL linear solver demonstrated the best overall performance for an experimental IPOPT implementation. Our work is ongoing and we intend to add support for more optimization solvers and benchmark test suites in the future.

97 MATHEMATICS AND COMPUTING↗

Phase-field predictions of the influence of cooling rates during AM on the Evolution of Microstructures in Nickel-Based Single Crystal Superalloys

Additive manufacturing of single crystals made of Ni-based superalloys offers major cost savings for gas turbine engines with the inclusion of internal cooling channels. However, the lack of understanding of the effect of transient thermal conditions on solidification grain structure during additive manufacturing hinders the potential for process control to maintain the single crystal quality. The use of high-fidelity simulations through high performance computing to predict the evolution of the solidification microstructure will enhance the abilities to tailor the microstructures through process optimization. Phase field simulations are used to determine the effect local thermal conditions and defects on the stability of the solidification morphology, specifically with respect to the onset of columnar-to-equiaxed transition that results in the loss of the single crystal. The results are expected to be instrumental for developing future surrogate models to speed up the integration of design and manufacturing of turbine blades under the harsh in-service conditions.

36 MATERIALS SCIENCE↗

S4PST: Stewardship and Advancement for Programming Systems and Tools 2024-2025 Project Report

We present the "Stewardship and Advancement of Programming Systems and Tools" (S4PST) project report for the calendar years 2024 and 2025. S4PST is dedicated to the stewardship and advancement of Programming Systems and Tools (PST) mainly targeting high-performance computing (HPC) for the scientific community. The project is part of the funded software stewardship organizations (SSOs) selected by ASCR as part of the NGSST program, and a member of CASS: the Consortiumfor the Advancement of Scientific Software.

97 MATHEMATICS AND COMPUTING↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗