Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

2026 Site Development Plan

The 2026 edition of the Site Development Plan focuses on 30 years of infrastructure plans as they strengthen the core Science, Technology, and Engineering (ST&E) pillars of Lawrence Livermore National Laboratory (LLNL): High-Energy-Density (HED) Science, Advanced Materials and Manufacturing (AMM), and High Performance Computing (HPC).

Tran, Lucas [Lawrence Livermore National Laborator↗

Evaluation of LLVM Flang for Production HPC Applications and Modern Fortran Features

In 2025, LLVM released its first Flang Fortran compiler version considered ready for widespread evaluation. We know of no published assessment of Flang compiling a workload- derived portfolio of high-performance computing (HPC) applications. We address this gap using workload data from the National Energy Research Scientific Computing Center (NERSC), which supports more than 10,000 scientists on approximately 1,000 projects. The NERSC workload analyses identify many Fortran components in heavily used applications. We selected 10 such packages with available source code. We compiled them with Flang 22.1.3 on NERSC’s Perlmutter system. Six compiled without code modifications, though some required build-system changes. Three compiled after minor source edits, mostly to address Fortran standard violations. One built only without OpenMP enabled. We evaluated seven additional packages selected for their use of, or enablement of, standard Fortran parallel features: multi-image execution and do concurrent. Six such codes compiled with most or all unit tests passing.

Rasmussen, Katherine↗

Characterization of Peripheral Neurophotonic Systems for High-Performance Human-Computer Interfaces (CRADA Final Report)

As part of the Cyclotron Road program, Morphosis Inc. sought to investigate a non-invasive neuromuscular sensing approach for use as an intuitive and secure human–computer interface. These highly miniaturized, wearable neural interfaces were completely non-invasive and maintained stable, high-bandwidth, long-term access to a user’s actions, intent, and identity, while offering an exceptionally high signal-to-noise ratio compared to contemporary neural recording technologies. The widespread adoption of neural interfaces had the potential to reshape how people interact with technology, with profound societal impacts. Millions worldwide suffered from movement and/or speech disabilities, and these tools had the potential to democratize access to technology to enhance autonomy and quality of life. More broadly, interfaces capable of accurately conveying intentions and safeguarding identities could serve as a cornerstone for privacy, trust, and personal authenticity in digital environments. The use of thought-driven control of digitally enabled devices and governance of digital identities had the potential to revolutionize relationships with technology, transforming how people learn, communicate, and interact with the world.

42 ENGINEERING↗

Openpronghorn

OpenPronghorn is a simulation tool specifically tailored for modeling thermal-hydraulic phenomena in advanced nuclear reactors. It is built on the Multiphysics Object-Oriented Simulation Environment (MOOSE), an open-source platform that facilitates the development of high-performance scientific computing applications. OpenPronghorn solves the Navier-Stokes equations, which describe the conservation of mass, momentum, and energy in fluid flows, using the finite volume numerical method. The code supports a wide range of fluid flow conditions that are applicable to nuclear reactors, including incompressible and weakly compressible flows, as well as single-phase and multiphase flows. It is capable of modeling diverse flow regimes, including laminar and turbulent flows, using various turbulence models such as the standard k-epsilon models, the v2f model, and the mixing length model. For multiphase flows, OpenPronghorn employs a mixture a Eulerian modeling approach with mixture, drift-flux, and full Eulerian models, and includes open-sourced interfacial transfer correlations for drag, exchange, and heat transfer coming from the scientific literature. OpenPronghorn's modular design allows it to handle multiscale simulations, ranging from detailed Reynolds-Averaged Navier Stokes (RANS) simulations to coarse-mesh and lumped parameter models. This flexibility enables users to perform high-fidelity simulations of specific reactor components as well as system-level analyses of entire reactor circuits. The code can be coupled with other MOOSE-based tools using the MultiApp system, allowing for the transfer of coupling quantities such as mass flow rates, heat fluxes, and boundary conditions between different simulation scales. One of the main features of OpenPronghorn is the it includes built-in validation cases from the open-source scientific literature and supports the implementation of user-defined models and correlations through MOOSE's FunctorMaterial system. OpenPronghorn is designed to be computationally efficient, leveraging the SIMPLE projection method for large-scale problems, and can be run on high-performance computing systems to handle the extensive computational demands of detailed reactor simulations. Overall, OpenPronghorn is a versatile and robust tool that provides critical insights into the thermal-hydraulic behavior of advanced nuclear reactors, supporting the design, safety, and optimization of next-generation nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Brochure for the DOE Office of Science Workshop on Envisioning Frontiers in AI and Computing for Biological Research

In February of 2025 a joint ASCR/BER workshop was held to identify key transformational research directions for understanding biology using artificial intelligence (AI), digital twins and high-performance (HPC) computational methods to facilitate scientific discovery and innovation in support of the Department of Energy mission. AI technologies offer exciting new groundbreaking methods to analyze large volumes of complex biological data, thereby greatly accelerating the ability to understand, predict, and design biological processes for beneficial purposes. In the laboratory, the bridging of AI-enabled automated experimental technologies, HPC and digital twins will provide potent tools for researchers to explore the fundamental nature of biology and harness its inherent metabolic potential for a variety of beneficial purposes. The focus of this workshop was on how high-performance computational methods can impact this objective by exploring digital twins, foundational models, and data-driven approaches with applications to advance automated laboratory experiments, modeling of complex living systems and engineering new functions into plants and microbial systems relevant to DOE mission. Workshop attendees with expertise in plant science, microbiology, mathematics, computer science, and AI assessed the current state of the science, trends, and AI challenges at the interface of plant and microbial systems biology and computational science to identify opportunities for high-impact research. This collaborative effort capitalized on ASCR's advancements in applied mathematics, computer science, and Exascale systems, and BER's expertise in basic genomics-enabled research on DOE relevant plant and microbial systems. The workshop culminated in four key priority research directions to guide future research and development within DOE Office of Science programs.

59 BASIC BIOLOGICAL SCIENCES↗

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING↗

New Deformation Mechanisms in Nanocrystalline Nano-porous Small Scale Metals as Defined by Kinetically-driven Microstructures

This report describes recent advances in the design and fabrication of nanostructured metallic pillars with hierarchical microstructures using a novel nanoscale additive manufacturing approach. Through a hydrogel-infusion-based two-photon lithography (TPL) method, we successfully fabricated 3D nickel nanopillars exhibiting both nanocrystalline and nanoporous features. The resulting “bamboo-like” internal architecture comprises 30–50 nm grains and voids with similarly scaled pores. These geometrically tunable pillars, with diameters ranging from ~130 to 550 nm, serve as an ideal platform for probing deformation mechanisms in structurally heterogeneous metals at the nanoscale.

36 MATERIALS SCIENCE↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Alternative mixed integer linear programming optimization for joint job scheduling and data allocation in grid computing

This paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as a mixed integer quadratically constrained program. To tackle the nonlinearity in the constraint, we alternatively fix a subset of decision variables and optimize the remaining ones via Mixed Integer Linear Programming (MILP). We solve the MILP problem at each iteration via an off-the-shelf MILP solver. Our experimental results show that our method significantly outperforms existing heuristic methods, employing either independent optimization or joint optimization strategies. We have also verified the generalization ability of our method over grid environments with various sizes and its high robustness to the algorithm setting.

97 MATHEMATICS AND COMPUTING↗

Hydrodynamic modeling of plasma channel systems for laser plasma accelerators

Structured plasma channels are an essential technology for driving high-gradient, plasma-based acceleration and control of electron and positron beams for advanced concepts accelerators. Laser and gas technologies can permit the generation of long plasma columns known as hydrodynamic, optically-field-ionized (HOFI) channels, which feature low on-axis densities and steep walls. By carefully selecting the background gas and laser properties, one can generate narrow, tunable plasma channels for guiding high intensity laser pulses. Here, we present on the development of simulations of HOFI channels using the FLASH code, a publicly available radiation hydrodynamics code. We explore sensitivities of the channel evolution to laser profile, intensity, and background gas conditions, and identify relevant scalings with laser intensity through a range of practical channel delays.

Accelerators↗

Machine learning-driven predictive resource management in complex science workflows

Here, the collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.

97 MATHEMATICS AND COMPUTING↗

Metal additive manufacturing simulation across length, time, and computing scales

Metal additive manufacturing (AM) offers a unique opportunity for production of advanced materials and complex geometries. However, variability in microstructure and properties challenges conventional approaches to design, process optimization, qualification, and materials selection. Modeling and simulation can improve understanding of AM processing and materials, but also poses major challenges for existing computational methods. Simultaneously, modern scientific computing hardware has become increasingly complex, most notably with the adoption of hybrid architectures such as Graphical Processing Units (GPUs). If appropriately utilized, emerging computational capabilities provide an opportunity to reveal new insight into AM processing and the resulting material structure and properties. In this review we describe the computational AM landscape, identify critical gaps, and highlight opportunities to impact the development and application of AM. First, the requirements and challenges of representative AM problem statements will be defined. Here, these problems range from scientific studies to industrial applications and are designed to capture the breadth of challenges facing the AM community. Next, the current state of AM modeling and simulation is evaluated, broken down by enabling hardware and software, process simulation, microstructure simulation, and property simulation. Each section describes the diversity of simulation approaches and associated trade-offs in physical fidelity and computational expense. Each area is then assessed based on their suitability and readiness for current and developing computational architectures. Lastly, the greatest opportunities for future research and application are highlighted, including gaps in modeling capabilities, opportunities for near-term application, and key scientific challenges.

additive manufacturing↗

ComPort: Rigorous Testing Methods to Safeguard Software Porting (Final Technical Report)

This is a technical report from the lead institution – University of Utah, Kahlert School of Computing – funded under the Department of Energy, Office of Science, Office of Advanced Scientific Computing Research under award number DE-SC0022252. We summarize our work done over the three years of funding received. The relevant papers and software have already been uploaded at the DOE site.

97 MATHEMATICS AND COMPUTING↗

Geometrical optics without singularities: using the ray time as the coordinate space

Geometrical optics (GO) is widely used for reduced modelling of waves in plasmas, but it fails near reflection points, where it predicts a spurious singularity of the wave amplitude. We show how to avoid this singularity by adopting a different representation of the wave equation. Instead of the physical coordinate 𝑥 and the wavevector 𝑘, we use the ray time 𝜏 as the new canonical coordinate and the ray energy ℎ as the associated canonical momentum. To derive the envelope equation in the 𝜏-representation, we construct the Weyl symbol calculus on the (𝜏,ℎ) space and show that the corresponding Weyl symbols are related to their (𝑥,𝑘) counterparts by the Airy transform. This allows us to express the coefficients in the envelope equation through the known properties of the original dispersion operator. When necessary, solutions of this equation can be mapped to the 𝑥-space using a generalised metaplectic transform. However, the field per se might not even be needed in practice. Instead, knowing the corresponding Wigner function usually suffices for linear and quasilinear calculations. As a Weyl symbol itself, the Wigner function can be mapped analytically, using the aforementioned Airy transform. We show that the standard Airy patterns that form in regions where conventional GO fails are successfully reproduced within metaplectic GO (MGO) simply by remapping the field from the 𝜏-space to the 𝑥-space. An extension to mode-converting waves is also presented. This formulation, which we call generalised MGO, can be particularly useful, for example, for reduced modelling of the O–X conversion in inhomogeneous plasma near the critical density, an effect that is important for fusion applications and also occurs in the ionosphere. Overall, MGO can replace GO for any practical purposes, because it better handles cutoffs and is similar otherwise.

plasma waves↗

Bringing HPE Slingshot 11 support to Open MPI

The Cray HPE Slingshot 11 network is used on the new exascale systems arriving at the U.S. Department of Energy (DoE) laboratories (e.g., Frontier, Aurora, Perlmutter). As such, the support of this network is an important capability to meet the needs of exascale applications. Here, this article highlights recent work to develop supporting infrastructure to enable Open MPI to efficiently support these new platforms. A key component of this effort involves development of a new Open Fabrics Interface (OFI) provider, LinkX. We discuss the design and development of enhancements that take advantage of the new Slingshot 11 network and AMD GPUs. We include performance data from tests on the Frontier supercomputer using synthetic communication benchmarks, and the vendor provided MPI as a baseline for comparison. The tests demonstrate full functionality of Open MPI on the system and initial results show favorable performance when compared to the highly tuned vendor implementation.

97 MATHEMATICS AND COMPUTING↗

The dynamical state of eROSITA clusters and its impact on the brightest cluster galaxy luminosity

The first Spectrum-Roentgen-Gamma (SRG) eROSITA public release contains 12 247 clusters and groups from its first 6 months of operation. We used the offset between the brightest cluster galaxy (BCG) and the X-ray peak ( D BCG − X ) to classify the cluster dynamical state of 3946 galaxy clusters and groups. We aim to investigate the evolution of the merger and relaxed cluster distributions with redshift and mass, and the distributions’ impact on the BCG. We used the X-ray peak from the eROSITA survey and the BCG position from the LS DR10 optical data, which includes the DECam eROSITA Survey optical data, to measure the D BCG − X offset. We modelled the distribution of D BCG − X , in units of R 500 , as the sum of two Rayleigh distributions representing the cluster’s relaxed and disturbed populations, and explored their evolution with redshift and mass. To explore the impact of the cluster’s dynamical state on the BCG luminosity, we separated the main sample according to the dynamical state. We defined clusters as relaxed if D BCG − X < 0.25R500, disturbed if D BCG − X > 0.5R500, and ‘diverse’ if 0.25R500 < D BCG − X < 0.5R500. We find no evolution of the merging fraction with redshift or mass. The width of the relaxed distribution increases with redshift, while the width of the two Rayleigh distributions decreases with mass. The analysis reveals that BCGs in relaxed clusters are brighter than BCGs in both the disturbed and diverse cluster populations. The most significant differences are found for high-mass clusters at higher redshifts. The results suggest that BCGs in low-mass clusters are less centrally bound than those in high-mass systems, irrespective of the dynamical state. Over time, BCGs in relaxed clusters progressively align with the potential centre. This alignment correlates with their luminosity growth relative to BCGs in dynamically disturbed clusters, underscoring the critical role of the cluster’s dynamical state in regulating BCG evolution.

galaxy clusters↗

Small tensor product distributed active space (STP-DAS) framework for relativistic and non-relativistic multiconfiguration calculations: Scaling from 10 9 on a laptop to 10 12 determinants on a supercomputer

Despite the power and flexibility of configuration interaction (CI) based methods in computational chemistry, their broader application is limited by an exponential increase in both computational and storage requirements, particularly due to the substantial memory needed for excitation lists that are crucial for scalable parallel computing. Here, the objective of this work is to develop a new CI framework, namely, the small tensor product distributed active space (STP-DAS) framework, aimed at drastically reducing memory demands for extensive CI calculations on individual workstations or laptops, while simultaneously enhancing scalability for extensive parallel computing. Moreover, the STP-DAS framework can support various CI-based techniques, such as complete active space (CAS), restricted active space, generalized active space, multireference CI, and multireference perturbation theory, applicable to both relativistic (two- and four-component) and non-relativistic theories, thus extending the utility of CI methods in computational research. We conducted benchmark studies on a supercomputer to evaluate the storage needs, parallel scalability, and communication downtime using a realistic exact-two-component CASCI (X2C-CASCI) approach, covering a range of determinants from 10 9 to 10 12 . Additionally, we performed large X2C-CASCI calculations on a single laptop and examined how the STP-DAS partitioning affects performance.

Complete-active space self-consistent field↗

Symbolic construction of the chemical Jacobian of quasi-steady state (QSS) chemistries for Exascale computing platforms

The Quasi-Steady State Approximation (QSSA) can be an effective tool for reducing the size and stiffness of chemical mechanisms for implementation in computational reacting flow solvers. However, for many applications, the resulting model still requires implicit methods for efficient time integration. Here, in this paper, we outline an approach to formulating the QSSA reduction that is coupled with a strategy to generate C++ source code to evaluate the net species production rates, and the chemical Jacobian. The code-generation component employs a symbolic approach enabling a simple and effective strategy to analytically compute the chemical Jacobian. For computational tractability, the symbolic approach needs to be paired with common subexpression elimination which can negatively affect memory usage. Several solutions are outlined and successfully tested on a 3D multipulse ignition problem, thus allowing portable application across chemical model sizes and GPU capabilities. The implementation of the proposed method is available at https://github.com/AMReX-Combustion/PelePhysics under an open-source license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗