Search NASA⌕ Search

SEARCH · Search NASA

Results for “High performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

Small arms suppression project (LLNL final report)

US Special Operations Command (USSOCOM) was seeking a technological leap in small firearms weapon suppressor technology, because anticipated enemy capabilities are requiring the operators to have smaller detection cross sections to ensure the safe execution of missions. Suppressors have been developed almost exclusively through trial-and-error methods since the time of the original design by Hiram Maxim over one hundred years ago. Consequently USSOCOM deemed it prudent to perform a physicsbased study of weapon suppression to understand performance limits and possibly identify breakthrough technologies. Lawrence Livermore National Laboratory’s (LLNL’s) high performance production level computational tool called ALE3D (Arbitrary Lagrangian-Eulerian 3D and 2D) has unique physics models and numerical algorithms for modeling suppressor dynamics. The flexible and extendable code framework supports fully integrated hydrodynamics, heat transfer, solid and fluid dynamics, and chemistry that can be applied to simulating propellant-driven motion of a bullet down a gun barrel, the transfer of heat from the burning propellant to the barrel and suppressor, the chemistry of muzzle flash, and the shock/acoustic/optical signatures in the near-field. LLNL’s originally anticipated role was to augment ALE3D for this task, by developing the software and analysis methodologies specific to the simulation of blast and muzzle flash phenomena. It was believed that insights provided by our ALE3D simulations in tandem with a coordinated experimental component by our other team members from Oak Ridge National Laboratory (ORNL) and the U. S. Army Armament, Research, Development and Engineering Center (ARDEC), would have excellent prospects of yielding useful suppressor design improvements that could be transitioned to industry and utilized by US Special Operations Command. The three year effort has come to fruition with the development of revolutionary suppressor designs that far outperform any previous or current design by anyone outside this multi-lab team.

42 ENGINEERING↗

CO2 Capture Using Amines Bound to Silica

One of the main culprits of global warming is the increased amount of carbon dioxide, or CO2, in the atmosphere. NASA's global climate reports a 13% increase in atmospheric CO2 from 2000 to present. Adsorptive CO2 capture by nitrogen groups of amine-containing solvents is one of the most mature technologies deployed in petrochemical and natural gas processing plants to purify industrial gases. However, key challenges in widespread application include solvent induced reactor corrosion, amine degradation, and high regeneration energy for repeated cycling. An alternative approach is to immobilize amines on solid supports. Key performance metrics of solid amine-based CO2 adsorbents include the CO2 adsorption capacity and stability to degradation over hundreds of thousands of regeneration cycles. Our research aims to develop descriptors for CO2 capture capacity and stability against oxygen-induced degradation for amines bound to porous silica supports using experimental and computational techniques. We experimentally measure the change in CO2-uptake using solid amine adsorbents with varying chemical compositions and exposure to varying gas streams and use high-performance computers to simulate the nature and strength of CO2-adsorption and oxidative degradation reaction mechanisms. Insights from our work can facilitate the development of stable solid amine adsorbents for large-scale CO2 capture processes.

amines↗

Large-scale simulation-based parametric analysis of an optimal precooling strategy for demand flexibility in a commercial office building

Achieving success with grid-interactive efficient buildings (GEBs) is closely tied to the utilization of flexible loads. A valuable strategy involves the implementation of precooling techniques before high-demand events, such as peak hours, by adjusting zone air temperature setpoints. This leads to a reduction in thermal loads and peak electricity demand during these times, as the building’s thermal mass stores and subsequently releases thermal energy. However, the effectiveness of the pre-cooling optimization is highly contingent on specific conditions such as building thermal properties, weather conditions, utility rate structure, HVAC equipment sizing, etc. Therefore, investigating the impacts of these condition-specific factors is crucial, especially when considering precooling strategies that utilize thermal mass in commercial buildings. In this paper, we first devised a novel heuristic control approach that incorporates parameterized optimal precooling thermostat schedules to enhance demand flexibility in a commercial office building. Subsequently, we conducted a thorough performance evaluation of this control strategy. Here, the optimal thermostat schedule was parameterized using three optimization variables: the precooling start time, the precooling end time, and the precooling temperature setpoint. Utilizing the DOE medium-sized office building as the virtual testbed, we showed that the parameterized schedule effectively approximates model predictive control and requires drastically reduced computational overhead. In addition, we investigated the impact of different influencing factors on the optimal precooling strategy. These factors include building thermal mass, outdoor air conditions, and energy price profiles. Using high-performance computing, we simulated a total of 225 scenarios, consisting of three levels of thermal mass, five typical outdoor air temperature profiles, and fifteen time-of-use price plans. The results demonstrate that optimal thermostat scheduling could save substantial energy cost in medium-sized office buildings with heavy thermal mass but with some energy penalty. Although the potential for cost savings is lower in buildings with low and medium thermal mass, the energy penalty remains consistent in all three thermal mass scenarios. The study also highlights the need to account for zone diversity and recognize that a one-size-fits-all-zone setpoint schedule may not be suitable for all zones and can lead to unnecessary energy wastage. Furthermore, the results highlight that while outdoor air conditions play a role in cost and energy performance, the cooling load exerts a more immediate and substantial influence on cost savings in precooling strategies. Although cost savings are comparable under certain conditions with the same cooling load, observed deviations in energy penalty indicate potential disparities in the efficiency of the HVAC system during the load-shifting process. In addition, the duration of peak pricing and the ratio between peak and off-peak times exhibit clear correlations with cost savings and energy consumption, aligning with intuitive expectations. These findings offer valuable insights for optimizing precooling strategies in office buildings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Regional-scale fault-to-structure earthquake simulations with the EQSIM framework: Workflow maturation and computational performance on GPU-accelerated exascale platforms

Continuous advancements in scientific and engineering understanding of earthquake phenomena, combined with the associated development of representative physics-based models, is providing a foundation for high-performance, fault-to-structure earthquake simulations. However, regional-scale applications of high-performance models have been challenged by the computational requirements at the resolutions required for engineering risk assessments. The EarthQuake SIMulation (EQSIM) framework, a software application development under the US Department of Energy (DOE) Exascale Computing Project, is focused on overcoming the existing computational barriers and enabling routine regional-scale simulations at resolutions relevant to a breadth of engineered systems. This multidisciplinary software development—drawing upon expertise in geophysics, engineering, applied math and computer science—is preparing the advanced computational workflow necessary to fully exploit the DOE’s exaflop computer platforms coming online in the 2023 to 2024 timeframe. Achievement of the computational performance required for high-resolution regional models containing upward of hundreds of billions to trillions of model grid points requires numerical efficiency in every phase of a regional simulation. This includes run time start-up and regional model generation, effective distribution of the computational workload across thousands of computer nodes, efficient coupling of regional geophysics and local engineering models, and application-tailored highly efficient transfer, storage, and interrogation of very large volumes of simulation data. This article summarizes the most recent advancements and refinements incorporated in the workflow design for the EQSIM integrated fault-to-structure framework, which are based on extensive numerical testing across multiple graphics processing unit (GPU)-accelerated platforms, and demonstrates the computational performance achieved on the world’s first exaflop computer platform through representative regional-scale earthquake simulations for the San Francisco Bay Area in California, USA.

58 GEOSCIENCES↗

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

COLUMBUS─An Efficient and General Program Package for Ground and Excited State Computations Including Spin–Orbit Couplings and Dynamics

The COLUMBUS program system provides the tools for performing high-level multireference (MR) computations, including the multireference configuration interaction (MRCI) method and its multireference averaged quadratic coupled cluster (MR-AQCC) extension, allowing computations on a wide range of fascinating atomic and molecular systems, including the treatment of open-shells and complicated excited state phenomena. The inclusion of spin−orbit coupling (SOC) directly within the MRCI step enables the description of systems containing heavy elements, such as lanthanides and actinides, whose properties are strongly influenced by SOC. Analytic energy gradients and nonadiabatic couplings at the correlated MRCI level provide the foundation for a variety of dynamics studies, giving insight into ultrafast photochemistry. New and ongoing method developments in COLUMBUS include the computation of spin densities, improved descriptions of ionic states, enhancements to the AQCC method, and the porting of COLUMBUS to graphical processing units (GPUs). New external interfaces enable an enhanced description of electronic resonances and molecules in strong laser fields. This work highlights these new developments while providing a detailed account of the diverse applications of COLUMBUS in recent years.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A performant energy-conserving particle reweighting method for Particle-in-Cell simulations

A new particle-based reweighting method is developed and demonstrated in the Aleph Particle-in-Cell with Direct Simulation Monte Carlo (PIC-DSMC) program. Novel splitting and merging algorithms ensure that modified particles maintain physically consistent positions and velocities. This method allows a single reweighting simulation to efficiently model plasma evolution over orders of magnitude variation in density, while accurately preserving energy distribution functions (EDFs). Demonstrations on electrostatic sheath and collisional rate dynamics show that reweighting simulations achieve accuracy comparable to fixed weight simulations with substantial computational time savings. This highly performant reweighting method is recommended for modeling plasma applications that require accurate resolution of EDFs or exhibit significant density variations in time or space.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Dayflow-PR: High-Resolution Streamflow Reanalysis for Puerto Rico, Version 1.0

This dataset presents a high-resolution historical streamflow reanalysis for NHDPlusV2 stream reaches across Puerto Rico (PR) spanning 1950 - 2019. The reanalysis is generated using the calibrated VIC-RAPID hydrologic modeling framework at the Hydrologic Unit Code Sub-basin (HUC08) scale, forced with sub-daily and daily meteorological forcings from Daymet. Runoff is simulated on 1- and 6-km grids, and the resulting total runoff is routed through the NHDPlusV2 river network using the RAPID routing model to produce Naturalized Streamflow Reanalysis. Where complete observational records are available over 1980 - 2019, streamflows are assimilated (substituted) and subsequently routed downstream through the river network to produce Assimilated Streamflow Reanalysis. The dataset includes streamflow outputs from eight distinct hydrologic modeling configurations along with key performanc evaluation metrics at daily and monthly scales, supporting a wide range of water resource applications. This dataset is derived to support the Non-Powered Dam Assessment, as well as 9505 Secure Water Assessment projects for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). For further details, refer to Ghimire et al. (2023), Kao et al. (2024), and Ghimire et al. (2025).

13 HYDRO ENERGY↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

GENTANGLE: integrated computational design of gene entanglements

The design of two overlapping genes in a microbial genome is an emerging technique for adding more reliable control mechanisms in engineered organisms for increased stability. The design of functional overlapping gene pairs is a challenging procedure, and computational design tools are used to improve the efficiency to deploy successful designs in genetically engineered systems. GENTANGLE (Gene Tuples ArraNGed in overLapping Elements) is a high-performance containerized pipeline for the computational design of two overlapping genes translated in different reading frames of the genome. This new software package can be used to design and test gene entanglements for microbial engineering projects using arbitrary sets of user-specified gene pairs.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Quantum Computing for Energy Materials: Opportunities and Challenges

Developing high-performance materials is critical for diverse energy applications to increase efficiency, improve sustainability and reduce costs. Classical computational methods have enabled important breakthroughs in energy materials development, but they face scaling and time-complexity limitations, particularly for high-dimensional or strongly correlated material systems. Quantum computing (QC) promises to offer a paradigm shift by exploiting quantum bits with their superposition and entanglement to address challenging problems intractable for classical approaches. This Perspective discusses the opportunities in leveraging QC to advance energy materials research and the challenges QC faces in solving complex and high-dimensional problems. We present cases on how QC, when combined with classical computing methods, can be used for the design and simulation of practical energy materials. We also outline the outlook for error-corrected, fault-tolerant QC capable of achieving predictive accuracy and quantum advantage for complex material systems.

Algorithms↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Effects of renormalon scheme and perturbative scale choices on determinations of the strong coupling from e + e − event shapes

We study the role of renormalon cancellation schemes and perturbative scale choices in extractions of the strong coupling constant α s ( m Z ) and the leading nonperturbative shift parameter Ω 1 from resummed predictions of the e + e − event shape thrust. We calculate the thrust distribution to N L 3 L ′ resummed accuracy in soft-collinear effective theory (SCET) matched to the fixed-order O ( α s 2 ) prediction, and perform a new high-statistics computation of the O ( α s 3 ) matching in , although we do not include the latter in our final α s fits due to some observed systematics that require further investigation. We are primarily interested in testing the phenomenological impact sourced from varying amongst three renormalon cancellation schemes and two sets of perturbative scale profile choices. We then perform a global fit to available data spanning center-of-mass energies between 35–207 GeV in each scenario. Relevant subsets of our results are consistent with prior SCET-based extractions of α s ( m Z ) , but we are also led to a number of novel observations. Notably, we find that the combined effect of altering the renormalon cancellation scheme and profile parameters can lead to few-percent-level impacts on the extracted values in the α s − Ω 1 plane, indicating a potentially important systematic theory uncertainty that should be accounted for. We also observe that fits performed over windows dominated by dijet events are typically of a higher quality than those that extend into the far tails of the distributions, possibly motivating future fits focused more heavily in this region. Finally, we discuss how different estimates of the three-loop soft matching coefficient c S ˜ 3 can also lead to measurable changes in the fitted { α s , Ω 1 } values. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Revealing the evolution of order in materials microstructures using multi-modal computer vision

The development of high-performance materials for microelectronics, energy storage, and extreme environments depends on our ability to describe and direct property-defining microstructural order. Our present understanding is typically derived from laborious manual analysis of imaging and spectroscopy data, which is difficult to scale, challenging to reproduce, and lacks the ability to reveal latent associations needed for mechanistic models. Here, we demonstrate a multi-modal machine learning (ML) approach to describe order from electron microscopy analysis of the complex oxide La 1−x Sr x FeO 3 . We construct a hybrid pipeline based on fully and semi-supervised classification, allowing us to evaluate both the characteristics of each data modality and the value each modality adds to the ensemble. We observe distinct differences in the performance of uni- and multi-modal models, from which we draw general lessons in describing crystal order using computer vision.

36 MATERIALS SCIENCE↗

Three-dimensional modeling of hyphal fusion, branching, and nutrient transport in filamentous fungi

Fungi exhibit behaviors distinct from other microbes. Filamentous fungi grow by extending complex networks of branched filaments collectively referred to as the mycelium. These networks can expand over large distances and traverse low-nutrient areas by translocating nutrients through the filament network. This spatial characteristic makes filamentous fungi crucial for soil ecosystems, supporting stable microbial communities and promoting plant growth. However, simulating these behaviors is complex. The elongated nature of fungal compartments results in different mechanical interactions compared to the commonly modeled spherical bacteria. These detailed hyphal mechanics require specialized consideration and are often excluded from conventional fungal simulation packages. Additionally, the extensive fungal networks in nature demand computationally intensive simulations, necessitating high-performance algorithms. Therefore, realistic fungi simulations require specialized software. Here, we introduce a fungal modeling expansion to the high-performance biological modelling and interface exchange (bmx) software suite. bmx leverages adaptive mesh refinement in AMReX for chemical diffusion and incorporates a full mechanical model for bacterial cells, accelerated by GPUs. By extending bmx to model filamentous particles, we demonstrate the formation of complex filament networks through interactions like hyphal branching and fusion (anastomosis). We show that the networks produced match real-world fungal structures through various metrics. This work supports computational studies of fungal growth dynamics and can be adapted to investigate the growth of other filamentous structures in biology or materials science. The expanded-BMX package is open-sourced and is available online.

Cell mechanics↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗