Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer systems performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Non-Equilibrium Actinide Radiation Chemistry and the Nuclear Fuel Cycle

Actinides are inherently unstable elements that frequently coexist with other radioisotopes, generating intense ionizing radiation fields that drive the formation of non equilibrium oxidation states. These transient species exert a profound mechanistic influence on the radiation response of actinide containing systems due to their unique redox chemistry. Despite their importance, they remain poorly understood, yet such insight is essential for advancing actinide science and accurately predicting radiation driven behavior. Actinide separations—critical for nuclear energy technologies, strategic deterrence, space exploration, and nuclear medicine—depend on precise control of actinide oxidation states to recover targeted elements from complex matrices such as used nuclear fuel. However, during these processes, actinides, their coordination complexes, and the separation media are all exposed to intense, multicomponent (alpha, beta, gamma, etc.) radiation fields that can alter process efficiency, selectivity, and chemical stability. Understanding, controlling, and mitigating radiation induced reactions is therefore key to innovating and optimizing next generation separation technologies. This seminar will provide an overview of the nuclear fuel cycle and non equilibrium actinide radiation chemistry in the context of recovering actinides from used nuclear fuel, with a particular emphasis on direct dissolution–based reprocessing strategies. We will explore time resolved electron pulse radiolysis and alpha and gamma dose accumulation studies, integrated with multiscale computational modeling, to elucidate the molecular level roles of radiation driven, non equilibrium actinide species in process performance and in the radiolytic stability of organic ligands used for actinide recovery. These insights offer new pathways for designing advanced separation methods and next generation solvent systems, with broad implications for the future of the nuclear fuel cycle.

37 - INORGANIC, ORGANIC, PHYSICAL AND ANALYTICAL C↗

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

kilic, Ozgur Ozan [Brookhaven National Laboratory ↗

Outcomes of HPC User Support using a Science Gateway AI Assistant

High Performance Computing (HPC) is a vital resource for nuclear energy research, facilitating advanced simulations and complex modeling of the quantification and qualification of advanced reactor technology. However, a common gap in knowledge exists around utilizing HPC systems, particularly for nuclear energy researchers unfamiliar with specific HPC systems. A researcher may be well-versed in using one HPC system and understanding its associated processes. Yet, they might struggle when faced with a different HPC system and its unique processes. HPC support staff play a crucial role in addressing these challenges by providing educational resources and assisting users. However, they also face the challenge of maintaining these systems and ensuring they run efficiently for all users, a responsibility that can be challenging to scale effectively with the increasing demand and expansion of HPC systems. This paper addresses this knowledge gap with an artificial intelligence (AI) assistant that offers on-demand, site-specific HPC support for researchers. Idaho National Laboratory (INL) has deployed an AI assistant that is intended to supplement expert HPC support staff and assist nuclear energy researchers. This paper reports on a four-and-a-half-month study evaluating the integration of an AI assistant within a science gateway, with the goal of enhancing existing HPC support.

97 MATHEMATICS AND COMPUTING↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Computational Design of Interlayers for Thermally Stable Compositionally Graded Coatings on Nickel Alloys

To extend the service life of Ni-based superalloys, refractory metal coatings are often used. However, direct bonding between metals with dissimilar crystal structure promotes brittle intermetallic phase formation. This work presents a computational thermodynamic framework for high throughput design of functionally graded interlayers to suppress deleterious phases that may form at the interlayer. The Thermo-Calc software package was used to screen candidate metallic interlayer elements based on stability of solid-solution phases. Vanadium was identified as a promising interlayer due to its consistent suppression of intermetallic phases. Temperature-dependent phase diagram mapping between 600 and 1000 °C guided selection of a compositional pathway that significantly reduced intermetallic formation compared to directly joining the Ni-based and Nb refractory alloys. Time–temperature–transformation analysis was performed to assess whether equilibrium-predicted phases are kinetically accessible along regions of the graded path where non-solid-solution phases are not fully suppressed. The methodology was further applied to additional Ni-based alloy and coating systems, illustrating its transferability as an approach for rapid computational design of graded interlayers in dissimilar high-temperature materials.

36 MATERIALS SCIENCE↗

Transient Modeling and Simulation of a Generic Stable Salt Reactor

A SAM system-level model of a generic stable salt reactor has been developed to investigate thermal-hydraulic behavior and safety performance under steady and transient conditions. The model integrates information generated from a reactor physics analysis using PROTEUS and PERSENT, and a computation fluid dynamics (CFD) analysis using STAR-CCM+. A loose, iterative coupling scheme between PROTEUS and SAM is implemented to calculate the equilibrium power and temperature distributions in the steady-state critical core condition. The converged steady-state model is then used in PERSENT to calculate the four reactivity feedback temperature coefficients (Doppler, fuel density, coolant density, and core radial expansion) and kinetic parameters that are needed in SAM to model the temperature feedback effects in transient simulations. Within the fully enclosed liquid fuel pins, natural convection is the dominant heat transfer mechanism. The STAR-CCM+ model of the fuel pin considers conjugate heat transfer from the liquid fuel salt to the pin cladding and external reactor coolant. The CFD results of the axial and radial temperature profiles are used to empirically determine an effective fuel salt thermal conductivity in the SAM fuel pin model so that the temperatures predicted by the SAM model match as closely as possible the CFD results. In the central region of the fuel pin, the effective thermal conductivity is as high as similar to 60 times the physical fuel salt thermal conductivity. The whole-plant SAM model is then used to simulate an unprotected station blackout transient. The results of this simulation showed that the large negative fuel axial expansion reactivity feedback reduces fission power to similar to 2.4% nominal power. The core is cooled by natural circulation, which removes heat in the core to the emergency heat removal system, and ultimately, to the ambient. However, peak fuel salt and cladding temperatures can potentially reach as high as 1500 K, albeit briefly, if the shutdown mechanism fails to operate.

stable salt reactor; transient simulations; system↗

Bringing HPE Slingshot 11 support to Open MPI

The Cray HPE Slingshot 11 network is used on the new exascale systems arriving at the U.S. Department of Energy (DoE) laboratories (e.g., Frontier, Aurora, Perlmutter). As such, the support of this network is an important capability to meet the needs of exascale applications. Here, this article highlights recent work to develop supporting infrastructure to enable Open MPI to efficiently support these new platforms. A key component of this effort involves development of a new Open Fabrics Interface (OFI) provider, LinkX. We discuss the design and development of enhancements that take advantage of the new Slingshot 11 network and AMD GPUs. We include performance data from tests on the Frontier supercomputer using synthetic communication benchmarks, and the vendor provided MPI as a baseline for comparison. The tests demonstrate full functionality of Open MPI on the system and initial results show favorable performance when compared to the highly tuned vendor implementation.

97 MATHEMATICS AND COMPUTING↗

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Preventing Loss of Selectivity during the Oxidative Dehydrogenation of Propane over Supported Vanadium Catalysts

Supported vanadium materials are promising catalysts for the oxidative dehydrogenation of propane to propylene (ODHP), but a lack of mechanistic understanding limits the rational design of catalysts with improved propylene selectivity. Adding Ta to V/SiO 2 increases the propylene selectivity, as well as the activity, leading to superior performance compared to state-of-the-art boron-based systems. In this contribution, we utilize this surprising promotional effect of Ta to elucidate key elements of the mechanistic cycle. Through a combination of characterization techniques, computational modeling, and kinetic experiments, we show that the catalytic cycle over V/SiO 2 likely involves the formation of an isopropyl alcohol intermediate, the fate of which is in kinetic competition between subsequent dehydration to propylene or further oxidation. Furthermore, we show that the relatively facile propylene overoxidation observed for these materials occurs via the epoxidation of propylene by a proposed peroxovanadium intermediate, rather than the abstraction of propylene’s allylic C–H bond as previously assumed. Using these key mechanistic features, we rationalize the enhanced selectivity and activity of Ta promotion. In conclusion, our mechanistic framework offers avenues for future catalyst development to improve supported vanadium materials for ODHP.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Probabilistic Resilience-Oriented Assessment Approach for Transmission Networks Under Wildfires

The rising threat of wildfires poses significant challenges to power transmission networks, particularly in areas prone to such disasters. Traditional approaches for wildfire risk assessment neglect some potential wildfire scenarios. Here, this paper introduces a probabilistic resilience-oriented assessment approach for power transmission networks to address this gap. Initially, a probabilistic wildfire model is developed to capture uncertainties in ignition, intensity, and fire spread. Next, a spatiotemporal fragility model is constructed to assess the impact of wildfires on transmission corridors, incorporating Thermal Aging (TA) and Dynamic Thermal Rate (DTR) change. Finally, a comprehensive resilience metric is defined to evaluate system performance, leveraging the fragility model to determine component and system-level resilience. The approach employs a combinatorial enumeration method to generate potential wildfire scenarios, enhanced by an impact-increment-based state enumeration (IISE) method for computational efficiency. The proposed method provides critical insights for identifying system vulnerabilities and developing robust strategies to protect transmission networks from wildfires. The efficacy of this approach is validated through extensive scenarios of the RTS-GMLC system across Southern California, Nevada and Arizona.

Vahedi, Soroush [Univ. of Connecticut, Storrs, CT ↗

"Forward" Projects Boost U.S. Leadership in Advanced Computing and Artificial Intelligence

High-performance computing (HPC) has been an indispensable research tool for accessing physical realms difficult, or impossible, achieve with experiment alone. For several decades, the Department of Energy’s (DOE’s) Office of Science has deployed sophisticated HPC systems for solving the nation’s most pressing grand challenge problems in energy, climate change, and human health. In addition, DOE’s National Nuclear Security Administration (NNSA) has adeptly applied HPC in support of key national security objectives, such as nuclear science and stockpile modernization and stewardship. Over time, HPC systems have become increasingly more complex and capable, and as each new machine has come online, scientists and engineers have taken advantage of vast increases in compute power to accelerate scientific discoveries and engineering innovation.

42 ENGINEERING↗

Validation of Numerical Tools for Calculating Reactivity Feedback in Sodium Fast Reactors Using SEFOR Experimental Data

The Southwest Experimental Fast Oxide Reactor (SEFOR) was an experimental sodium-cooled fast breeder reactor operated from 1969 to 1972 with experiments designed to measure Doppler reactivity feedback in a wide temperature range from around 350 °F to temperatures approaching the melting point of mixed oxide fuel of around 5000 °F, providing valuable data for code validations. Co-supported by the Department of Energy (DOE) Fast Reactor Program (FRP) and the DOE Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, the SEFOR benchmark project focused on using the experimental data to validate numerical tools that are used in industry and academia to design and license sodium-cooled fast reactors (SFRs). By the end of FY-25, substantial progress was achieved in the SEFOR benchmark study. A variety of numerical tools commonly used for modeling SFRs were applied to develop models for SEFOR core configurations I-D, I-E, I-I, and I-J. These included Monte Carlo codes such as MCNP, Serpent, and Shift; deterministic codes such as the legacy Argonne Reactor Computation (ARC) suite and the high-fidelity NEAMS code Griffin; and the system analysis code SAS4A/SASSYS-1 (SAS). Using these models, both SEFOR zero-power experiments and power-ascending tests were successfully simulated. Comparisons were performed against experimental measurements of core criticalities, reflector worth, kinetics parameters (Λ/βeff), isothermal reactivity feedback (from 350 °F to 760 °F at zero power), and power-ascending reactivity feedback (as power increased from 0.4 MW to 17 MW). In general, these comparisons demonstrated very good agreement between numerical results and experimental data. In Fiscal Year 26 (FY-26), the SEFOR benchmark project will continue to address the modeling issues identified in FY-25. Effort will focus on the simulation of reactivity insertion transients in SEFOR core II using the ARC/SAS model. Future work will also focus on incorporating BISON into the SEFOR core modeling process to enable the first Multiphysics simulations of the isothermal tests based on the MOOSE framework.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Automated Inspection of Criticality Control Overpacks for Surplus Plutonium Disposition: Qualification Update – 25313

In an effort to reduce the amount of nuclear waste in South Carolina, the Department of Energy (DOE) tasked the Savannah River Site (SRS) with diluting and disposing of the amount of plutonium in the state. This process involves the movement and shipment of over 100,000 criticality control overpacks (CCOs) throughout the project lifespan, lending itself to the use of automation to reduce worker radiation exposure and more efficiently utilize human capital. Due to the large scope, this overarching process was broken down into several different “automation projects” to be developed. The first opportunity pursued was the receipt and inspection of empty CCO drums coming into SRS, identified as Automation Project 1 (AP1), and is the focus of this paper. AP1 was developed to unpack incoming CCOs and inspect them for unwanted foreign objects and any damage to the drum or its contents. This process is accomplished by the combination of an automated guided vehicle (AGV) that delivers CCOs to a robotic arm which uses a suite of custom tools to disassemble a CCO, inspect the inside and outside of the CCO and its inner criticality control container (CCC), reassemble the CCC and CCO, and apply a tamper indicating device (TID) to the inspected drum. In past years, the robotic work cell had been developed in a small-scale testing facility for proof-of-concept. This year, major improvements were made to the robotic work cell to perform the process, including integration into the final facility where CCOs will be inspected. Other technical improvements include the implementation of sensor feedback and safety relays into the control system to allow the state of the work cell to be better tracked, and additional development of the TID application process to complete the robotic inspection. Further enhancements were made to the robotic vision processes and robot pathing, as well as development on a computer vision inspection process to detect inspection criteria anomalies in CCOs. In addition to developmental improvements, the work cell underwent a six-month testing period to ensure the project requirements were met. Results of this testing period demonstrate the work cell’s capability to meet project throughput goals at an acceptable level, successfully document the status of each CCO inspected, and reduce the toll on technical operations’ human power by two thirds. At the time of this paper, the work cell is capable of autonomously handling up to eight CCOs with an AGV, delivering CCOs to and from the robot work cell, and having a robotic arm perform a full receipt and inspection procedure on each CCO. Moving forward, repeatability will be improved so that these CCOs can be run back-to-back seamlessly, as well as improving the system to handle more significant edge cases and failure modes.

Spivey, Nicholas↗

Numerical assessment of triply periodic minimal surfaces for direct air capture of carbon dioxide

Direct air capture (DAC) systems often consist of packing material wetted by a capture fluid that reacts with CO 2 in the airstream. The efficiency of the contactor is determined by a complex relationship of fluid dynamics, heat and mass transfer, contactor geometry, and chemical properties. The efficiency of the contactor must be balanced with other factors, primarily pressure drop through the system. Triply periodic minimal surfaces (TPMS) are a class of differential surfaces that have been explored in multiple engineering applications and have been shown to exhibit excellent performance when used in heat exchangers. Their tortuous path provides a high surface-to-volume ratio and favorable trade-off between contact area and pressure drop. In this work, a gyroid-type TPMS contactor was evaluated using computational fluid dynamics for a variety of geometric parameters to explore the potential benefit of TPMS shapes for DAC applications. A thin-film model was employed to model the flow and distribution of the capture solvent, allowing efficient simulations of TPMS structures at scale by eliminating the need for a computationally intensive interface capturing method. A liquid-gas mass transfer model was implemented in the commercial software STAR-CCM+ and used to predict the CO 2 capture efficiency and study the trade-off between capture performance and pressure drop through analysis of capture rates, mass transfer coefficients, and other relevant variables. TPMS contactors with a variety of geometric parameters and two capture solvent options were investigated to determine the effect of design choices on the operational performance of DAC systems. In conclusion, results showed that while contactor geometry is the dominant factor in efficiency and pressure drop, the physiochemical properties of the solvent are an important secondary influence on the contactor performance.

CFD↗

Energy dataset of Frontier supercomputer for waste heat recovery

The Hewlett Packard Enterprise–Cray EX Frontier is the world’s first and fastest exascale supercomputer, hosted at the Oak Ridge Leadership Computing Facility in Tennessee, United States. Frontier is a significant electricity consumer, drawing 8–30 MW; this massive energy demand produces significant waste heat, requiring extensive cooling measures. Although harnessing this waste heat for campus heating is a sustainability goal at Oak Ridge National Laboratory (ORNL), the 30 °C–38 °C waste heat temperature poses compatibility issues with standard HVAC systems. Heat pump systems, prevalent in residential settings and some industries, can efficiently upgrade low-quality heat to usable energy for buildings. Thus, heat pump technology powered by renewable electricity offers an efficient, cost-effective solution for substantial waste heat recovery. However, a major challenge is the absence of benchmark data on high-performance computing (HPC) heat generation and waste heat profiles. This paper reports power demand and waste heat measurements from an ORNL HPC data centre, aiming to guide future research on optimizing waste heat recovery in large-scale data centres, especially those of HPC calibre.

97 MATHEMATICS AND COMPUTING↗