Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Chlorine Worth Study Nuclear Data

The Chlorine Worth Study (CWS) was a series of experiments that took place at the National Criticality Experiments Research Center (NCERC), operated by Los Alamos National Laboratory (LANL). The focus of the experiments was to develop new integral benchmarks for the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook with high sensitivity to chlorine in the thermal neutron energy region and which match sensitivities of aqueous chloride operations at LANL. This work discusses the experiment and how the experimental and simulated results using different nuclear data libraries compare to each other.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CRCNS22 Learning Rules in the Hippocampus and their Mapping to Neuromorphic Systems (Final Technical Report)

Large scale biologically-realistic computational models are key to investigating the interplay between structure and function in nervous systems, thus paving the way to new clinical methods and neuro-inspired computing solutions. This project focuses on the hippocampus, in particular the CA3-CA1 regions, due to their role in associative learning and memory, pattern separation and completion, and spatial navigation. Investigations into the neuronal organization and learning rule(s) of this circuit can shed light into how declarative memories are formed, stored, recalled and forgotten and inform computational, experimental and clinical neuroscience work. Our project aims at developing a novel data-driven methodology supported by a broad heterogeneous base of neuroscience experimental knowledge and inspired from advances in computer science and engineering. Specifically, this work will benchmark existing and new learning rules within a full-scale spiking neural network simulation of the CA3-CA1 region. The model will be based on an open-source repository, called the Hippocampome, which contains neuronal morphologies, firing patterns, synapse probabilities, and most other required parameters for all known neuron types in the rodent hippocampal formation. The model will be first trained in a supervised fashion for associative memory tasks using backpropagation through time traditionally used in computer science, enhanced with a new technique called the surrogate gradient method. This optimization method will be used to obtain a global loss minimization, but it is not biologically inspired as it assumes the use of data not locally available to the synapses. However, we propose its use as a benchmarking tool, to compare the training performance of local biologically plausible and hardware-mappable learning rules at scale. New rules or combinations will be proposed and tested as needed, based on the obtained results. Progress in this area will also drive the development of novel hardware-mappable algorithms for continual lifelong learning and categorization of new events from few presented examples. This project goes beyond the existing state-of-the-art by looking at large scale realistic neuronal circuits as networks trainable via global optimization methods such as surrogate gradient descent. The objective function of the brain that supports learning is largely unknown, but it is likely that it operates through local learning rules. Studying network trajectories around local minima as proposed in this work represents a useful strategy for understanding whether a network is training by using a specific (set of) learning rule(s). Starting from a completely untrained network is a challenging test since it is difficult to determine how the learning rule affects the trajectory of the network. This interdisciplinary project will help understand what rule governs learning in these regions or if multiple learning rules are involved. The work will develop a robust methodology to measure if the network is converging to the target solution, oscillating around it, or diverging away.

59 BASIC BIOLOGICAL SCIENCES↗

Wilkins: HPC in situ workflows made easy

In situ approaches can accelerate the pace of scientific discoveries by allowing scientists to perform data analysis at simulation time. Current in situ workflow systems, however, face challenges in handling the growing complexity and diverse computational requirements of scientific tasks. In this work, we present Wilkins, an in situ workflow system that is designed for ease-of-use while providing scalable and efficient execution of workflow tasks. Wilkins provides a flexible workflow description interface, employs a high-performance data transport layer based on HDF5, and supports tasks with disparate data rates by providing a flow control mechanism. Wilkins seamlessly couples scientific tasks that already use HDF5, without requiring task code modifications. We demonstrate the above features using both synthetic benchmarks and two science use cases in materials science and cosmology.

HPC↗

A Bayesian desmearing algorithm for Bonse–Hart USANS with anisotropic scattering

Ultra-small-angle neutron scattering (USANS) using Bonse–Hart optics provides micrometer-scale structural insights but suffers from severe slit-geometry smearing. While well-established for isotropic systems, quantitative desmearing of anisotropic data remains a challenge because conventional corrections break down for non-radial scattering. In this work, we address this by developing a resolution-aware Bayesian framework that explicitly incorporates anisotropy via an affine deformation to the scattering pattern, guided by the principle of parsimony. This results in orientation-resolved point-spread functions that enable a self-consistent determination of both the resolution and deformation parameters. Using Gaussian process regression with uncertainty quantification and a probabilistic correction for multiple scattering, we demonstrate the framework’s effectiveness through numerical benchmarks and experimental studies of a stretched polymer melt. Our approach enables the seamless integration of SANS and USANS data, facilitating quantitative structural analysis of deformed materials at nanometer to micrometer scales.

36 MATERIALS SCIENCE↗

Integration and Quantitative Comparison of Up-Scaled Molecular Observation Network Data with Existing Soil Databases

MONet provides novel soil molecular data to the research community for understanding biogeochemical processes and complementing other soil datasets. This study assesses MONet's ability to replicate known soil patterns via a comparative analysis of soil respiration (Rs), pH, and clay content against benchmark datasets. Results show moderate agreement for pH and clay content, highlighting MONet's strengths in capturing soil biogeochemical variation in underrepresented regions like urban areas. Rs data are marked by the appropriate trends relative to other datasets, but direct comparison is impractical due to methodological differences in underlying data. Strategic sampling is recommended to improve MONet's coverage and eventual utility in bridging molecular observations with global datasets.

54 ENVIRONMENTAL SCIENCES↗

AutoTandemML: Active Learning Enhanced Tandem Neural Networks for Inverse Design Problems

Inverse design in science and engineering involves determining optimal design parameters that achieve desired performance outcomes, a process often hindered by the complexity and high dimensionality of design spaces, leading to significant computational costs. To tackle this challenge, we propose a novel hybrid approach that combines active learning with Tandem Neural Networks to enhance the efficiency and effectiveness of solving inverse design problems. Active learning allows to selectively sample the most informative data points, reducing the required dataset size without compromising accuracy. We investigate this approach using three benchmark problems: airfoil inverse design, photonic surface inverse design, and scalar boundary condition reconstruction in diffusion partial differential equations. We demonstrate that integrating active learning with Tandem Neural Networks outperforms standard approaches across the benchmark suite, achieving better accuracy with fewer training samples.

97 MATHEMATICS AND COMPUTING↗

Systematic Benchmarking of Climate Models: Methodologies, Applications, and New Directions

As climate models become increasingly complex, there is a growing need to comprehensively and systematically assess model performance with respect to observations. Given the increasing number and diversity of climate model simulations in use, the community has moved beyond simple model intercomparison and toward developing methods capable of benchmarking a large number of simulations against a suite of climate metrics. Here, we present a detailed review of evaluation and benchmarking methods and approaches developed in the last decade, focusing primarily on scientific implications for Coupled Model Intercomparison Project (CMIP) simulations and CMIP6 results that contributed to the Intergovernmental Panel on Climate Change (IPCC) Sixth Assessment Report (AR6). Based on this review, we explain the resulting contemporary philosophy of model benchmarking, and provide clear distinctions and definitions of the terms model verification, process validation, evaluation, and benchmarking. While significant progress has been made in model development based on systematic evaluation and benchmarking efforts, some climate system biases still remain. The development of open‐source community software packages has played a fundamental role in identifying areas of significant model improvement and bias reduction. We review the key features of several software packages that have been commonly used over the past decade to evaluate and benchmark global and regional climate models. Additionally, we discuss best practices for the selection of evaluation and benchmarking metrics and for interpreting the obtained results, the importance of selecting suitable sources of reference data and accurate uncertainty quantification.

Environmental sciences↗

Benchmark of the Fe xvv 𝓡 ratio in photoionized plasma during eclipse of Centaurus X-3 with XRISM/Resolve

The $\mathcal {R}$ ratio is a useful diagnostic of the X-ray emitting astrophysical plasmas and is defined as the intensity ratio of the forbidden over the inter-combination lines in the K$\alpha$ line complex of He-like ions. The value is altered by excitation processes (electron impact or UV photoexcitation) from the metastable upper level of the forbidden line, thereby constraining the electron density or UV field intensity. The diagnostic has been applied mostly in electron density constraints in collisionally ionized plasmas using low-Z elements, as was originally proposed for the Sun (Gabriel & Jordan, 1969a, MNRAS, 145, 241), but it can also be used in photoionized plasmas. To make use of this diagnostic, we need to know its value in the limit of no excitation of metastables ($\mathcal {R}_{0}$), which depends on the element, how the plasmas are formed, how the lines are propagated, and the spectral resolution affecting line blending principally with satellite lines from Li-like ions. We benchmark $\mathcal {R}_0$ for photoionized plasmas by comparing calculations using radiative transfer codes and observation data taken with the Resolve X-ray microcalorimeter onboard XRISM. We use the Fe xxv He$\alpha$ line complex of the photo-ionized plasma in Centaurus X-3 observed during eclipse, in which the plasma is expected to be in the limit of no metastable excitation. The measured $\mathcal {R} = 0.65 \pm 0.08$ is consistent with the value calculated using xstar for the plasma parameters derived from other line ratios of the spectrum. We conclude that the $\mathcal {R}$ ratio diagnostic can be used for high-Z elements such as Fe in photoionized plasmas, which has wide applications in plasmas around compact objects at various scales.

X-rays: binaries↗

ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [ORNL] (ORCID:0000000165451943)↗

ORBIT-2 Weather and Climate Downscaling Software Repository

ORBIT-2 is a scalable foundation model for global, hyper-resolution climate and weather downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with 𝑅2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [Oak Ridge National Laboratory]↗

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗

HEU Metal Delayed Critical Experiments with 10 to 19 Inch Thick Graphite Reflectors

Approximately 100 graphite-reflected highly enriched uranium (HEU, 93.14 wt % 235 U) metal annular and cylindrical critical experiments were performed in the early 1960s at the Oak Ridge Critical Experiments Facility (ORCEF). This report presents details from experiment logbooks, experimental data sheets and the author's memory for 44 HEU metal (93.14 wt % 235 U) critical assemblies with graphite reflectors varying from 10 to 19 in. thick, outside diameters varying from 7 to 15 in., inside diameters varying from 7 to 13 in. and critical HEU metal masses varying from 20.4 to 69.0 kg. The data from the 44 experiments described in this report are acceptable for use as criticality safety benchmark experiments for the International Criticality Safety Evaluation Program (ICSBEP) once the uncertainty analysis on the measured k eff is completed. Based on previous ICSBEP benchmarks with this HEU metal at ORCEF, the uncertainties in the measured k eff are expected to be as low as ±0.0004. Preparation of this report is part of an effort at Oak Ridge National Laboratory (ORNL) to document more than 15 undocumented series of critical and subcritical experiments enumerated in Critical and Subcritical NEA Benchmark Possibilities for Measurements at ORCEF and Other US DOE Facilities (Mihalzo, ORNL/TM-2019/1188, 2019) and performed by ORNL at ORCEF and other US Department of Energy critical experiments facilities. More than 500 operational days of critical facility time were used, not including setup and dismantlement time. This documentation for a part of one series of graphite reflected highly enriched uranium metal critical experiments, that used 50 operational days of ORCEF time, was performed using funding received from the DOE Office of Nuclear Energy’s Nuclear Energy University Programs at the University of Tennessee Nuclear Engineering Department. This documentation was also supported by the Nuclear Criticality, Radiation Transport, and Safety programs at ORNL.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Temporal sequence transformer to advance long-term streamflow prediction

Accurate streamflow prediction is crucial for understanding climate change impacts on water resources and for effective management of extreme hydrological events. While Long Short-Term Memory (LSTM) networks have been the dominant data-driven approach for streamflow forecasting, recent advancements in transformer architectures for time series tasks have shown promise in outperforming traditional LSTM models. This study introduces a transformer-based model that integrates historical streamflow data with climatic variables to enhance streamflow prediction accuracy. We evaluated our transformer model against a benchmark LSTM across five diverse basins in the United States. Results demonstrate that the transformer architecture consistently outperforms the LSTM model across all evaluation metrics, highlighting its potential as a more effective tool for hydrological forecasting. This research contributes to the ongoing development of advanced AI techniques for improved water resource management and climate change adaptation strategies.

Singh, Ruhaan [Farragut High School]↗

Monitoring of Liquid Metal Reactor Heater Zones with Recurrent Neural Network Learning of Temperature Time Series

Advanced high-temperature fluid reactors (ARs), such as sodium fast reactors (SFRs) and molten salt cooled reactors (MSCRs) utilize high-temperature fluids at ambient pressure. To melt the fluid during reactor startup and prevent fluid freezing during cooldown, the thermal–hydraulic systems of such ARs include heater zones consisting of specific heaters with controllers, temperature sensors, and thermal insulation. The failure of heater zones due to insulation material degradation or improper installation, resulting in parasitic heat losses, can lead to fluid freezing. The detection of faults using a heat-transfer model is difficult because of a lack of knowledge of the experimental details. Data-driven machine learning of heater zone temperature time series offers a viable alternative. In this study, we benchmarked the performance of recurrent neural networks (RNNs) in an analysis of heat-up transient temperature time series of heater zones installed on a liquid sodium vessel. The RNN models include long short-term memory (LSTM) and gated recurrent unit (GRU) networks, as well as their bi-directional variants, BiLSTM and BiGRU. Anomalous temperature points were designated using a percentile-based threshold applied to residual fluctuations in the detrended temperature time series. Additionally, the impact of the exponentially weighted moving average (EWMA) method on detection accuracy was examined. The RNN models’ performance was assessed using precision, recall, and F 1 score metrics. Results demonstrated that RNN models effectively detect anomalies in temperature time series with the best models for each heater zone achieving F 1 scores of over 93%. To explain the variations in RNN model performance across different heater zones, we used Kullback–Leibler (KL) divergence to quantify the relative entropy between training and testing data, and the Detrended Fluctuation Analysis (DFA) to assess long-range temporal correlations. For datasets with strong long-range correlations and minimal relative entropy between training and testing data, GRU is the best-performing model. When the data exhibits weaker long-term correlations and a significant relative entropy between training and testing distributions, BiGRU shows the best performance. For the data sets with intermediate values of both KL divergence and DFA, the best performance is obtained with LSTM and BiLSTM, respectively.

gated recurrent unit↗

Detecting Process Equipment Failures Using Acoustic Data and Machine Learning

Nuclear power plant (NPP) process equipment such as fans, motors, valves, and pumps generate frequent or continuous noise, and deviations from the normal operational sounds made by this equipment can indicate potential issues. These deviations can be identified via automated acoustic anomaly detection, which involves using acoustic sensors (i.e., microphones) alongside detection algorithms to continuously monitor for changes in acoustic signatures. This task is made challenging by the substantial background noise that exists, such as operators opening and closing doors, manipulating valves, and conversing—in addition to typical plant noises. In collaboration with a nuclear power utility partner, this effort assessed the efficacy of acoustic anomaly detection when using a specific acoustic sensor that compresses data into a fixed set of features that are transferable over a standard Internet of Things communication protocol, thereby improving usability but potentially degrading detection performance. Two methods of performing automated acoustic anomaly detection were evaluated: one-class support vector machine (OC-SVM) and isolation forest (iForest). To enable the use of high-quality acoustic data encompassing both normal and anomalous conditions, the study utilized the publicly available Malfunctioning Industrial Machine Investigation and Inspection dataset, which includes real measured acoustic sensor data for a range of equipment types, model numbers, and signal-to-noise ratios (SNRs), along with a benchmark set of detection results. Using this dataset, the methods were tested and then compared against the benchmark results. The results indicated that although the specific acoustic sensor did not enable as rich a feature set extraction, the proposed methods with the limited feature set performed just as well. This provides solid justification for both the methods and the use of the proposed acoustic sensor.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

The Verification and Validation of a Magnetic Plasma Fluid Model Utilizing the MOOSE (Multiphysics Object Oriented Simulation Environment) Framework

As the goal of achieving fusion power on the grid comes closer to fruition, fully coupled multiphysics models of fusion devices will be crucial. Currently, there are two main approaches to developing these platforms: (1) loosely coupled, where one couples existing codes and solvers together through input and output parameters and data, and (2) tightly coupled, where one develops the necessary models within a singular, integrated framework. This work focuses on the latter approach for magnetically confined fusion devices by developing a fluid-based plasma-edge model within the Multiphysics Object Oriented Simulation Environment (MOOSE) Framework. This effort is coordinated with other efforts to develop, test, demonstrate, and deploy fusion relevant multiphysics capabilities including electromagnetics, particle-in-cell plasma, tritium transport, and fusion blanket design. This new model is an expansion of the MOOSE-based plasma application, Zapdos, which was originally formulated to model low-temperature, non-magnetized plasma processes. Verification, benchmarking, and validation studies have been conducted. Verification studies involved utilizing the method of manufactured solutions and comparing the convergence slope of a known solution to the theoretical slope. Benchmarking consists of comparisons to existing edge codes, namely BOUT++ and SOLEDGE3X. Validation efforts focused on comparisons against open-source data from the TCV tokamak.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗