Search NASA⌕ Search

SEARCH · Search NASA

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Novel artificial neural network model for instantaneous power losses and operational efficiency mapping of MW-scale vanadium redox flow battery for improved technoeconomic analysis

A novel data-driven, machine-learning-based method for modeling the instantaneous power losses of a distribution-sited 2 MW/8MWh vanadium redox flow battery (VRFB), a grid-scale electrochemical storage technology, is introduced and compared against benchmark empirical modeling approaches, including symmetric and asymmetric models, as well as a recent convex hull modeling approach. The novel loss modeling method introduces several advantages over the benchmark models and over simplistic efficiency estimates, the most significant of which is that the model can accurately reflect the stepwise and non-linear parasitic losses associated with the duty cycles of mechanical auxiliary systems like pump motor drives and blower fans. Residuals of the models are compared; the proposed data driven model features significantly improved accuracy over the benchmark models. The model's coefficient of determination is also improved relative to that of the benchmark models. Furthermore, a novel method for visualization of operational efficiency of the grid-scale storage technology is introduced. To demonstrate the benefits of the novel data-driven method for modeling the VRFB, the benchmark models and the proposed models are embedded into an Open DSS distribution network model to study two applications of the grid-scale electrical storage system: load leveling for grid support and energy arbitrage. This article demonstrates that the accuracy of the instantaneous power loss model significantly impacts the understanding of the state of charge of the VRFB. In turn, the accuracy of the efficiency modeling of the VRFB impacts the understanding of the potential economic value and technical benefits to the distribution network operators. In conclusion, the presented power loss modeling approach is, therefore, highly relevant for utility-stakeholders, battery asset owners, system engineers, system designers, and financial planners interested in evaluating or optimizing the operation of grid-scale VRFBs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Comparison of Results between the Legacy and Refined RELAP5-3D Models of the High Temperature Test Facility in Exercises 1 and 2 of the HTTF Benchmark

Work conducted in FY23 identified that RELAP5-3D was capable of reproducing trends in HTTF data during experiment PG-27 but was incapable of reproducing measured values. The primary cause of this discrepancy between RELAP5-3D results and experimental data was hypothesized to be a distortion in power density that was introduced by the radial nodalization of the model. We further hypothesized that a new model would provide better results when compared to the experiments PG-27 and PG-29. Work this FY developed a new model that is better capable of capturing local heat generation rates and contains a representation of each 1/6 azimuthal sector of the core. We used this model to develop a new set of solutions to Exercises 1 and 2 of Problems 2 and 3 in the benchmark. In this report, we present the first comprehensive comparison of the results between the two models. We see that in Exercise 1A, which is common between problems 2 and 3, the results are similar, though the results from the new model show greater detail than those from the legacy model. In Problem 2 Exercise 1B and Problem 3 Exercise 1B, we see that heat removal is slower in the new model than the legacy model. Problem 3 Exercise 1C shows temperatures that are lower in most places in the new model than the legacy model, but the area with active heat generation has higher block temperatures in the new model than the legacy model. Problem 3 Exercise 1D further shows that long-term heat removal is lower in the new model. Problem 2 Exercise 1C demonstrated that the new model observes higher temperatures in the core regions than the legacy model, justifying the need to preserve the power density in HTTF. The validation of PG-27 and PG-29 also demonstrated the improved temperature agreement in the core regions, particularly with a calibrated model that implements an effective thermal conductivity for the core material. Overall, PG-27 models show reasonable to excellent agreement for steady-state temperatures and minimal to reasonable agreement for transients. PG-29 models showed minimal to insufficient agreement with the data.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

SoK: What does it Mean to Benchmark Database Forensics?

Relational Database Management Systems are the backbone of modern enterprises and public-sector services, and are thus frequent targets of security incidents, insider threats, and thorough regulatory audits. Consequently, databases have become key sources of digital evidence, requiring investigators to reconstruct past activity from audit logs, transaction logs, and backups. Although benchmarking frameworks such as those developed by the Transaction Processing Performance Council (TPC) are widely used to evaluate database performance, they do not capture forensic requirements such as evidentiary completeness, tamper-evidence, chain of custody, or regulatory compliance under GDPR and CCPA. This survey examines the emerging domain of forensic database benchmarking. We gathered prior research on database forensics, secure logging, and tamper-evident data structures; we analyze modern forensic-ready features in commercial and open-source systems (SQL Server Ledger, Oracle Blockchain Tables, PostgreSQL pgAudit, Db2 Audit, Aurora Database Activity Streams, Oracle Real Application Security and IBM Guardium) and assess why existing benchmarks are insufficient. We propose forensic workloads, metrics, and methodologies that incorporate adversarial stressors, deleted-record recovery, and backup analysis. We also identify open research problems and call for a community-driven forensic benchmark suite. The result is an idea for evaluating not only database performance but also forensic soundness, bridging the gap between system engineering, compliance, and digital investigations.

Lenard, Ben↗

U.S. Average Market Carbon Dioxide Production Baseline Documentation For 45Q Life Cycle Analysis: Version 1.0 (2018-2022)

U.S. Average Benchmark life cycle inventory documentation for market carbon dioxide production for calendar years 2018 - 2022. The documentation includes a technical report describing data sources and statistical methods used and an Excel file showing calculations. The calculations performed are used in the 45Q benchmark json-ld dataset, compatible with the NETL CO2U LCA Guidance Toolkit. Summary impact results will be included in the NETL CO2U LCA Documentation Spreadsheet within the Toolkit. To access the model referenced in the report, please visit https://www.netl.doe.gov/energy-analysis/details?id=21915ba8-d2cf-43ea-b9b2-1d76363a46a7

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Demonstration of Optimal Benchmark Selection Website and Validation of the q c Coverage Metric Using HEU-SOL-THERM-013-003 Experiment

In the work documented in this interim report, the experiment selection toolkit web site was demonstrated and q C coverage metric methodology was validated for IEU-MET-FAST-002-001, MIX-COMP-THERM 004-004, and HEU-SOL-THERM-013-003 experiments. 𝑞 𝐶 is an information-theoretic measure based on mutual information that quantifies the ability of candidate benchmark experiments to reduce the bias and uncertainty of a target criticality safety application. The metric and an accompanying open-source Python toolkit with a web-based interface were tested against a benchmark set of 425 experiments drawn from the International Criticality Safety Benchmark Evaluation Project Handbook. The interface is hosted at https://edim.covdef.com. It accepts sensitivity data files produced by the TSUNAMI-IP module of the SCALE code system and supports both (i) deterministic analysis using the ENDF/B-VII.0 covariance library and (ii) stochastic analysis based on user-supplied keff samples. Demonstrations on representative applications across a range of material composition, spectrum, and form show that q C -guided benchmark selection achieves greater uncertainty reduction with fewer experiments and yields more stable posterior bias and uncertainty estimates than traditional similarity coefficient ( c k )–based selection, while also capturing valuable low-ck experiments that one-to-one metrics overlook.

Abdel-khalik, Hany S. [Indiana Univ.-Purdue Univ. ↗

Validation of ENDF/B-VIII.1 Nuclear Data Files [Slides]

ENDF-6 formatted files were processed into A Compact ENDF (ACE) files using NJOY2016. Several validation tests were performed: (1) LANL Legacy Benchmark Suite (2) “Modern” Benchmark Suite (3) HEU Benchmark Suite (4) LEU Benchmark Suite (5) Mixed (U+Pu) Benchmark Suite (6) Pu Benchmark Suite (7) 233 U Benchmark Suite.

97 MATHEMATICS AND COMPUTING↗

SMR safety through HTTF modeling and benchmark efforts for code validation for gas-cooled reactor applications

Accurate modeling and simulation tools for thermal-hydraulics calculations are a key element needed to design and license new advanced reactors including Small Modular Reactors (SMR) and Microreactors. Uncertainties in modeling and simulation can have significant safety and economic implications. The High Temperature Test Facility (HTTF) at Oregon State University (OSU) is a scaled integral effects experiment designed to investigate transient behavior in high-temperature gas-cooled prismatic-block nuclear reactors. High-quality measurement data is available from the HTTF that is suitable for a thermal-hydraulics code validation benchmark for gas-cooled reactor simulations. Here, this paper summarizes individual HTTF modeling efforts to date for tool validation at Idaho National Laboratory (INL), Argonne National Laboratory (ANL), Oregon State University (OSU) and Canadian Nuclear Laboratories (CNL) using system thermal-hydraulics codes, Computational Fluid Dynamics (CFD) codes and system-CFD code couplings. Also, the paper introduces the ongoing OECD Nuclear Energy Agency (NEA) High Temperature Gas Reactor Thermal-Hydraulics (HTGR T/H) benchmark that allows for better comparisons of results between different international modeling teams. The benchmark provides well defined computational problems that include code-to-code comparisons and comparisons to measured data. These problems provide an avenue for quantifying accuracy and identifying sources of uncertainty in thermal-hydraulics calculations, including in measured thermophysical properties, as part of validation for gas-cooled reactor simulation tools.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

New constraint on the Np 237 ( n , γ ) Np 238 integral cross section using the Godiva-IV critical assembly

Accurate knowledge of the 237 Np(n, γ) 238 Np cross section at fast neutron energies is important for applied nuclear science. The presently available experimental data has large disagreements in the fast neutron region. Perform a model-independent measurement of the 237 Np(n, γ) 238 Np integral cross section using a well characterized fast neutron source and compare the result with previous measurements and current nuclear data evaluations. Provide an integral measurement that can be used as a benchmark for current evaluations. Multiple samples of 237 Np were irradiated in the Godiva-IV critical assembly. Following the irradiation, the samples placed in a γ-ray counting setup and the γ-rays emitted from the decay of 238 Np were measured over a time period of approximately 7 days. Multiple γ-ray decay branches of 238 Np were observed. The observed activity of 238 Np was used to calculate the amount of 238 Np produced during the irradiation via the 237 Np(n, γ) 238 Np reaction and an integral cross section of 342(11) mb was measured for the Godiva-IV neutron spectrum. Further, the 238 Np half-life has been measured with a result of 50.31(5) hours. The 237 Np(n, γ) 238 Np integral cross section measured in this work is in agreement with overlapping 1σ error bands to ENDF/B-VIII.0. However, the measured value is 3σ away from the calculated integral cross section using JENDL-5. This measurement offers a reliable benchmark for future 237 Np(n, γ) 238 Np cross section evaluations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Structural Changes in Metal Chalcogenide Nanoclusters Associated with Single Heteroatom Incorporation

Atomically precise nanoclusters (NCs) are promising building blocks for designing materials and interfaces with unique properties. By incorporating heteroatoms into the core, the electronic and magnetic properties of NCs can be precisely tuned. To accurately predict these properties, density functional theory (DFT) is often employed, making the rigorous benchmarking of DFT results particularly important. In this study, we present a benchmarking approach based on metal chalcogenide NCs as a model system. We synthesized a series of bimetallic, iron-cobalt chalcogenide NCs [Co 6-x Fe x S 8 (PEt 3 ) 6 ] + (x = 0-6) (PEt = triethyl phosphine) and investigated the effect of heteroatoms in the octahedral metal chalcogenide core on their size and electronic properties. Using ion mobility-mass spectrometry (IM-MS), we observed a gradual increase in the collision cross section (CCS) with an increase in the number of Fe atoms in the core. DFT calculations combined with trajectory method CCS simulations successfully reproduced this trend, revealing that the increase in cluster size is primarily due to changes in metal-ligand bond lengths, while the electronic properties of the core remain largely unchanged. Moreover, this method allowed us to exclude certain multiplicity states of the NCs, as their CCS values were significantly different from those predicted for the lowest-energy structures. Here, this study demonstrates that gas-phase IM-MS is a powerful technique for detecting subtle size differences in atomically precise NCs, which are often challenging to observe using conventional NC characterization methods. Accurate CCS measurements are established as a benchmark for comparison with theoretical calculations. The excellent correspondence between experimental data and theoretical predictions establishes a robust foundation for investigating structural changes of transition metal NCs of interest to a broad range of applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring the Frontiers of Energy Efficiency using Power Management at System Scale

In the face of surging power demands for exascale HPC systems, this work tackles the critical challenge of understanding the impact of software-driven power management techniques like Dynamic Voltage and Frequency Scaling (DVFS) and Power Capping. These techniques have been actively developed over the past few decades. By combining insights from GPU benchmarking to understand application power profiles, we present a telemetry data-driven approach for deriving energy savings projections. This approach has been demonstrably applied to the Frontier supercomputer at scale. Our findings based on three months of telemetry data indicate that, for certain resource-constrained jobs, significant energy savings (up to 8.5%) can be achieved without compromising performance. This translates to a substantial cost reduction, equivalent to 1438 MWh of energy saved. The key contribution of this work lies in the methodology for establishing an upper limit for these best-case scenarios and its successful application. This work enables HPC professionals to optimize the power-performance trade-off within constrained power budgets, not only for the exascale era but also beyond.

Karimi, Ahmad Maroof↗

NCSP Outlook and Interest for Collaboration on HST Experiments [Slides]

The majority of the NCSP budget goes to Integral Experiments. The goal is to produce needed integral data for criticality safety needs in DOE, largely resulting in ICSBEP benchmarks. NCSP has a well defined process for allocating funding through proposals and expert review. NCSP is a fairly small program and funding is prioritized for experiments that would address DOE criticality safety needs. The majority of the currently identified DOE criticality safety needs are HEU and Pu systems. NCSP has a formal mechanism to ensure quality and benefit through the phase gates and approvals within the CED process.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

Qualitative and Quantitative Evaluation for Representative Human Reliability Analysis Methods

The Korea Institute of Nuclear Safety (KINS) is the regulatory expert organization established by the Korean government to strengthen the nation’s technical capabilities relating to nuclear safety regulation. KINS oversees the technical aspects of nuclear safety regulation, including safety reviews, inspections, education, and safety research—all conducted based on technical knowledge and accumulated regulatory experience. In 2023, KINS requested that Idaho National Laboratory (INL) validates representative human reliability analysis (HRA) methods used throughout the world, thus affording KINS with a basis for determining an HRA method adequate for its domestic regulatory purposes. The present paper mainly examines INL’s efforts in this regard. The resulting INL study covered four representative HRA methods widely used by nuclear utilities and regulatory institutes. These methods were qualitatively evaluated by applying specific evaluation criteria and determining how well each method reflected critical HRA issues. For this assessment, INL benchmarked the Halden International HRA Empirical Study. Using the Halden empirical data, along with information on human failure events (HFEs), the present study employed the selected HRA methods to estimate human error probabilities (HEPs) for the HFEs. It also performed statistical analyses to compare the HEPs predicted via the HRA methods against those from the Halden empirical data.

99 - GENERAL AND MISCELLANEOUS↗

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference↗

The Surface-Topography Challenge: A Multi-Laboratory Benchmark Study to Advance the Characterization of Topography

Surface performance is critically influenced by topography in virtually all real-world applications. The current standard practice is to describe topography using one of a few industry-standard parameters. The most commonly reported number is Ra, the average absolute deviation of the height from the mean line (at some, not necessarily known or specified, lateral length scale). However, other parameters, particularly those that are scale-dependent, influence surface and interfacial properties; for example the local surface slope is critical for visual appearance, friction, and wear. The present Surface-Topography Challenge was launched to raise awareness for the need of a multi-scale description, but also to assess the reliability of different metrology techniques. In the resulting international collaborative effort, 153 scientists and engineers from 64 research groups and companies across 20 countries characterized statistically equivalent samples from two different surfaces: a “rough” and a “smooth” surface. The results of the 2088 measurements constitute the most comprehensive surface description ever compiled. We find wide disagreement across measurements and techniques when the lateral scale of the measurement is ignored. Consensus is established through scale-dependent parameters while removing data that violates an established resolution criterion and deviates from the majority measurements at each length scale. Our findings suggest best practices for characterizing and specifying topography. The public release of the accumulated data and presented analyses enables global reuse for further scientific investigation and benchmarking.

42 ENGINEERING↗