Search NASASearch

SEARCH · Search NASA

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Development of a C-ELS Specimen-Based Numerical Benchmark for Mode II Delamination and Assessment of Two VCCT-Based Propagation Strategies

A finite element (FE) benchmark example inspired by the calibrated end-loaded split (C-ELS) specimen is developed and used to assess the performance of delamination propagation capabilities based on linear elastic fracture mechanics (LEFM). The C-ELS specimen has the advantage of a longer region of stable delamination propagation compared to the existing mode II benchmark case. The new benchmark example may therefore provide a better assessment tool by enabling more stable crack growth in regions further away from the boundary conditions or load application. First, a benchmark result is created manually using two-dimensional finite element models of the C-ELS specimen with different delamination lengths. Second, the performance of the virtual crack closure technique (VCCT) delamination propagation capabilities in the Abaqus/Standard®1 FE code and the recently developed Progressive Release eXplicit-VCCT (PRX-VCCT) method are assessed by comparing the results to the benchmark case. Two examples with different starter delamination lengths are studied. A shorter starter length is chosen to create a scenario with unstable delamination propagation. A longer delamination encourages stable delamination propagation. Detailed results from three-dimensional analyses with aligned and misaligned meshes and two levels of mesh refinement are provided. In general, good agreement can be achieved between the results obtained from the quasi-static propagation analysis and the benchmark analysis. Numerical artifacts including anomalous unreleased nodes in the crack wake and zig-zag crack fronts occur for propagation analyses using Abaqus/Standard VCCT. In comparison, continuous, smooth, delamination fronts are observed for PRX-VCCT. The use of the benchmark case to assess different VCCT-based propagation strategies illustrates the value of establishing benchmark cases.

Composite Materials

Results of a Geant4 benchmarking study for bio‐medical applications, performed with the G4‐Med system

Geant4, a Monte Carlo Simulation Toolkit extensively used in bio-medical physics, is in continuous evolution to include newest research findings to improve its accuracy and to respond to the evolving needs of a very diverse user community. In 2014, the G4-Med benchmarking system was born from the effort of the Geant4 Medical Simulation Benchmarking Group, to benchmark and monitor the evolution of Geant4 for medical physics applications. The G4-Med system was first described in our Medical Physics Special Report published in 2021. Results of the tests were reported for Geant4 10.5. Purpose In this work, we describe the evolution of the G4-Med benchmarking system. Methods The G4-Med benchmarking suite currently includes 23 tests, which benchmark Geant4 from the calculation of basic physical quantities to the simulation of more clinically relevant set-ups. New tests concern the benchmarking of Geant4-DNA physics and chemistry components for regression testing purposes, dosimetry for brachytherapy with a 125 I source, dosimetry for external x-ray and electron FLASH radiotherapy, experimental microdosimetry for proton therapy, and in vivo PET for carbon and oxygen beams. Regression testing has been performed between Geant4 10.5 and 11.1. Finally, a simple Geant4 simulation has been developed and used to compare Geant4 EM physics constructors and physics lists in terms of execution times. Results In summary, our EM tests show that the parameters of the multiple scattering in the Geant4 EM constructor G4EmStandardPhysics_option3 in Geant4 11.1, while improving the modeling of the electron backscattering in high atomic number targets, are not adequate for dosimetry for clinical x-ray and electron beams. Therefore, these parameters have been reverted back to those of Geant4 10.5 in Geant4 11.2.1. The x-ray radiotherapy test shows significant differences in the modeling of the bremsstrahlung process, especially between G4EmPenelopePhysics and the other constructors under study (G4EmLivermorePhysics, G4EmStandardPhysics_option3, and G4EmStandardPhysics_option4). These differences will be studied in an in-depth investigation within our Group. Improvement in Geant4 11.1 has been observed for the modeling of the proton and carbon ion Bragg peak with energies of clinical interest, thanks to the adoption of ICRU90 to calculate the low energy proton stopping powers in water and of the Linhard–Sorensen ion model, available in Geant4 since version 11.0. Nuclear fragmentation tests of interest for carbon ion therapy show differences between Geant4 10.5 and 11.1 in terms of fragment yields. In particular, a higher production of boron fragments is observed with Geant4 11.1, leading to a better agreement with reference data for this fragment. Conclusions Based on the overall results of our tests, we recommend to use G4EmStandardPhysics_option4 as EM constructor and QGSP_BIC_HP with G4EmStandardPhysics_option4, for hadrontherapy applications. The Geant4-DNA physics lists report differences in modeling electron interactions in water, however, the tests have a pure regression testing purpose so no recommendation can be formulated.

62 RADIOLOGY AND NUCLEAR MEDICINE

A New Shutdown Dose Rate Benchmark Problem for Representative Fusion Applications

Here, this work introduces a new benchmark problem for calculating shutdown dose rates (SDDRs) aimed at fusion reactor applications. The model is designed to represent a simplified version of a typical ITER port plug. The responses of interest include neutron flux, gamma flux, and gamma SDDR at 12 different locations scattered throughout the port. This article outlines the geometry specifications of the problem, provides material definitions for the components, specifies the required responses to be calculated, and presents the source definition information. The need for this benchmark arises from the limited availability of publicly accessible references, with only one benchmark representing the typical dimensions and materials found in fusion systems. This existing benchmark has been cited extensively, reflecting the demand within the scientific community to test both established and novel workflows for SDDR calculations. However, since its presentation at a conference in 2011, the results have become increasingly well known. Moreover, the absence of formal publication and peer review has led to the details of this benchmark being extracted from secondary sources, such as subsequent studies that reference it. As a result, analysts are left with significant flexibility in interpreting the key parameters, which can be adjusted to account for unknown systematic errors, ultimately reproducing the already well-known responses. This new benchmark serves as an updated version of that earlier work, with the aim of providing a more reliable description of the materials and their impurities, which is crucial for assessing activation and subsequent gamma emission. Additionally, it seeks to provide a geometry that more closely represents an ITER port plug. The improvements in the problem definition will lead to a more reproducible benchmark problem, while also presenting the radiation transport community with a completely new challenge. The results will be published in a future article to allow analysts adequate time to analyze this problem independently.

Benchmark

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking

Verification of the Uniformly-Ordered Binary Decision Algorithm in Correlated-Benchmark Whisper Calculations

Whisper is a nuclear criticality safety code package that aids analysts in validation exercises by computing upper subcritical limits (USL) for applications of interest. To obtain statistically meaningful, significant, and conservative USLs, the analyst must ensure that Whisper selects a sufficient number of benchmarks that are neutronically similar to the application. Many of the available benchmarks are correlated but are currently treated as independent, leading to an artificially small sample size, as their individual information contributions will be overestimated. To aid the analyst in obtaining a sufficient sample size, prior work [2] demonstrated application of the Uniformly-Ordered Binary Decision (UOBD) algorithm in adjusting benchmark weights to account for benchmark correlations. This work provides verification of the Whisper implementation and considers the impact of updated benchmark correlations compared to those available previously. We demonstrate that the UOBD algorithm performs as expected with an analytic example. With HEU-SOL-THERM-001 cases 1 through 10 as the applications, we compare the USLs computed with benchmark correlations available in the Whisper 1.1 release only to those computed with additional benchmark correlations from DICE 2023 and demonstrate substantive differences.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Systematic Benchmarking of Climate Models: Methodologies, Applications, and New Directions

As climate models become increasingly complex, there is a growing need to comprehensively and systematically assess model performance with respect to observations. Given the increasing number and diversity of climate model simulations in use, the community has moved beyond simple model intercomparison and toward developing methods capable of benchmarking a large number of simulations against a suite of climate metrics. Here, we present a detailed review of evaluation and benchmarking methods and approaches developed in the last decade, focusing primarily on scientific implications for Coupled Model Intercomparison Project (CMIP) simulations and CMIP6 results that contributed to the Intergovernmental Panel on Climate Change (IPCC) Sixth Assessment Report (AR6). Based on this review, we explain the resulting contemporary philosophy of model benchmarking, and provide clear distinctions and definitions of the terms model verification, process validation, evaluation, and benchmarking. While significant progress has been made in model development based on systematic evaluation and benchmarking efforts, some climate system biases still remain. The development of open‐source community software packages has played a fundamental role in identifying areas of significant model improvement and bias reduction. We review the key features of several software packages that have been commonly used over the past decade to evaluate and benchmark global and regional climate models. Additionally, we discuss best practices for the selection of evaluation and benchmarking metrics and for interpreting the obtained results, the importance of selecting suitable sources of reference data and accurate uncertainty quantification.

Environmental sciences

Ecological Benchmark for Chemicals

The Ecological Benchmark Tool dataset serves as a comprehensive repository of benchmarks designed to assess ecological risks at contaminated sites. This tool facilitates the evaluation of various environmental media and contaminants, supporting regulatory compliance and ecological protection. Benchmarks are available for air, biota, sediment, soil, and surface water. The dataset also provides species-specific benchmarks for fish, plants, birds, mammals, and invertebrates. Users can select benchmark sources, media, individual chemicals, and retrieve results in tabular or spreadsheet formats for analysis. Benchmarks are derived from authoritative sources, including government agencies, scientific councils, and academic publications. The dataset supports ecological risk assessments, regulatory decision-making, and environmental planning, with tools for benchmarking against chemical thresholds, sensitive species protection, and habitat impact evaluations. This structured approach ensures a robust evaluation of ecological risks tailored to site-specific and regulatory needs.

Stewart, Debra [Oak Ridge National Laboratory (ORN

Ecological Benchmark for Radionuclides

The Ecological Benchmark Tool for radionuclides dataset serves as a comprehensive repository of benchmarks designed to assess ecological risks at contaminated sites. This tool facilitates the evaluation of various environmental media and contaminants, supporting regulatory compliance and ecological protection. Benchmarks are available for sediment, soil, and surface water. The dataset also provides species-specific benchmarks for fish, plants, birds, mammals, and invertebrates. Users can select benchmark sources, media, individual radionuclides, and retrieve results in tabular or spreadsheet formats for analysis. Benchmarks are derived from authoritative sources, including government agencies, scientific councils, and academic publications. The dataset supports ecological risk assessments, regulatory decision-making, and environmental planning, with tools for benchmarking against radiological thresholds, sensitive species protection, and habitat impact evaluations. This structured approach ensures a robust evaluation of ecological risks tailored to site-specific and regulatory needs.

Stewart, Debra [Oak Ridge National Laboratory (ORN

Development of Benchmark Examples for Static Delamination Propagation and Fatigue Growth Predictions

The development of benchmark examples for static delamination propagation and cyclic delamination onset and growth prediction is presented and demonstrated for a commercial code. The example is based on a finite element model of an End-Notched Flexure (ENF) specimen. The example is independent of the analysis software used and allows the assessment of the automated delamination propagation, onset and growth prediction capabilities in commercial finite element codes based on the virtual crack closure technique (VCCT). First, static benchmark examples were created for the specimen. Second, based on the static results, benchmark examples for cyclic delamination growth were created. Third, the load-displacement relationship from a propagation analysis and the benchmark results were compared, and good agreement could be achieved by selecting the appropriate input parameters. Fourth, starting from an initially straight front, the delamination was allowed to grow under cyclic loading. The number of cycles to delamination onset and the number of cycles during stable delamination growth for each growth increment were obtained from the automated analysis and compared to the benchmark examples. Again, good agreement between the results obtained from the growth analysis and the benchmark results could be achieved by selecting the appropriate input parameters. The benchmarking procedure proved valuable by highlighting the issues associated with the input parameters of the particular implementation. Selecting the appropriate input parameters, however, was not straightforward and often required an iterative procedure. Overall, the results are encouraging but further assessment for mixed-mode delamination is required.

Kruger, Ronald

Development and Application of Benchmark Examples for Mode II Static Delamination Propagation and Fatigue Growth Predictions

The development of benchmark examples for static delamination propagation and cyclic delamination onset and growth prediction is presented and demonstrated for a commercial code. The example is based on a finite element model of an End-Notched Flexure (ENF) specimen. The example is independent of the analysis software used and allows the assessment of the automated delamination propagation, onset and growth prediction capabilities in commercial finite element codes based on the virtual crack closure technique (VCCT). First, static benchmark examples were created for the specimen. Second, based on the static results, benchmark examples for cyclic delamination growth were created. Third, the load-displacement relationship from a propagation analysis and the benchmark results were compared, and good agreement could be achieved by selecting the appropriate input parameters. Fourth, starting from an initially straight front, the delamination was allowed to grow under cyclic loading. The number of cycles to delamination onset and the number of cycles during delamination growth for each growth increment were obtained from the automated analysis and compared to the benchmark examples. Again, good agreement between the results obtained from the growth analysis and the benchmark results could be achieved by selecting the appropriate input parameters. The benchmarking procedure proved valuable by highlighting the issues associated with choosing the input parameters of the particular implementation. Selecting the appropriate input parameters, however, was not straightforward and often required an iterative procedure. Overall the results are encouraging, but further assessment for mixed-mode delamination is required.

Krueger, Ronald

Development of Benchmark Examples for Quasi-Static Delamination Propagation and Fatigue Growth Predictions

The development of benchmark examples for quasi-static delamination propagation and cyclic delamination onset and growth prediction is presented and demonstrated for Abaqus/Standard. The example is based on a finite element model of a Double-Cantilever Beam specimen. The example is independent of the analysis software used and allows the assessment of the automated delamination propagation, onset and growth prediction capabilities in commercial finite element codes based on the virtual crack closure technique (VCCT). First, a quasi-static benchmark example was created for the specimen. Second, based on the static results, benchmark examples for cyclic delamination growth were created. Third, the load-displacement relationship from a propagation analysis and the benchmark results were compared, and good agreement could be achieved by selecting the appropriate input parameters. Fourth, starting from an initially straight front, the delamination was allowed to grow under cyclic loading. The number of cycles to delamination onset and the number of cycles during delamination growth for each growth increment were obtained from the automated analysis and compared to the benchmark examples. Again, good agreement between the results obtained from the growth analysis and the benchmark results could be achieved by selecting the appropriate input parameters. The benchmarking procedure proved valuable by highlighting the issues associated with choosing the input parameters of the particular implementation. Selecting the appropriate input parameters, however, was not straightforward and often required an iterative procedure. Overall the results are encouraging, but further assessment for mixed-mode delamination is required.

Krueger, Ronald

Structural Benchmark Creep Testing for Microcast MarM-247 Advanced Stirling Convertor E2 Heater Head Test Article SN18

This report provides test methodology details and qualitative results for the first structural benchmark creep test of an Advanced Stirling Convertor (ASC) heater head of ASC-E2 design heritage. The test article was recovered from a flight-like Microcast MarM-247 heater head specimen previously used in helium permeability testing. The test article was utilized for benchmark creep test rig preparation, wall thickness and diametral laser scan hardware metrological developments, and induction heater custom coil experiments. In addition, a benchmark creep test was performed, terminated after one week when through-thickness cracks propagated at thermocouple weld locations. Following this, it was used to develop a unique temperature measurement methodology using contact thermocouples, thereby enabling future benchmark testing to be performed without the use of conventional welded thermocouples, proven problematic for the alloy. This report includes an overview of heater head structural benchmark creep testing, the origin of this particular test article, test configuration developments accomplished using the test article, creep predictions for its benchmark creep test, qualitative structural benchmark creep test results, and a short summary.

life (durability)

A Benchmark Example for Delamination Propagation Predictions Based on the Single Leg Bending Specimen Under Quasi-Static and Fatigue Loading

Benchmark examples based on Single Leg Bending (SLB) specimens with equal and unequal bending arm thicknesses were used to assess the performance of delamination prediction capabilities in finite element codes. First, the development of the quasi-static benchmark cases using the Virtual Crack Closure Technique (VCCT) is discussed in detail. Second, based on the quasi-static benchmark results, additional benchmark cases to assess delamination propagation under fatigue loading are created. Third, the application is demonstrated for the commercial finite element code Abaqus Standard 2018. The benchmark cases are compared to results obtained from VCCT-based, automated quasi-static propagation analysis. A comparison with results from automated fatigue propagation analysis was not performed at this point since the current version of Abaqus does not include this capability under variable mixed-mode conditions. In general, good agreement between the results obtained from the quasi-static propagation analysis and the benchmark results were achieved. Overall, the benchmarking procedure proved valuable for analysis verification.

Krueger, Ronald

A Benchmark Example for Delamination Propagation Predictions Based on the Single Leg Bending Specimen under Quasi-static and Fatigue Loading

Benchmark examples based on Single Leg Bending (SLB) specimens with equal and unequal bending arm thicknesses were used to assess the performance of delamination prediction capabilities in finite element codes. First, the development of the quasi-static benchmark cases using the Virtual Crack Closure Technique (VCCT) is discussed in detail. Second, based on the quasi-static benchmark results, additional benchmark cases to assess delamination propagation under fatigue loading are created. Third, the application is demonstrated for the commercial finite element code Abaqus Standard 2018. The benchmark cases are compared to results obtained from VCCT-based, automated quasi-static propagation analysis. A comparison with results from automated fatigue propagation analysis was not performed at this point since the current version of Abaqus does not include this capability under variable mixed-mode conditions. In general, good agreement between the results obtained from the quasi-static propagation analysis and the benchmark results were achieved. Overall, the benchmarking procedure proved valuable for analysis verification.

Ronald Krueger

Discrete fracture network model benchmarks developed and applied in a DECOVALEX-2023 repository performance assessment study

This study presents newly developed benchmarks for modeling flow and transport within discrete fracture networks (DFNs) and useful methods for analyzing the results. The new benchmarks are designed to test modeling approaches for use in probabilistic performance assessment models of deep geologic repositories in fractured rock. The benchmarks simulate flow and transport through a 1 km 3 block of fractured rock. The first simulates migration of a short pulse of tracer through a simple network of four intersecting fractures. The second adds 1089 stochastically generated fractures. The third changes the pulse to a continuous point source. Evaluation of model performance relies on moment analysis and comparison of the results of different models. The expected nondimensional first moment of the conservative tracer for each benchmark is 1. The benchmarks were simulated by teams from Canada, Czechia, Germany, Korea, Sweden, Taiwan, and the United States as part of a DECOVALEX-2023 study (decovalex.org). The teams used various approaches, including explicit DFN modeling, DFN upscaling to an equivalent continuous porous medium (ECPM), and a combination of both methods. Transport mechanisms are modeled using either the advection-dispersion equation or particle tracking. Results demonstrate strong agreement among the models in breakthrough behavior up to the 75th percentile. Significant deviations in first moments and well-clustered outputs led to the identification of inaccuracies in several models. Such findings exemplify the benefit of exercising these benchmarks and using the presented methods to test DFN flow and transport models.

Benchmark

Depletion Benchmark of the AFIP-7 Experiment in the Advanced Test Reactor

Reactor physics depletion benchmarks for low-enriched uranium fuel are limited in number. In particular, there is very limited data for LEU benchmarks for U-10Mo (Uranium-10% Molybdenum) plate fuel developed for use in U.S. high-performance research reactors (USHPRR). USHPRR includes the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), University of Missouri Research Reactor (MURR), Massachusetts Institute of Technology Reactor (MITR), and National Bureau of Standards Reactor (NBSR) at the National Institute of Science and Technology. These reactors are fueled with high-enriched uranium dispersed fuel in a silicon/aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiated fueled specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This report documents the development of a benchmark model of the irradiation of the ATR Full -size plate In center flux trap Position 7 (AFIP-7) experiment.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Depletion benchmark for a high-assay low-enriched uranium fuel experiment in the advanced test reactor

Reactor physics depletion benchmarks for high-assay low-enriched uranium (HALEU) fuel are limited in number. In particular, there is limited data for HALEU benchmarks for U-10Mo (uranium-10% molybdenum) plate fuel that is being developed for use in the United States’ high performance research reactors including the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), Massachusetts Institute of Technology Reactor (MITR), University of Missouri Research Reactor (MURR), National Bureau of Standards Reactor (NBSR). These six reactors currently operate with highly enriched uranium dispersed fuel in an aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiating fuel specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This paper documents the development of a benchmark model of the irradiation of the ATR Full-size plate In center flux trap Position 7 (AFIP-7) experiment using the depletion codes MC21 and Advanced Dimensional Depletion for Engineering of Reactors (ADDER).

Nielsen, Joseph W. [Idaho National Laboratory (INL

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris