Search NASA⌕ Search

SEARCH · Search NASA

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Benchmarking quantum trial wavefunctions for phaseless auxiliary-field quantum Monte Carlo

The phaseless auxiliary-field quantum Monte Carlo (ph-AFQMC) method is a stochastic imaginary-time projection technique for computing ground-state properties of strongly correlated quantum systems, with accuracy that depends critically on the choice of trial wavefunction. Here, we investigate ph-AFQMC with trial states prepared using parameterized quantum circuits. In this work, we present a comprehensive benchmarking study of quantum trial wavefunctions spanning unitary coupled-cluster, Hamiltonian-informed, Jastrow-inspired, and adaptively constructed ansatze. The benchmarking evaluates accuracy, expressibility, and scalability of these ansatze within the QC-AFQMC framework. We test these ansatze on linear hydrogen chains under bond stretching and find that several ansatz families produce chemically accurate ph-AFQMC energies across the dissociation curve. We have performed simulations using the CUDA-Q quantum development platform on the GPU partition of the Perlmutter supercomputer. When comparing ansatze at similar numbers of variational parameters, we find that different ansatz families yield comparable ph-AFQMC results despite exhibiting substantially different variational energies, optimization costs, and circuit depths. Our results indicate that the variational energy of an ansatz is not always a reliable indicator of its quality for ph-AFQMC and reveal instances of over-parameterization. In the strongly correlated regime, trial wavefunctions obtained from adaptive ansatze, exemplified here by ADAPT-VQE with the UCCSD operator pool, can outperform their fixed-ansatz counterparts (UCCSD) in terms of projected energies while using substantially more compact circuits, providing a flexible route to optimize quantum resources within the ph-AFQMC framework.

Rofougaran, Rod [LBNL, Berkeley; Columbia U.; PNL,↗

Hydrogen Bond Benchmark: Focal‐Point Analysis and Assessment of DFT Functionals

We performed a hierarchical, convergent ab initio benchmark study and systematically analyzed the performance of density functional approximations for describing hydrogen bonds in small neutral, cationic, and anionic complexes, as well as in larger systems involving amide, urea, deltamide, and squaramide moieties. Focal point analyses (FPA), extrapolating to the ab initio limit, were carried out using correlated wave function methods up to CCSDT(Q) for the small complexes and CCSD(T) for the larger systems, together with correlation-consistent Gaussian basis sets up to the complete basis set limit. Optimized geometries and vibrational frequencies were obtained at the CCSD(T) level. The resulting FPA hydrogen-bond energies converge within a few tenths of a kcal mol −1 . These reference data were used to evaluate 60 density functionals (including 12 dispersion-corrected), spanning the local-density approximation (LDA), generalized gradient approximations (GGAs), meta-GGAs, hybrids, meta-hybrids, double-hybrids, and range-separated hybrids. Overall, the meta-hybrid M06-2X provides the best performance for both hydrogen bond energies and geometries, while the dispersion-corrected GGAs BLYP-D3(BJ) and BLYP-D4 also yield accurate hydrogen-bond data and can serve as cost-effective options for studying large and complex systems.

coupled cluster theory↗

Benchmark study of the DTU OWC chamber with both two-way and one-way absorption

This paper reports on a benchmark study based on small-scale (1:50) measurements of a single, oscillating water column chamber mounted sideways in a long flume. The geometry of the OWC chamber is extracted from a barge-like, attenuator-type floating concept “KNSwing” with 40 chambers targeted for deployment in the Danish part of the North Sea. In addition to traditional two-way energy extraction we also consider one-way energy extraction with passive venting and compare chamber response, pressures and total absorbed energy between the two methods. A blind study was established for the numerical modeling, with participants applying several implementations of weakly nonlinear potential flow theory and commercial Navier–Stokes solvers (CFD). Both compressible and incompressible models were used for the air phase. Potential flow calculations predict more energy absorption near the chamber resonance for one-way absorption than for two-way absorption, but the opposite is found from the experimental measurements. This outcome is mainly attributed to energy losses in the experimental passive valve system, but this conclusion must be confirmed by better experimental measurements. Modeling the one-way valve in CFD proved to be very challenging and only one team was able to provide results which were generally closer to the experiments. The study illustrates the challenges associated with both numerical and experimental analysis of OWC chambers. Air compressibility effects were not found to be important at this scale, even with the large volume of additional air used for the one-way case.

16 TIDAL AND WAVE POWER↗

Using Electrochemistry to Benchmark, Understand, and Develop Noble Metal Nanoparticle Syntheses

The complex chemical nature of metal nanoparticle synthesis presents obstacles for the mechanistic understanding of nanoparticle growth and predictive synthesis design, despite significant progress in this area. Real-time characterization of the chemical processes that take place throughout nanoparticle growth will enable progress toward addressing outstanding challenges in metal nanoparticle synthesis, such as mitigating synthetic reproducibility issues, defining chemical mechanisms that direct nanoparticle growth, and designing synthetic conditions for previously unachievable combinations of nanoparticle shape and composition. In this Perspective, we present open-circuit potential (OCP) measurements as an in situ, real-time method for characterizing chemical changes during nanoparticle growth and discuss the method’s strengths in comparison to and in combination with other characterization techniques. We propose the use of OCP measurements as benchmarks for troubleshooting irreproducibility and streamlining synthetic optimization. Finally, we explore possibilities for using the increased parameter space accessible by electrodeposition to accelerate the development of shape-selective nanoparticle syntheses.

benchmarking↗

Integral Nuclear Data and Benchmarking Needs for Fusion Energy Systems

Fusion energy systems are currently being designed and optimized using radiation transport codes. To deal with the unique environment inside a fusion-based system, many of these designs incorporate novel materials able to withstand the high radiation fields, ensure adequate cooling and thermal protection, and produce tritium. Validation plays a vital role in building trust in the predictive power of these models and computational methods. Validation of a code consists of modeling documented real-world experiments and comparing the code-predicted response to the measured response. Adequate validation requires measured responses from real-world experiments, also known as integral data, that mimic the system being designed, including materials, impinging radiation, and temperature, among other variables. The most trusted integral data are experimental responses that have been through a rigorous benchmarking process that develops a recommended computational model and evaluates all experimental uncertainties. Finally, there are a few research groups around the world that have been producing integral data for fusion applications, but a substantial investment is needed to address the unique validation needs of the fusion community.

Fusion↗

A code-to-code benchmark for magneto-convection in a horizontal duct

Liquid metals and magnetic fields are used in many technical applications such as metallurgy, crystal growth and nuclear fusion reactors. When an electrically conducting fluid moves in a magnetic environment, electric currents and electromagnetic forces are generated that affect velocity and pressure losses in the flow. These magnetohydrodynamic (MHD) interactions have to be investigated to optimize the engineering processes. The characteristics of MHD flows depend on the geometrical configuration, the strength of the applied magnetic field, the electrical properties of fluid and structural materials and the thermal conditions. In the so-called blankets for fusion reactors, where liquid metals are used to breed the plasma fuel component tritium and to extract the generated heat, magneto-convective flows play a crucial role in determining heat and mass transfer. Therefore, the availability of numerical codes to simulate this type of flow is mandatory and their validation is a necessary step to guarantee the reliability of the results. For that reason, a benchmark problem has been defined to simulate liquid metal flows in a horizontal rectangular duct heated from below and exposed to a non-uniform magnetic field. Results obtained by five research groups using different codes are compared.

benchmark↗

Benchmarking of three DWM-based wake models at below-rated wind speeds

Wind turbine wake models are essential tools for predicting power losses and structural loads in wind farms. Among these, the dynamic wake meandering (DWM) model, included as a recommended approach in the International Electrotechnical Commission design standard, is a widely used engineering-fidelity method that balances accuracy and computational cost. This study compares the performance of three DWM-based wake model implementations (from the Technical University of Denmark, the National Renewable Energy Laboratory, and the Institute for Energy Technology) under below-rated wind speed conditions. Model predictions of wake flow, power output, and structural loads for a four-turbine row are evaluated across different ambient turbulence levels and wind-direction misalignments and compared against high-fidelity large-eddy simulation results. All three models captured the overall wake evolution and mean turbine performance with reasonable accuracy; their predicted time-averaged thrust and power were typically within 5 %–10 % of the large-eddy simulation benchmark. However, notable differences emerged in wake structure and unsteady load predictions, with discrepancies increasing for turbines further downstream. These differences highlight the importance of modelling choices such as wake summation and turbulence treatment, which strongly influence power-deficit and fatigue-load predictions. Comparison with large-eddy simulations reveals each approach's strengths and weaknesses, indicating where improvements are needed. Overall, the findings point to specific refinements for DWM models to improve their fidelity, ultimately enabling more robust wake predictions for wind farm design and operation.

17 WIND ENERGY↗

NASA/Navy Benchmarking Exchange (NNBE). Volume 1. Interim Report. Navy Submarine Program Safety Assurance

The NASA/Navy Benchmarking Exchange (NNBE) was undertaken to identify practices and procedures and to share lessons learned in the Navy's submarine and NASA's human space flight programs. The NNBE focus is on safety and mission assurance policies, processes, accountability, and control measures. This report is an interim summary of activity conducted through October 2002, and it coincides with completion of the first phase of a two-phase fact-finding effort.In August 2002, a team was formed, co-chaired by senior representatives from the NASA Office of Safety and Mission Assurance and the NAVSEA 92Q Submarine Safety and Quality Assurance Division. The team closely examined the two elements of submarine safety (SUBSAFE) certification: (1) new design/construction (initial certification) and (2) maintenance and modernization (sustaining certification), with a focus on: (1) Management and Organization, (2) Safety Requirements (technical and administrative), (3) Implementation Processes, (4) Compliance Verification Processes, and (5) Certification Processes.

SUBSAFE PROGRAM↗

A Segmentation Algorithm for Characterizing Rise and Fall Segments in Seasonal Cycles: an Application to Xco2 to Estimate Benchmarks and Assess Model Bias

There is more useful information in the time series of satellite-derived column-averaged carbon dioxide (XCO2) than is typically characterized. Often, the entire time series is treated at once without considering detailed features at shorter timescales, such as nonstationary changes in signal characteristics – amplitude, period and phase. In many instances, signals are visually and analytically differentiable from other portions in a time series. Each rise (increasing) and fall (decreasing) segment in the seasonal cycle is visually discernable in a graph of the time series. The rise and fall segments largely result from seasonal differences in terrestrial ecosystem production, which means that the segment's signal characteristics can be used to establish observational benchmarks because the signal characteristics are driven by similar underlying processes. We developed an analytical segmentation algorithm to characterize the rise and fall segments in XCO2 seasonal cycles. We present the algorithm for general application of the segmentation analysis and emphasize here that the segmentation analysis is more generally applicable to cyclic time series. We demonstrate the utility of the algorithm with specific results related to the comparison between satellite- and model-derived XCO2 seasonal cycles (2009–2012) for large bioregions across the globe. We found a seasonal amplitude gradient of 0.74–0.77 ppm for every 10∘ of latitude in the satellite data, with similar gradients for rise and fall segments. This translates to a south–north seasonal amplitude gradient of 8 ppm for XCO2, about half the gradient in seasonal amplitude based on surface site in situ CO2 data (∼19 ppm). The latitudinal gradients in the period of the satellite-derived seasonal cycles were of opposing sign and magnitude (−9 d per 10∘ latitude for fall segments and 10 d per 10∘ latitude for rise segments) and suggest that a specific latitude (∼2∘ N) exists that defines an inversion point for the period asymmetry. Before (after) the point of asymmetry inversion, the periods of rise segments are lesser (greater) than the periods of fall segments; only a single model could reproduce this emergent pattern. The asymmetry in amplitude and the period between rise and fall segments introduces a novel pattern in seasonal cycle analyses, but, while we show these emergent patterns exist in the data, we are still breaking ground in applying the information for science applications. Maybe the most useful application is that the segmentation analysis allowed us to decompose the model biases into their correlated parts of biases in amplitude, period and phase independently for rise and fall segments. We offer an extended discussion on how such information about model biases and the emergent patterns in satellite-derived seasonal cycles can be used to guide future inquiry and model development.

segmentation algorithm↗

Current Status of Problem 2 in the HTGR T/H Benchmark

Problem 2 of the HTGR T/H Benchmark is for modeling the depressurized conduction cooldown (DCC) transient. This presentation shares the status of code-to-code and code-to-date comparisons (Exercises 1 and 2) of Problem 2 with results from several participants. A key factor in the discrepancy between results and data in Exercise 2 is the assumed power distribution that is used to model PG-29. This talk highlights both the similarities and differences in solutions and discusses areas where further investigation is merited.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

CI/CD Efforts for Validation, Verification and Benchmarking OpenMP Implementations

Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.

Jarmusch, Aaron↗

The Surface-Topography Challenge: A Multi-Laboratory Benchmark Study to Advance the Characterization of Topography

Surface performance is critically influenced by topography in virtually all real-world applications. The current standard practice is to describe topography using one of a few industry-standard parameters. The most commonly reported number is Ra, the average absolute deviation of the height from the mean line (at some, not necessarily known or specified, lateral length scale). However, other parameters, particularly those that are scale-dependent, influence surface and interfacial properties; for example the local surface slope is critical for visual appearance, friction, and wear. The present Surface-Topography Challenge was launched to raise awareness for the need of a multi-scale description, but also to assess the reliability of different metrology techniques. In the resulting international collaborative effort, 153 scientists and engineers from 64 research groups and companies across 20 countries characterized statistically equivalent samples from two different surfaces: a “rough” and a “smooth” surface. The results of the 2088 measurements constitute the most comprehensive surface description ever compiled. We find wide disagreement across measurements and techniques when the lateral scale of the measurement is ignored. Consensus is established through scale-dependent parameters while removing data that violates an established resolution criterion and deviates from the majority measurements at each length scale. Our findings suggest best practices for characterizing and specifying topography. The public release of the accumulated data and presented analyses enables global reuse for further scientific investigation and benchmarking.

42 ENGINEERING↗

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE↗