Search NASA⌕ Search

SEARCH · Search NASA

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Error metrics for partially coherent wave fields

Lensless imaging methods that account for partial coherence have become very common in the past decade. However, there are no metrics in use for comparing partially coherent light fields, despite the widespread use of such metrics to compare fully coherent objects and wave fields. Here, we show how reformulating the mean squared error and Fourier ring correlation in terms of quantum state fidelity naturally generalizes them to partially coherent wave fields. We report these results fill an important gap in the lensless imaging literature and will enable quantitative assessments of the reliability and resolution of reconstructed partially coherent wave fields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Short-term electricity load forecasting: Application-driven evaluation of machine learning models across spatial and temporal scales

As we transition towards a decarbonized economy, the integration of variable renewable energy resources and new demands (e.g., electric vehicles, heat pumps) into the electricity grid places unprecedented pressure on grid operators to effectively anticipate and manage peak load. In this context, machine learning algorithms are proving to be indispensable for accurate short-term load forecasting, a crucial task to address these challenges. This study benchmarks 6 machine learning algorithms, including three neural networks and three tree-based algorithms, across various levels of spatial aggregation and time horizons (1, 4, 8, 24, and 48 h). The central contribution of this work is the comparison and analysis of load forecasting models not only based on statistical metrics, but also based on a novel error metric, which evaluates the cost implications of forecast errors for power system stakeholders. Results show that tree-based models outperform neural networks, based on statistical metrics, and yield less skewed error distributions for most spatial scales. However, through the lens of the novel error metric, neural networks are the more competitive choice, especially for forecast horizons that exceed 8 h. The study concludes with actionable recommendations to grid operators and highlights the need for the development of error metrics that link forecasting accuracy to operational costs. To promote transparency and open science, the datasets and Python code are open-sourced via a supplementary repository.

Houben, Nikolaus↗

Physics-informed transformation toward improving the machine-learned NLTE models of ICF simulations

The integration of machine-learning techniques into inertial confinement fusion (ICF) simulations has emerged as a powerful approach for enhancing computational efficiency. By replacing the costly nonlocal thermodynamic equilibrium (NLTE) model with machine-learning models, significant reductions in calculation time have been achieved. However, determining how to optimize machine-learning-based NLTE models in order to match ICF simulation dynamics remains challenging, underscoring the need for physically relevant error metrics and strategies to enhance model accuracy with respect to these metrics. Thus, we propose novel physics-informed transformations designed to emphasize energy transport, use these transformations to establish new error metrics, and demonstrate that they yield smaller errors within reduced principal-component spaces compared to conventional transformations. Published by the American Physical Society 2025

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure↗

Validating automated resonance evaluation with synthetic data

The integrity and precision of nuclear data are crucial for a broad spectrum of applications, from national security and nuclear reactor design to medical diagnostics, where the associated uncertainties can significantly impact outcomes. A substantial portion of uncertainty in nuclear data originates from the subjective biases in the evaluation process, a crucial phase in the nuclear data production pipeline. Recent advancements indicate that automation of certain routines can mitigate these biases, thereby standardizing the evaluation process and enhancing reproducibility. This research aims to provide a methodology, framework, and metrics for the validation of automated nuclear data evaluation software leveraging high-quality synthetic data that closely mimic real experimental observables. An introduced error metric provides a scale and intuitive measure of the evaluation quality by quantifying the estimate’s accuracy and performance across the specified energy range. Synthetic data provides access to experimental observables and underlying resonance parameters, enabling comparison of different evaluations. The methodology is demonstrated using Ta-181 isotope data in the resolved resonance region. The Automated Resonance Identification Subroutine (ARIS), which operates without prior resonance information, was used to test and showcase the framework’s capabilities utilizing the proposed error metrics. The results demonstrate the effectiveness of the proposed approach and framework for optimizing software parameters and testing hypotheses through “what-if” controlled experiments, such as modifying assumptions about experimental conditions or average resonance parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Development of a Method for Shape Optimization for a Gas Turbine Fuel Injector Design Using Metal-Additive Manufacturing

Adjoint shape optimization has enabled physics-based optimal designs for aerodynamic surfaces. Additive manufacturing (AM) makes it possible to manufacture complex shapes. However, there has been a gap between optimal and manufacturable surfaces due to the inherent limitations of commercial computational fluid dynamics (CFD) codes to implement geometric constraints during adjoint computation. In such cases, the design sensitivities are exported and used to perform constrained shape modifications using parametric information stored in computer aided design (CAD) files to satisfy manufacturability constraints. However, modifying the design using adjoint methods in CFD solvers and performing constrained shape modification in CAD can lead to inconsistencies due to different shape parameterization schemes. This paper describes a method to enable the simultaneous optimization of the fluid domain and impose AM manufacturability constraints, resolving one of the key issues of geometry definition for isogeometric analysis. Similar to a grid convergence study, the proposed method verifies the consistencies between shape parameterization techniques present within commercial CAD and CFD software during mesh movement as a part of the adjoint shape optimization routine. By identifying the appropriate parameters essential to a shape optimization study, the error metric between the different parameterization techniques converges to demonstrate sufficient consistencies for justifiable exchange of data between CAD and CFD. For the identified shape optimization parameters, the error metric to measure the deviation between the two parameterization schemes lies within the AM laser-powder bed fusion (L-PBF) process tolerance. Additionally, comparison for subsequent objective function calculations between iterations of the optimization loop showed acceptable differences within 1% variation between the modified geometries obtained using the two parameterization schemes. This method provides justification for the use of multiphysics guided adjoint design sensitivities computed in CFD software to perform shape modifications in CAD to incorporate AM manufacturability constraints during the shape optimization loop such that optimal designs are also additively manufacturable.

33 ADVANCED PROPULSION SYSTEMS↗

Validation of wind resource and energy production simulations for small wind turbines in the United States

Abstract. Due to financial and temporal limitations, the small wind community relies upon simplified wind speed models and energy production simulation tools to assess site suitability and produce energy generation expectations. While efficient and user-friendly, these models and tools are subject to errors that have been insufficiently quantified at small wind turbine heights. This study leverages observations from meteorological towers and sodars across the United States to validate wind speed estimates from the Wind Integration National Dataset (WIND) Toolkit, the European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis v5 (ERA5), and the Modern-Era Retrospective Analysis for Research and Applications, version 2 (MERRA-2), revealing average biases within ±0.5 m s−1 at small wind hub heights. Observations from small wind turbines across the United States provide references for validating energy production estimates from the System Advisor Model (SAM), Wind Report, MyWindTurbine.com, and Global Wind Atlas 3 (GWA3), which are seen to overestimate actual annual capacity factors by 2.5, 4.2, 11.5, and 7.3 percentage points, respectively. In addition to quantifying the error metrics, this paper identifies sources of model and tool discrepancies, noting that interannual fluctuation in the wind resource, wind speed class, and loss assumptions produces more variability in estimates than different horizontal and vertical interpolation techniques. The results of this study provide small wind installers and owners with information about these challenges to consider when making performance estimates and thus possible adjustments accordingly. Looking to the future, recognizing these error metrics and sources of discrepancies provides model and tool researchers and developers with opportunities for product improvement that could positively impact small wind customer confidence and the ability to finance small wind projects.

17 WIND ENERGY↗

Improved Representation of Horizontal Variability and Turbulence in Mesoscale Simulations of an Extended Cold-Air Pool Event

Abstract Cold-air pools (CAPs), or stable atmospheric boundary layers that form within topographic basins, are associated with poor air quality, hazardous weather, and low wind energy output. Accurate prediction of CAP dynamics presents a challenge for mesoscale forecast models in part because CAPs occur in regions of complex terrain, where traditional turbulence parameterizations may not be appropriate. This study examines the effects of the planetary boundary layer (PBL) scheme and horizontal diffusion treatment on CAP prediction in the Weather Research and Forecasting (WRF) Model. Model runs with a one-dimensional (1D) PBL scheme and Smagorinsky-like horizontal diffusion are compared with runs that use a new three-dimensional (3D) PBL scheme to calculate turbulent fluxes. Simulations are completed in a nested configuration with 3-km/750-m horizontal grid spacing over a 10-day case study in the Columbia River basin, and results are compared with observations from the Second Wind Forecast Improvement Project. Using event-averaged error metrics, potential temperature and wind speed errors are shown to decrease both with increased horizontal grid resolution and with improved treatment of horizontal diffusion over steep terrain. The 3D PBL scheme further reduces errors relative to a standard 1D PBL approach. Error reduction is accentuated during CAP erosion, when turbulent mixing plays a more dominant role in the dynamics. Last, the 3D PBL scheme is shown to reduce near-surface overestimates of turbulence kinetic energy during the CAP event. The sensitivity of turbulence predictions to the master length-scale formulation in the 3D PBL parameterization is also explored. Significance Statement In this article, we demonstrate how a new framework for modeling atmospheric turbulence improves cold pool predictions, using a case study from January 2017 in the Columbia River basin (U.S. Pacific Northwest). Cold pools are regions of cold, stagnant air that form within valleys or basins, and improved forecasts could help to mitigate the risks they pose to air quality, transportation, and wind energy production. For the chosen case study, our tests show a reduction in temperature and wind speed errors by up to a factor of 2–3 relative to standard model options. These results strongly motivate continued development of the framework as well as its application to other complex weather events.

17 WIND ENERGY↗

A North Sea in Situ Evaluation of the Fitch Wind Farm Parameterization Within the Mellor-Yamada-Nakanishi-Niino and 3D Planetary Boundary Layer Schemes

Wind resource assessments and wind power forecasts that account for wind farm wakes are sensitive to the choice of planetary boundary layer (PBL) scheme. This work compares the one-dimensional Mellor-Yamada-Nakanishi-Niino (MYNN) PBL scheme with a three-dimensional PBL (3DPBL) scheme, evaluating predictions made with both schemes against two sets of North Sea in situ observations of wind farm wakes. The optimal PBL scheme varies based on the observations (FINO1 tower vs. aircraft), the quantity of interest (wind speed vs. turbulence kinetic energy [TKE]), and the error metric (bias, centered root mean square error [cRMSE], R2, and earth mover's distance [EMD]). Whereas 3DPBL wind speeds outperform MYNN wind speeds with respect to the cRMSE at the FINO1 site located at a single point within the turbine rotor layer, 3DPBL TKE bias is larger than MYNN TKE bias when compared to aircraft observations taken 100 m above a wind farm. Wind speeds in the aircraft region are ambiguous with regard to which PBL scheme is optimal. Aircraft MYNN wind speeds outperform 3DPBL wind speeds with respect to R2 and cRMSE but underperform with respect to bias and EMD. Future evaluations across broader temporal and spatial scales may offer further insight into model differences.

17 WIND ENERGY↗

Model Calibration with Markov Chain Monte Carlo Tutorial

The purpose of this tutorial is to demonstrate how to use Markov chain Monte Carlo (MCMC) to calibrate a model. By calibration, we mean the selection of model parameters (and, when relevant, structures). A common goal in model development and diagnostics is calibration, or the identification of model structures and parameters which are consistent with data. While models can be calibrated through hand-tuning parameters or minimizing simple error metrics such as root-mean-square-error (RMSE), these approaches can underrepresent the probabilistic nature of the data-generating process, as well as the potential for multiple model configurations to be consistent with the data. Probabilistic uncertainty quantification, which is the topic of this notebook, can address these concerns. This tutorial is presented as an appendix to the e-book: Addressing Uncertainty in MultiSector Dynamics Research.

Markov chain Monte Carlo↗

Multiresolution Quantum Chemistry: Nonlinear Response Properties at the Basis Set Limit

We benchmark the accuracy of Dunning correlation-consistent Gaussian basis sets for computing frequencydependent second-order hyperpolarizabilities relevant to second-harmonic generation (SHG), using multiresolution analysis (MRA) as a reference. Basis set errors are analyzed using a unit-sphere representation of the effective hyperpolarizability vector, enabling direct assessment of directional error structure. We introduce a relative RMS total error metric that integrates directional deviations over the unit sphere and complement it with signed projection errors that distinguish over- and underestimation. Unsupervised clustering based on these signed directional metrics reveals four distinct convergence behaviors across a set of 68 molecules. Unitsphere visualizations of representative systems show that basis set errors are often highly anisotropic and localized along specific bond directions, even when global error measures appear small. Doubly augmented basis sets consistently outperform singly augmented ones, and core-polarization functions are required for uniform convergence in second-row systems. Overall, this work demonstrates that directional analysis combined with clustering provides a robust framework for understanding basis set convergence in nonlinear optical response properties.

Basis sets↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Assessing Clouds Using Satellite Observations Through Three Generations of Global Atmosphere Models

Abstract Clouds are parameterized in climate models using quantities on the model grid‐scale to approximate the cloud cover and impact on radiation. Because of the complexity of processes involved with clouds, these parameterizations are one of the key challenges in climate modeling. Differences in parameterizations of clouds are among the main contributors to the spread in climate sensitivity across models. In this work, the clouds in three generations of an atmosphere model lineage are evaluated against satellite observations. Satellite simulators are used within the model to provide an appropriate comparison with individual satellite products. In some respects, especially the top‐of‐atmosphere cloud radiative effect, the models show generational improvements. The most recent generation, represented by two distinct branches of development, exhibits some regional regressions in the cloud representation; in particular the southern ocean shows a positive bias in cloud cover. The two branches of model development show how choices during model development, both structural and parametric, lead to different cloud climatologies. Several evaluation strategies are used to quantify the spatial errors in terms of the large‐scale circulation and the cloud structure. The Earth mover's distance is proposed as a useful error metric for the passive satellite data products that provide cloud‐top pressure‐optical depth histograms. The cloud errors identified here may contribute to the high climate sensitivity in the Community Earth System Model, version 2 and in the Energy Exascale Earth System Model, version 1.

54 ENVIRONMENTAL SCIENCES↗

A variational framework for residual-based adaptivity in neural PDE solvers and operator learning

Residual-based adaptive strategies are widely used in scientific machine learning yet remain largely heuristic. We introduce a variational framework that formalizes these methods through convex transformations of the residual, where different transformations correspond to distinct objective functionals. For instance, exponential weights target uniform error minimization, while linear weights recover quadratic error minimization. This perspective reveals adaptive weighting as a means of selecting sampling distributions that optimize a primal objective, directly linking discretization choices to error metrics. This principled approach yields three key benefits: it enables systematic design of adaptive schemes, reduces discretization error by lowering estimator variance, and enhances learning dynamics by improving gradient signal-to-noise ratio. Extending the framework to operator learning, we demonstrate substantial performance gains across diverse optimizers and architectures. Our results provide a theoretical perspective for residual-based adaptivity and establish a foundation for principled discretization and training.

97 MATHEMATICS AND COMPUTING↗

Worldwide benchmarking of cost-effective radiometers for direct and diffuse irradiance

Solar energy projects can benefit from direct normal irradiance (DNI) and diffuse horizontal irradiance (DHI) measurements during all project phases. Several commercial measurement systems for DNI and DHI are available. Sun trackers with pyranometers and pyrheliometers can provide highly accurate measurements but are often impractical in solar energy applications. For less expensive and more robust sensors, it is often unclear which accuracy can be expected under a project site's specific atmospheric conditions. We address this challenge through our dedicated experimental comparison of relevant sensor systems (rotating shadowband irradiometer [short RSI], Delta-T SPN1, EKO MS-90, PyranoCam, Sunto CaptPro, Kipp & Zonen CSD3) at up to six sites worldwide. The RSI systems (rRMSD 3 to 8.6%, DNI; 4.8 to 7.6%, DHI) and PyranoCam (rRMSD 2.6 to 5.2%, DNI; 4.4 to 5.8%, DHI) exhibit similar error metrics and are the most accurate systems in the test. Delta-T SPN1 and EKO MS-90 (rRMSD 6.8 to 15%, DNI; 10.6 to 20.1%, DHI) but especially Kipp & Zonen CSD3 and Sunto CaptPro show significant deviations (rRMSD 17.7 to 20%, DNI; 33 to 58%, DHI). We evaluate the influence of relevant atmospheric parameters on the sensors' accuracies by a rather unique measurement setup. MS-90's DNI errors depend on DNI itself, with overestimations for low reference DNI. The deviations of SPN1's DHI and DNI measurements increase sharply in situations with high circumsolar irradiance. Also CaptPro and CSD3's increased measurement errors are related to circumsolar irradiance. For RSI and PyranoCam, only moderate influences on the measurements are identified, indicating a general applicability of these instruments.

14 SOLAR ENERGY↗

Understanding Biases in Simulated Cloud Radiative Effects in E3SMv3

This study systematically investigates biases in cloud radiative effects (CREs) within the recently released Energy Exascale Earth System Model version 3 (E3SMv3). Compared to its previous version (E3SMv2), E3SMv3 shows excessively strong shortwave CRE over tropical and subtropical oceans, the Southern Ocean, and the Northern Hemisphere storm tracks, which is primarily caused by an overestimation of optically intermediate low clouds. The model also displays excessive longwave CRE over the Maritime Continent and other tropical deep convection regions, resulting from an overestimation of optically thick high clouds. Implementing the Predicted Particle Properties (P3) scheme for stratiform clouds played the most significant role in these cloud changes. In addition, using a double-moment scheme for convective clouds contributed to the increase in intermediate low clouds and optically thick clouds in tropical deep convection regions. Error metrics for total cloud amount (E TCA ), cloud properties (E ctp-τ ), and cloud properties weighted by their SW and LW radiative impacts (E SW and E LW ) indicate that E3SMv3's ability to reproduce observed cloud radiative effect of low, middle, and high clouds has not been improved compared to E3SMv2. Nevertheless, the performance of E3SMv3 remains well within the spread of CMIP6 models, with E SW and E LW values smaller than those in most CMIP6 models. This study underscores the importance of integrating diverse satellite observations for robust cloud evaluation and using cloud-radiation relationships as consistency check for model errors. It suggests that further model developments focus on improving cloud microphysics and their interactions with radiation.

Environmental sciences↗

Active learning for SNAP interatomic potentials via Bayesian predictive uncertainty

Bayesian inference with a simple Gaussian error model is used to efficiently compute prediction variances for energies, forces, and stresses in the linear SNAP interatomic potential. Here, the prediction variance is shown to have a strong correlation with the absolute error over approximately 24 orders of magnitude. Using this prediction variance, an active learning algorithm is constructed to iteratively train a potential by selecting the structures with the most uncertain properties from a pool of candidate structures. The relative importance of the energy, force, and stress errors in the objective function is shown to have a strong impact upon the trajectory of their respective net error metrics when running the active learning algorithm. Batched training of different batch sizes is also tested against singular structure updates, and it is found that batches can be used to significantly reduce the number of retraining steps required with only minor impact on the active learning trajectory.

97 MATHEMATICS AND COMPUTING↗