Search NASASearch

SEARCH · Search NASA

Results for “test metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Improving Climate Projections Using "Intelligent" Ensembles

Recent changes in the climate system have led to growing concern, especially in communities which are highly vulnerable to resource shortages and weather extremes. There is an urgent need for better climate information to develop solutions and strategies for adapting to a changing climate. Climate models provide excellent tools for studying the current state of climate and making future projections. However, these models are subject to biases created by structural uncertainties. Performance metrics-or the systematic determination of model biases-succinctly quantify aspects of climate model behavior. Efforts to standardize climate model experiments and collect simulation data-such as the Coupled Model Intercomparison Project (CMIP)-provide the means to directly compare and assess model performance. Performance metrics have been used to show that some models reproduce present-day climate better than others. Simulation data from multiple models are often used to add value to projections by creating a consensus projection from the model ensemble, in which each model is given an equal weight. It has been shown that the ensemble mean generally outperforms any single model. It is possible to use unequal weights to produce ensemble means, in which models are weighted based on performance (called "intelligent" ensembles). Can performance metrics be used to improve climate projections? Previous work introduced a framework for comparing the utility of model performance metrics, showing that the best metrics are related to the variance of top-of-atmosphere outgoing longwave radiation. These metrics improve present-day climate simulations of Earth's energy budget using the "intelligent" ensemble method. The current project identifies several approaches for testing whether performance metrics can be applied to future simulations to create "intelligent" ensemble-mean climate projections. It is shown that certain performance metrics test key climate processes in the models, and that these metrics can be used to evaluate model quality in both current and future climate states. This information will be used to produce new consensus projections and provide communities with improved climate projections for urgent decision-making.

Baker, Noel C.

Improving Climate Projections Using "Intelligent" Ensembles

Recent changes in the climate system have led to growing concern, especially in communities which are highly vulnerable to resource shortages and weather extremes. There is an urgent need for better climate information to develop solutions and strategies for adapting to a changing climate. Climate models provide excellent tools for studying the current state of climate and making future projections. However, these models are subject to biases created by structural uncertainties. Performance metrics-or the systematic determination of model biases-succinctly quantify aspects of climate model behavior. Efforts to standardize climate model experiments and collect simulation data-such as the Coupled Model Intercomparison Project (CMIP)-provide the means to directly compare and assess model performance. Performance metrics have been used to show that some models reproduce present-day climate better than others. Simulation data from multiple models are often used to add value to projections by creating a consensus projection from the model ensemble, in which each model is given an equal weight. It has been shown that the ensemble mean generally outperforms any single model. It is possible to use unequal weights to produce ensemble means, in which models are weighted based on performance (called "intelligent" ensembles). Can performance metrics be used to improve climate projections? Previous work introduced a framework for comparing the utility of model performance metrics, showing that the best metrics are related to the variance of top-of-atmosphere outgoing longwave radiation. These metrics improve present-day climate simulations of Earth's energy budget using the "intelligent" ensemble method. The current project identifies several approaches for testing whether performance metrics can be applied to future simulations to create "intelligent" ensemble-mean climate projections. It is shown that certain performance metrics test key climate processes in the models, and that these metrics can be used to evaluate model quality in both current and future climate states. This information will be used to produce new consensus projections and provide communities with improved climate projections for urgent decision-making.

Baker, Noel C.

Short-Term Load Forecasting Considering EV Charging Loads with Prediction Interval Evaluation

Short-term load forecasting plays a critical role in power system planning and operation. Along with the electrification of various loads, electricity demands are becoming increasingly hard to predict. Notably, the recent rise in electric vehicles (EVs) has further contributed to this unpredictability. To address this issue, this paper proposes a probabilistic load forecasting strategy utilizing Gaussian process regression, structured in a day-ahead manner. While many works focus on deterministic prediction, probabilistic forecasting offers additional insights into variability and uncertainty, enabling more flexible and reliable operation for power systems. To enhance the accuracy of the load forecasting model, the inputs include features related to EV charging habits as well as commonly used weather information. The load forecasting results are evaluated using various metrics, including conventional ones that assess the accuracy of point forecasts, as well as additional metrics that test the reliability of prediction intervals. The proposed load forecasting method is finally tested on real residential power consumption data and EV charging data sampled from real-world sources. The results prove that the new features can greatly improve the performance of the load forecasting method.

electrical vehicle

A report on the gravitational redshift test for non-metric theories of gravitation

The frequencies of two atomic hydrogen masers and of three superconducting cavity stabilized oscillators were compared as the ensemble of oscillators was moved in the Sun's gravitational field by the rotation and orbital motion of the Earth. Metric gravitation theories predict that the gravitational redshifts of the two types of oscillators are identical, and that there should be no relative frequency shift between the oscillators; nonmetric theories, in contrast, predict a frequency shift between masers and SCSOs that is proportional to the change in solar gravitational potential experienced by the oscillators. The results are consistent with metric theories of gravitation at a level of 2%.

Source record

Microcomputer-based tests for repeated-measures: Metric properties and predictive validities

A menu of psychomotor and mental acuity tests were refined. Field applications of such a battery are, for example, a study of the effects of toxic agents or exotic environments on performance readiness, or the determination of fitness for duty. The key requirement of these tasks is that they be suitable for repeated-measures applications, and so questions of stability and reliability are a continuing, central focus of this work. After the initial (practice) session, seven replications of 14 microcomputer-based performance tests (32 measures) were completed by 37 subjects. Each test in the battery had previously been shown to stabilize in less than five 90-second administrations and to possess retest reliabilities greater than r = 0.707 for three minutes of testing. However, all the tests had never been administered together as a battery and they had never been self-administered. In order to provide predictive validity for intelligence measurement, the Wechsler Adult Intelligence Scale-Revised and the Wonderlic Personnel Test were obtained on the same subjects.

Kennedy, Robert S.

Aerocapture, Entry, Descent and Landing (AEDL) Human Planetary Landing Systems. Section 10: AEDL Analysis, Test and Validation Infrastructure

Contents include the following: 3 Listing of critical capabilities (knowledge, procedures, training, facilities) and metrics for validating that they are mission ready. Examples of critical capabilities and validation metrics: ground test and simulations. Flight testing to prove capabilities are mission ready. Issues and recommendations.

Arnold, J.

Psychoacoustic Test to Determine Sound Quality Metric Indicators of Rotorcraft Noise Annoyance

Noise certification metrics such as Effective Perceived Noise Level and Sound Exposure Level are used to ensure that helicopters meet regulations, but these metrics may not be good indicators of annoyance since noise complaints against helicopters persist. Sound quality (SQ) metrics, specifically fluctuation strength, tonality, impulsiveness, roughness, and sharpness, are explored to determine their relationship with annoyance. A psychoacoustic test was conducted at the NASA Langley Research Center Exterior Effects Room to assess annoyance to helicopter-like sounds over a range of SQ metric values. The amplitude, phase, and frequency of the AS350 helicopter main and tail rotor blade passage signal harmonics were manipulated to produce 105 unique helicopter-like sounds with prescribed values of SQ metrics. All sounds were set to roughly the same loudness level. These sounds were played to 40 subjects who rated each sound for annoyance. Analyses given in this paper point to which SQ metrics are important to the helicopter noise annoyance response.

Krishnamurthy, Siddhartha

A Superposed Metric for Spinning Black Hole Binaries Approaching Merger

We construct an approximate metric that represents the spacetime of spinning binary black holes (BBH) approaching merger. We build the metric as an analytical superposition of two Kerr metrics in harmonic coordinates, where we transform each black hole term with time-dependent boosts describing an inspiral trajectory. The velocities and trajectories of the boost are obtained by solving the post-Newtonian (PN) equations of motion at 3.5 PN order. We analyze the spacetime scalars of the new metric and we show that it is an accurate approximation of Einstein’s field equations in vacuum for a BBH system in the inspiral regime. Furthermore, to prove the effectiveness of our approach, we test the metric in the context of a 3D general relativistic magnetohydrodynamical (GRMHD) simulation of accreting minidisks around the black holes. We compare our results with a previous well-tested spacetime construction based on the asymptotic matching method. We conclude that our new spacetime is well-suited for long-term GRMHD simulations of spinning binary black holes on their way to the merger.

Luciano Combi

Real-Time Assessment of Robot Performance During Remote Exploration Operations

To ensure that robots are used effectively for exploration missions, it is important to assess their performance during operations. We are investigating the definition and computation of performance metrics for assessing remote robotic operations in real-time. Our approach is to monitor data streams from robots, compute performance metrics, and provide Web-based displays of these metrics for assessing robot performance during operations. We evaluated our approach for measuring robot performance with the K10 rovers from NASA Ames Research Center during a field test at Moses Lake Sand Dunes (WA) in June 2008. In this paper we present the results of evaluating our software for robot performance and discuss our conclusions from this evaluation for future robot operations.

SBIR TOPIC X7.02 PHASE 1

Evaluating Core Quality for a Mars Sample Return Mission

Sample return missions, including the proposed Mars Sample Return (MSR) mission, propose to collect core samples from scientifically valuable sites on Mars. These core samples would undergo extreme forces during the drilling process, and during the reentry process if the EEV (Earth Entry Vehicle) performed a hard landing on Earth. Because of the foreseen damage to the stratigraphy of the cores, it is important to evaluate each core for rock quality. However, because no core sample return mission has yet been conducted to another planetary body, it remains unclear as to how to assess the cores for rock quality. In this report, we describe the development of a metric designed to quantitatively assess the mechanical quality of any rock cores returned from Mars (or other planetary bodies). We report on the process by which we tested the metric on core samples of Mars analogue materials, and the effectiveness of the core assessment metric (CAM) in assessing rock core quality before and after the cores were subjected to shocking (g forces representative of an EEV landing).

Mars samples

Ensuring Success of Adaptive Control Research Through Project Lifecycle Risk Mitigation

Lessons Learne: 1. Design-out unnecessary risk to prevent excessive mitigation management during flight. 2. Consider iterative checkouts to confirm or improve human factor characteristics. 3. Consider the total flight test profile to uncover unanticipated human-algorithm interactions. 4. Consider test card cadence as a metric to assess test readiness. 5. Full-scale flight test is critical to development, maturation, and acceptance of adaptive control laws for operational use.

Pavlock, Kate M.

A call to standardize metrics for monitoring baleen whales near marine construction activities

Effective monitoring is necessary to protect marine mammal species during the construction of offshore infrastructure. The tools for detecting or monitoring marine mammals span traditional (e.g., visual observers, optical cameras), to newer (e.g., passive acoustic monitoring, infrared cameras, tags), and emerging (e.g., satellite imagery, environmental DNA, dimethyl sulfide concentration) technologies. Some are better suited for use during offshore development; however, peer-reviewed literature does not typically evaluate and report on the performance of these various technologies. We define a minimum set of metrics related to efficacy (i.e., confusion matrix, precision and recall, probability of missed mitigation), detection range (i.e., maximum and reliable detection range, spatial resolution), and data delivery (i.e., detection latency, system reliability, temporal resolution) that we recommend are needed to assess the utility of monitoring technologies for this purpose. Following a literature review of relevant studies, we highlight which publications reported these metrics and used multiple technologies to compare relative performance. We also emphasize the benefits of multi-modal approaches and recommend performance assessments through modeling or large-scale collaborative field testing. These metrics will standardize data collection, reporting, and analysis; promote consistent and comparable results; and foster collaboration among developers, regulatory agencies, and scientists. This may lead to the co-development of technology that achieves multiple goals, has greater application, and can answer research questions while collecting data to fulfill permitting requirements. These metrics may also inform decisions on what systems regulatory agencies might consider using and reduce monitoring costs, which is critical to support the marine sector's rapid growth alongside marine mammal conservation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Coverage Metrics for Model Checking

When using model checking to verify programs in practice, it is not usually possible to achieve complete coverage of the system. In this position paper we describe ongoing research within the Automated Software Engineering group at NASA Ames on the use of test coverage metrics to measure partial coverage and provide heuristic guidance for program model checking. We are specifically interested in applying and developing coverage metrics for concurrent programs that might be used to support certification of next generation avionics software.

Penix, John

Provider of Services for Urban Air Mobility (PSU) Prototype Simulation (X5) Final Report

Urban Air Mobility (UAM) is a new air transportation service concept to carry passengers or cargo in metropolitan areas, leveraged by innovative aircraft and air traffic automation technologies. NASA has conducted a series of simulations, called the X-series simulation, to evaluate the UAM concept of operations and support the development of airspace procedures and services for UAM operations. The simulation called “X5” was conducted in 2023 to test a Provider of Services for UAM (PSU) prototype developed by NASA for UAM flight planning, strategic conflict management support, and data exchange between UAM operators. In this simulation, two strategic conflict management capabilities, Demand-Capacity Balancing and Sequencing and Scheduling, were further investigated. This document describes the UAM system architecture modeled, the X5 simulation environment to be executed (e.g., traffic scenario and UAM airspace construct), and the strategic conflict management processes developed and evaluated in this study. Then, the simulation results are provided using several system performance metrics, such as the number of operations planned and activated, demand-capacity imbalances detected and resolved, and pre-departure delays. Based on these metrics, the test findings and lessons learned from this simulation are discussed. NASA developed a PSU prototype as part of a reference implementation of UAM system architecture and evolved strategic conflict management capabilities for UAM operations from the previous collaborative simulations with industry partners. Below is the summary of the achievements: - Aligned NASA’s UAM reference architecture with the FAA’s UAM ConOps notional architecture - Extended UAM airspace management capabilities to include 1) Demand-Capacity Balancing (DCB) to ensure operators coordinate planned usage of shared vertiports, and 2) Sequencing and Scheduling (S&S) at UAM corridor entry and exit points to help facilitate an orderly flow of traffic - Defined the PSU information exchange APIs and requirements towards informing industry standards - Developed and tested a NASA PSU prototype as reference implementation to validate the requirements and APIs - Developed a prototype service connecting NASA’s PSU and the FAA system for testing future PSU-ATM interface requirements - Tested NASA-developed assumptions for UAM operations such as airspace design, procedures, vehicle performance, and strategic conflict management methods to inform future Cooperative Operating Practices (COPs) development with industry - Evaluated system performance metrics such as number of simultaneous operations and ground delays that can help define system-level requirements. The simulation results showed that the UAM traffic demand could be managed to minimize the needs of tactical separation provision with ground delays assigned by DCB and S&S. These accomplishments and the lessons learned from the PSU Prototype X5 simulation activities will be valuable inputs for the Air Mobility Pathfinders (AMP) project, which is NASA’s new project to create and evaluate a reference architecture for safe, secure, and scalable UAM operations.

Simulation

Comment on “Advanced Testing of Low, Medium, and High ECS CMIP6 GCM Simulations Versus ERA5-T2m” by N. Scafetta (2022)

Scafetta (2022, https://doi.org/10.1029/2022gl097716) purports to test Coupled Model Intercomparison Project Phase 6 (CMIP6) climate models through a comparison of temperature changes over three decades. Unfortunately, the paper contains numerous conceptual and statistical errors that undermine all of the conclusions. First, no uncertainty is given for the observational temperature difference, making it impossible to assess compatibility with any model result. Second, the CMIP6 data are the ensemble means for each model, but the metric being tested is sensitive to the internal variability and so the full ensemble for each model must be used. When this is corrected, the conclusion that “all models with ECS > 3.0°C overestimate the observed global surface warming” is not sustained. Third, the statistical test in Section 2 would reject all models even in a perfect model setup given sufficient ensemble members, thus the second conclusion “that spatial t-statistics rejects the data-model agreement” is also not sustainable.

CMIP6

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Modeling Framework for Data Center

This chapter highlights the critical need for advanced modeling of data centers due to their rapidly increasing energy consumption and impact on grid reliability. Driven by the demand for AI applications, data centers are projected to consume a significant portion of US energy by 2028, putting stress on an already challenged power grid. The chapter emphasizes the importance of "fast" time-scale models to understand the dynamic interactions between data centers and the grid, especially given the rapid power fluctuations of AI workloads. It outlines a modeling framework that includes both offline and real-time EMT domain simulations, detailing the necessary representations for various components like utility interfaces, transformers, IT loads, UPS, cooling loads, Battery Energy Storage Systems (BESS), generators, protection systems, and higher-level control systems. While standard simulation tools like PSCAD offer basic models, custom development is often required to accurately capture the unique and fast-changing behaviors of modern data centers. The chapter also discusses key metrics and test cases for validating these models, focusing on transient load responses, protection relay coordination, and demand flexibility. Finally, it addresses the challenges of modeling large-scale data centers, such as computational complexity and the trade-off between model fidelity and practicality, suggesting hybrid modeling approaches as a solution. The overarching goal is to create a robust framework that helps assess data center impacts on grid stability, identify vulnerabilities, and inform the development of standards for reliable integration of these large loads into the bulk power system.

25 ENERGY STORAGE

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory