Search NASASearch

SEARCH · Search NASA

Results for “test metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Use of Spectral Analysis of Singular Values as a Test Metric for Impedance Matched Multi-Axis Test Trials

One of many challenges in the implementation of multiple exciter testing is establishing a reasonable set of test metrics to measure the quality of testing. This is especially true in the application of Impedance Matched Multi-Axis Testing; in that it is possible to have very large spectral density matrices that serve as reference criteria. While there exist plotting schemes to view a spectral density matrix, it is often necessary to break the overlay of reference and test results into subsections of the matrices to get sufficient resolution to interpret the data. In addition, as one attempts to control multiple locations on a structure, implementation of classical single degree-of freedom test tolerances across all channels and associated cross spectra is simply not feasible. Hence it is challenging to evaluate overall test quality. The use of spectral views of the dominant singular values from the singular value decomposition of the spectral density matrices and metrics based upon them is proposed for establishing a set of compact metrics for evaluating test quality. A laboratory experiment will be included to demonstrate this proposed technique.

Vibration Testing

Use of Spectral Analysis of Singular Values as a Test Metric for IMMAT Trials

One of many challenges in the implementation of multiple exciter testing is establishing a reasonable set of test metrics to measure the quality of testing. This is especially true in the application of Impedance Matched Multi-Axis Testing; in that it is possible to have very large spectral density matrices that serve as reference criteria. While there exist plotting schemes to view a spectral density matrix, it is often necessary to break the overlay of reference and test results into subsections of the matrices to get sufficient resolution to interpret the data. In addition, as one attempts to control multiple locations on a structure, implementation of classical single degree-of freedom test tolerances across all channels and associated cross spectra is simply not feasible. Hence it is challenging to evaluate overall test quality. The use of spectral views of the dominant singular values from the singular value decomposition of the spectral density matrices and metrics based upon them is proposed for establishing a set of compact metrics for evaluating test quality. A laboratory experiment will be included to demonstrate this proposed technique.

Impedance Matched Multi-Axis Testing

An evaluation of software testing metrics for NASA's mission control center

Software metrics are used to evaluate the software development process and the quality of the resulting product. Five metrics were used during the testing phase of the Shuttle Mission Control Center Upgrade at the NASA Johnson Space Center. All but one metric provided useful information. Based on the experience, it is recommended that metrics be used during the test phase of software development and additional candidate metrics are proposed for further study.

Stark, George E.

Rotorcraft Sound Quality Metric Test 1: Stimuli Generation and Supplemental Analyses

A psychoacoustic test was conducted at the NASA Langley Research Center Exterior Effects Room (EER) to assess annoyance to simulated helicopter sounds over a range of sound quality (SQ) metric values. Initial findings identified important SQ metrics as sharpness, tonality, and fluctuation strength. This document is a supplement to the initial findings in which the following are discussed: (i) a detailed treatment of the sound generation process, (ii) the impact of analyzing results with stimuli measured in the EER instead of the intended synthesized stimuli, (iii) an evaluation of annoyance responses with certification metrics, and (iv) adjunct analyses related to the test methodology.

Psychoacoustic test, psychoacoustics, sound qualit

EVM and Schedule Management

The objective of EVMS surveillance is to ensure that the management control processes that support the performance measurement baseline (PMB) are in place, compliant with the EVMS guidelines, are routinely being used, and provide timely and reliable data. The PMB is a triple constraint where the constraints are schedule, budget and scope. For Surveillance, NASA uses the DCMA EVM Compliance Metrics (DECM) Tests that are aligned with the EIA-748 EVM Standard. Guidelines 6 is Scheduling Work, and DECM has 23 Tests for evaluating if the IMS supports project goals in its planning, statusing and forecasting. This session will focus on the Test Metric that analyzes forecast start/finish dates riding the status date of the IMS for two or more consecutive months as an example of how surveillance works in concert with IMS health checks. It will cover how to run the test to recognize trends and how this test helps ensure that the forecast is credible in support of critical path analysis.

EVM

Exploring Sustainability in Scientific Software through Code Quality & Test Coverage Metrics

Context: Scientific open-source software (SciOSS) plays a foundational role in research and engineering, yet its long-term sustainability has often been overlooked and remains a significant concern. Objective: This study investigates the long-term sustainability of SciOSS through code and test quality metrics. Method: We analyze CASS Software Portfolio projects, classifying them by sustainability and comparing their code structure, test coverage, and links between code quality and testing across the dataset. Results: Sustainable projects show higher, more consistent test coverage and clearer code-test correlations, while unsustainable ones show weaker patterns. Overall, test coverage is low in scientific software, and high complexity and coupling reduce testability. Conclusion: In this study, we present a practical, data-driven approach for assessing sustainability in scientific software, offering a foundation for evaluating long-term software health and supporting future efforts in quality assurance and sustainability monitoring.

Md mushfiqur rahman, Sheikh [University of Tenness

Fighter agility metrics, research, and test

Proposed new metrics to assess fighter aircraft agility are collected and analyzed. A framework for classification of these new agility metrics is developed and applied. A completed set of transient agility metrics is evaluated with a high fidelity, nonlinear F-18 simulation provided by the NASA Dryden Flight Research Center. Test techniques and data reduction methods are proposed. A method of providing cuing information to the pilot during flight test is discussed. The sensitivity of longitudinal and lateral agility metrics to deviations from the pilot cues is studied in detail. The metrics are shown to be largely insensitive to reasonable deviations from the nominal test pilot commands. Instrumentation required to quantify agility via flight test is also considered. With one exception, each of the proposed new metrics may be measured with instrumentation currently available. Simulation documentation and user instructions are provided in an appendix.

Liefer, Randall K.

A Dynamic Testing Complexity Metric

This paper introduces a dynamic metric that is based on the estimated ability of a program to withstand the effects of injected "semantic mutants" during execution by computing the same function as if the semantic mutants had not been injected. Semantic mutants include: (1) syntactic mutants injected into an executing program and (2) randomly selected values injected into an executing program's internal states. The metric is a function of a program, the method used for injecting these two types of mutants, and the program's input distribution; this metric is found through dynamic executions of the program. A program's ability to withstand the effects of injected semantic mutants by computing the same function when executed is then used as a tool for predicting the difficulty that will be incurred during random testing to reveal the existence of faults, i.e., the metric suggests the likelihood that a program will expose the existence of faults during random testing assuming faults were to exist. If the metric is applied to a module rather than to a program, the metric can be used to guide the allocation of testing resources among a program's modules. In this manner the metric acts as a white-box testing tool for determining where to concentrate testing resources. Index Terms: Revealing ability, random testing, input distribution, program, fault, failure.

Voas, Jeffrey

Testing Strategies for Model-Based Development

This report presents an approach for testing artifacts generated in a model-based development process. This approach divides the traditional testing process into two parts: requirements-based testing (validation testing) which determines whether the model implements the high-level requirements and model-based testing (conformance testing) which determines whether the code generated from a model is behaviorally equivalent to the model. The goals of the two processes differ significantly and this report explores suitable testing metrics and automation strategies for each. To support requirements-based testing, we define novel objective requirements coverage metrics similar to existing specification and code coverage metrics. For model-based testing, we briefly describe automation strategies and examine the fault-finding capability of different structural coverage metrics using tests automatically generated from the model.

Heimdahl, Mats P. E.

A Framework for Evaluating Climate Model Performance Metrics

The CMIP5 archive contains future climate projections from over 50 models provided by dozens of modeling centers from around the world. Individual model projections, however, are subject to biases created by structural model uncertainties. As a result, ensemble averaging of multiple models is often used to add value to model projections: consensus projections have been shown to consistently outperform individual models. Previous reports for the IPCC establish climate change projections based on an equal-weighted average of all model projections. However, certain models reproduce climate processes better than other models. Should models be weighted based on performance? Unequal ensemble averages have previously been constructed using a variety of mean state metrics. What metrics are most relevant for constraining future climate projections? This project develops a framework for systematically testing metrics in models to identify optimal metrics for unequal weighting multi-model ensembles. A unique aspect of this project is the construction and testing of climate process-based model evaluation metrics. A climate process-based metric is defined as a metric based on the relationship between two physically related climate variables?e.g., outgoing longwave radiation and surface temperature. Metrics are constructed using high-quality Earth radiation budget data from NASA's Clouds and Earth's Radiant Energy System (CERES) instrument and surface temperature data sets. It is found that regional values of tested quantities can vary significantly when comparing weighted and unweighted model ensembles. For example, one tested metric weights the ensemble by how well models reproduce the time-series probability distribution of the cloud forcing component of reflected shortwave radiation. The weighted ensemble for this metric indicates lower simulated precipitation (up to .7 mm/day) in tropical regions than the unweighted ensemble: since CMIP5 models have been shown to overproduce precipitation, this result could indicate that the metric is effective in identifying models which simulate more realistic precipitation. Ultimately, the goal of the framework is to identify performance metrics for advising better methods for ensemble averaging models and create better climate predictions.

Noel C Baker

Progress in Flaps Down Flight Reynolds Number Testing Techniques at the NTF

A series of NASA/Boeing cooperative low speed wind tunnel tests was conducted in the National Transonic Facility (NTF) between 2003 and 2004 using a semi-span high lift model representative of the 777-200 aircraft. The objective of this work was to develop the capability to acquire high quality, low speed (flaps down) wind tunnel data at up to flight Reynolds numbers in a facility originally optimized for high speed full span models. In the course of testing, a number of facility and procedural improvements were identified and implemented. The impact of these improvements on key testing metrics data quality, productivity, and so forth - was significant, and is discussed here, together with the relevance of these metrics as applied to cryogenic wind tunnel testing in general. Details of the improvements at the NTF are discussed in AIAA-2006-0508 (Recent Improvements in Semi-span Testing at the National Transonic Facility). The development work at the NTF culminated with validation testing of a 787-8 semi-span model at full flight Reynolds number in the first quarter of 2006.

Payne, Frank

NASA biological and physical sciences databases: who’s the FAIRest of them all?

Conceptual models are a key part of the foundation of scientific study. Scientific data discovery and retrieval are often inaccurate and incomplete because these models are not sufficiently well-incorporated into data retrieval systems. Systems often don’t provide the necessary tools to those producing scientific data to fully and unambiguously annotate them and the result is consumers of the data cannot find them efficiently. The capability of data archives to provide these tools to link data to underlying conceptual models is one of dimensions of the recently developed “FAIR” principles (https://www.go-fair.org/fair-principles/ ), and is key to many automated processes being able to operate on these data, particularly analytics involving artificial intelligence. We used an open-source web service to measure the FAIR compliance of the three data archives operated by NASA for the biological and physical sciences: the Life Sciences Data Archive, the Physical Sciences Informatics database, and GeneLab. The service ingests references to data sets in these archives, and then executes domain-non-specific examinations of these data and metadata that test compliance to the FAIR principles. Of the 22 metrics tested, GeneLab passed 11 (50%), and PSI and LSDA each passed 7 (32%). These data were gathered using only one representative data set from each archive and we anticipate variability in results as we continue to apply these metrics to other data. A preliminary study of the failure traces for each metric suggests there is a wide range of effort and complexity in the enhancements required for each system to elevate FAIR compliance, and this is the subject of continued investigation. This information has been and will likely continue to be important information in planning these enhancements, with the goal of increased readiness of the data for automated processes.

database

Examining functional acceptance testing with structural coverage metrics

Functionally generated acceptance tests are examined using structural coverage metrics. A method of comparing acceptance tests and operational usage was generated. Acceptance tests are prepresentative of operational usage except for the mix of statement types. Structural coverage metrics may provide insight into software faults.

Ramsey, J.

Comparison of Image Restoration Methods for Lunar Epithermal Neutron Emission Mapping

Orbital measurements of neutrons by the Lunar Exploring Neutron Detector (LEND) onboard the Lunar Reconnaissance Orbiter are being used to quantify the spatial distribution of near surface hydrogen (H). Inferred H concentration maps have low signal-to-noise (SN) and image restoration (IR) techniques are being studied to enhance results. A single-blind. two-phase study is described in which four teams of researchers independently developed image restoration techniques optimized for LEND data. Synthetic lunar epithermal neutron emission maps were derived from LEND simulations. These data were used as ground truth to determine the relative quantitative performance of the IR methods vs. a default denoising (smoothing) technique. We review and used factors influencing orbital remote sensing of neutrons emitted from the lunar surface to develop a database of synthetic "true" maps for performance evaluation. A prior independent training phase was implemented for each technique to assure methods were optimized before the blind trial. Method performance was determined using several regional root-mean-square error metrics specific to epithermal signals of interest. Results indicate unbiased IR methods realize only small signal gains in most of the tested metrics. This suggests other physically based modeling assumptions are required to produce appreciable signal gains in similar low SN IR applications.

McClanahan, T. P.

Towards Scaling Law Analysis For Spatiotemporal Weather Data

Compute-optimal scaling laws are relatively well studied for NLP and CV, where objectives are typically single-step and targets are comparatively homogeneous. Weather forecasting is harder to characterize in the same framework: autoregressive rollouts compound errors over long horizons, outputs couple many physical channels with disparate scales and predictability, and globally pooled test metrics can disagree sharply with per-channel, late-lead behavior implied by short-horizon training. We extend neural scaling analysis for autoregressive weather forecasting from single-step training loss to long rollouts and per-channel metrics. We quantify (1) how prediction error is distributed across channels and how its growth rate evolves with forecast horizon, (2) if power law scaling holds for test error, relative to rollout length when error is pooled globally, and (3) how that fit varies jointly with horizon and channel for parameter, data, and compute-based scaling axes. We find strong cross-channel and cross-horizon heterogeneity: pooled scaling can look favorable while many channels degrade at late leads. We discuss implications for weighted objectives, horizon-aware curricula, and resource allocation across outputs.

Kiefer Jr, Alexander [ORNL] (ORCID:000000025398874

Development of a metric half-span model for interference free testing

A metric half-span model has been developed which allows measurement of aerodynamic forces and moments without support interference or model distortion. This is accomplished by combining the best features of the conventional sting/balance and half-span splitter plate supports, For example, forces and moments are measured on one-half of a symmetrical model which is mechanically supported by a sting on the nonmetric half. Tests were performed in the Langley Unitary Plan wind tunnel over a Mach range of 1.60 to 2.70 and an angle-of-attack range of 04 deg to 20 deg. Preliminary results on concept evaluation, and effect of fuselage modification to house a conventional balance and sting are presented.

Corlett, W. A.

Recrystallization driven softening and heating rate dependencies of FeCrAl nuclear fuel cladding during accident transients

A refined understanding of FeCrAl cladding behavior during rapid transients is critical for its potential deployment in light-water reactors. Current assessments focus on transient burst testing metrics such as balloon geometry, burst temperature, and hoop stress, often used as proxies for simpler conventional tensile properties. However, directly correlating isothermal tensile and creep data with accident transient scenarios remains a challenge, although it is essential for high fidelity model development. Recent modeling based on tensile tests up to 800 °C, conducted with both immediate loading and a 10-minute soak, showed that immediate loading better predicts experimental burst temperatures, indicating a thermal softening effect. Building upon this observation, the current study connects transient performance, microstructural evolution, and high-temperature tensile properties by leveraging results from C26M claddings burst tests performed at heating rates of 1–50 °C/s and hoop stresses from 25 to 100 MPa. At 25 MPa, rupture temperatures varied by only 6 °C, but at 100 MPa, the difference reached 116 °C, with faster heating yielding higher burst temperatures. Microstructural analysis identified recrystallization as the primary cause of heating rate-dependent softening, eliminating prior cold-working. In-situ thermomechanical data linked ballooning onset to localized instabilities, similar to ultimate tensile strength behavior in conventional tensile tests. High heating rates correlated with immediate loading tensile data, while lower rates matched soaked data. Furthermore, by linking burst performance to microstructural evolution and tensile properties, this work provides a foundation for more accurate modeling of FeCrAl claddings and potentially other Fe-based materials under accident conditions.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Spacecraft Tests of General Relativity

Current spacecraft tests of general relativity depend on coherent radio tracking referred to atomic frequency standards at the ground stations. This paper addresses the possibility of improved tests using essentially the current system, but with the added possibility of a space-borne atomic clock. Outside of the obvious measurement of the gravitational frequency shift of the spacecraft clock, a successor to the suborbital flight of a Scout D rocket in 1976 (GP-A Project), other metric tests would benefit most directly by a possible improved sensitivity for the reduced coherent data. For purposes of illustration, two possible missions are discussed. The first is a highly eccentric Earth orbiter, and the second a solar-conjunction experiment to measure the Shapiro time delay using coherent Doppler data instead of the conventional ranging modulation.

Anderson, John D.