Search NASASearch

SEARCH · Search NASA

Results for “test metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A comparison of surrogate constitutive models for viscoplastic creep simulation of HT-9 steel

Mechanistic microstructure-informed constitutive models for the mechanical response of polycrystals are a cornerstone of computational materials science. However, as these models become increasingly more complex – often involving coupled differential equations describing the effect of specific deformation modes – their associated computational costs can become prohibitive, particularly in optimization or uncertainty quantification tasks that require numerous model evaluations. To address this challenge, surrogate constitutive models that balance accuracy and computational efficiency are highly desirable. Data-driven surrogate models, that learn the constitutive relation directly from data, have emerged as a promising solution. In this work, we develop two local surrogate models for the viscoplastic response of a steel: a piecewise response surface method and a mixture of experts model. These surrogates are designed to adapt to complex material behavior, which may vary with material parameters or operating conditions. The surrogate constitutive models are applied to creep simulations of HT-9 steel, an alloy of considerable interest to the nuclear energy sector due to its high tolerance to radiation damage, using training data generated from viscoplastic self-consistent (VPSC) simulations. In conclusion, we define a set of test metrics to numerically assess the accuracy of our surrogate models for predicting viscoplastic material behavior, and show that the mixture of experts model outperforms the piecewise response surface method in terms of accuracy.

36 MATERIALS SCIENCE

Neural network-based classification and regression of magnetohydrodynamic modes in tokamaks

We present a machine learning-based magnetohydrodynamic (MHD) classifier and regressor that utilizes real or complex-valued 3D magnetic sensor array data to determine neoclassical tearing mode (NTM) onset times in tokamaks with millisecond accuracy. The input dataset consists of poloidal profiles of complex Fourier amplitudes with an n = 1 toroidal mode number from 144 human-labeled ITER Baseline Scenario discharges in the DIII-D tokamak, spanning both tearing-dominated and sawtooth-dominated regimes. Since m, n = 2,1 NTMs frequently emerge alongside sawteeth at the same frequency in this scenario, the focus is on isolating the m = 1 and m = 2 components of the n = 1 MHD mode near the tearing onset. To improve model regularization and prediction stability, singular value decomposition was applied to balance the sawtooth and tearing datasets. The enriched datasets facilitated training neural networks that learn the key distinguishing features of sawtooth and tearing modes in the poloidal profiles of their magnetic amplitude and phase. When the modes occur independently, the networks achieve perfect classification due to the modes’ distinct characteristics and low measurement noise. In the more experimentally relevant case where both modes coexist, the networks maintain exceptional performance across key metrics. Tests on synthetic data with known ground truth demonstrate the superior accuracy of the neural network trained on complex-valued input compared to models using real amplitude, phase, or pseudo-complex data, achieving both a mean time delay and standard deviation below 1 ms. Notably, standard linear regression methods fitting the dominant singular modes to the data closely match the neural network’s performance. Applying these methods across a broad range of H-mode scenarios will enable future studies to systematically identify dominant NTM triggers as scenario-specific variables, paving the way for more effective tearing mode avoidance strategies in future fusion reactor designs.

machine learning

Data-Driven Protection Software to classify fault locations by protective zone in distribution systems with high PV penetration

The software contains (a) the source codes to generate Point-on-Wave (PoW) transient data for any feeder model in Alternative Transient Program (ATP) format. Codes provide options to change different steady state settings, including the loading condition and PV capacity and transient state setting like faults type, location and initiation time (b) data post-processing source code to converted data from native format to COMTRADE, csv, HDF5 (c) Docker container to train CNN to classify fault locations by protective zone. The container takes dataset and other training parameters (sampling rate, training epochs, batch size etc) as input to train CNN. The container writes back the trained CNN model, training and testing metrics and plots to the local workstation

Ramesh, Meghana

REDESIGNING A PERFORMANCE MONITORING SOFTWARE FOR SUPERCOMPUTERS

The objective of this project was to improve upon the existing Watchr software that charts performance test metrics from the Trilinos project run on supercomputers at Sandia and elsewhere across the DOE complex. Software was iteratively designed and developed using Python Pandas and Dash data visualization to improve the extensibility and user experience of Watchr. Documentation is being maintained for future developers who want to extend the application.

Camacho, Dane Joseph [Sandia National Laboratories

Test and Evaluation Metrics of Crew Decision-Making And Aircraft Attitude and Energy State Awareness

NASA has established a technical challenge, under the Aviation Safety Program, Vehicle Systems Safety Technologies project, to improve crew decision-making and response in complex situations. The specific objective of this challenge is to develop data and technologies which may increase a pilot's (crew's) ability to avoid, detect, and recover from adverse events that could otherwise result in accidents/incidents. Within this technical challenge, a cooperative industry-government research program has been established to develop innovative flight deck-based counter-measures that can improve the crew's ability to avoid, detect, mitigate, and recover from unsafe loss-of-aircraft state awareness - specifically, the loss of attitude awareness (i.e., Spatial Disorientation, SD) or the loss-of-energy state awareness (LESA). A critical component of this research is to develop specific and quantifiable metrics which identify decision-making and the decision-making influences during simulation and flight testing. This paper reviews existing metrics and methods for SD testing and criteria for establishing visual dominance. The development of Crew State Monitoring technologies - eye tracking and other psychophysiological - are also discussed as well as emerging new metrics for identifying channelized attention and excessive pilot workload, both of which have been shown to contribute to SD/LESA accidents or incidents.

Bailey, Randall E.

Testing, Requirements, and Metrics

The criticality of correct, complete, testable requirements is a fundamental tenet of software engineering. Also critical is complete requirements based testing of the final product. Modern tools for managing requirements allow new metrics to be used in support of both of these critical processes. Using these tools, potential problems with the quality of the requirements and the test plan can be identified early in the life cycle. Some of these quality factors include: ambiguous or incomplete requirements, poorly designed requirements databases, excessive or insufficient test cases, and incomplete linkage of tests to requirements. This paper discusses how metrics can be used to evaluate the quality of the requirements and test to avoid problems later. Requirements management and requirements based testing have always been critical in the implementation of high quality software systems. Recently, automated tools have become available to support requirements management. At NASA's Goddard Space Flight Center (GSFC), automated requirements management tools are being used on several large projects. The use of these tools opens the door to innovative uses of metrics in characterizing test plan quality and assessing overall testing risks. In support of these projects, the Software Assurance Technology Center (SATC) is working to develop and apply a metrics program that utilizes the information now available through the application of requirements management tools. Metrics based on this information provides real-time insight into the testing of requirements and these metrics assist the Project Quality Office in its testing oversight role. This paper discusses three facets of the SATC's efforts to evaluate the quality of the requirements and test plan early in the life cycle, thus preventing costly errors and time delays later.

Rosenberg, Linda

Kepler Planet Detection Metrics: Statistical Bootstrap Test

This document describes the data produced by the Statistical Bootstrap Test over the final three Threshold Crossing Event (TCE) deliveries to NExScI: SOC 9.1 (Q1Q16)1 (Tenenbaum et al. 2014), SOC 9.2 (Q1Q17) aka DR242 (Seader et al. 2015), and SOC 9.3 (Q1Q17) aka DR253 (Twicken et al. 2016). The last few years have seen significant improvements in the SOC science data processing pipeline, leading to higher quality light curves and more sensitive transit searches. The statistical bootstrap analysis results presented here and the numerical results archived at NASAs Exoplanet Science Institute (NExScI) bear witness to these software improvements. This document attempts to introduce and describe the main features and differences between these three data sets as a consequence of the software changes.

Bootstrap

Testing of the Apollo 15 Metric Camera System.

Description of tests conducted (1) to assess the quality of Apollo 15 Metric Camera System data and (2) to develop production procedures for total block reduction. Three strips of metric photography over the Hadley Rille area were selected for the tests. These photographs were utilized in a series of evaluation tests culminating in an orbitally constrained block triangulation solution. Results show that film deformations up to 25 and 5 microns are present in the mapping and stellar materials, respectively. Stellar reductions can provide mapping camera orientations with an accuracy that is consistent with the accuracies of other parameters in the triangulation solutions. Pointing accuracies of 4 to 10 microns can be expected for the mapping camera materials, depending on variations in resolution caused by changing sun angle conditions.

Helmering, R. J.

An investigation of fighter aircraft agility

This report attempts to unify in a single document the results of a series of studies on fighter aircraft agility funded by the NASA Ames Research Center, Dryden Flight Research Facility and conducted at the University of Kansas Flight Research Laboratory during the period January 1989 through December 1993. New metrics proposed by pilots and the research community to assess fighter aircraft agility are collected and analyzed. The report develops a framework for understanding the context into which the various proposed fighter agility metrics fit in terms of application and testing. Since new metrics continue to be proposed, this report does not claim to contain every proposed fighter agility metric. Flight test procedures, test constraints, and related criteria are developed. Instrumentation required to quantify agility via flight test is considered, as is the sensitivity of the candidate metrics to deviations from nominal pilot command inputs, which is studied in detail. Instead of supplying specific, detailed conclusions about the relevance or utility of one candidate metric versus another, the authors have attempted to provide sufficient data and analyses for readers to formulate their own conclusions. Readers are therefore ultimately responsible for judging exactly which metrics are 'best' for their particular needs. Additionally, it is not the intent of the authors to suggest combat tactics or other actual operational uses of the results and data in this report. This has been left up to the user community. Twenty of the candidate agility metrics were selected for evaluation with high fidelity, nonlinear, non real-time flight simulation computer programs of the F-5A Freedom Fighter, F-16A Fighting Falcon, F-18A Hornet, and X-29A. The information and data presented on the 20 candidate metrics which were evaluated will assist interested readers in conducting their own extensive investigations. The report provides a definition and analysis of each metric; details of how to test and measure the metric, including any special data reduction requirements; typical values for the metric obtained using one or more aircraft types; and a sensitivity analysis if applicable. The report is organized as follows. The first chapter in the report presents a historical review of air combat trends which demonstrate the need for agility metrics in assessing the combat performance of fighter aircraft in a modern, all-aspect missile environment. The second chapter presents a framework for classifying each candidate metric according to time scale (transient, functional, instantaneous), further subdivided by axis (pitch, lateral, axial). The report is then broadly divided into two parts, with the transient agility metrics (pitch lateral, axial) covered in chapters three, four, and five, and the functional agility metrics covered in chapter six. Conclusions, recommendations, and an extensive reference list and biography are also included. Five appendices contain a comprehensive list of the definitions of all the candidate metrics; a description of the aircraft models and flight simulation programs used for testing the metrics; several relations and concepts which are fundamental to the study of lateral agility; an in-depth analysis of the axial agility metrics; and a derivation of the relations for the instantaneous agility and their approximations.

Valasek, John

Evaluation of Two Crew Module Boilerplate Tests Using Newly Developed Calibration Metrics

The paper discusses a application of multi-dimensional calibration metrics to evaluate pressure data from water drop tests of the Max Launch Abort System (MLAS) crew module boilerplate. Specifically, three metrics are discussed: 1) a metric to assess the probability of enveloping the measured data with the model, 2) a multi-dimensional orthogonality metric to assess model adequacy between test and analysis, and 3) a prediction error metric to conduct sensor placement to minimize pressure prediction errors. Data from similar (nearly repeated) capsule drop tests shows significant variability in the measured pressure responses. When compared to expected variability using model predictions, it is demonstrated that the measured variability cannot be explained by the model under the current uncertainty assumptions.

Horta, Lucas G.

Coverage Metrics for Requirements-Based Testing: Evaluation of Effectiveness

In black-box testing, the tester creates a set of tests to exercise a system under test without regard to the internal structure of the system. Generally, no objective metric is used to measure the adequacy of black-box tests. In recent work, we have proposed three requirements coverage metrics, allowing testers to objectively measure the adequacy of a black-box test suite with respect to a set of requirements formalized as Linear Temporal Logic (LTL) properties. In this report, we evaluate the effectiveness of these coverage metrics with respect to fault finding. Specifically, we conduct an empirical study to investigate two questions: (1) do test suites satisfying a requirements coverage metric provide better fault finding than randomly generated test suites of approximately the same size?, and (2) do test suites satisfying a more rigorous requirements coverage metric provide better fault finding than test suites satisfying a less rigorous requirements coverage metric? Our results indicate (1) only one coverage metric proposed -- Unique First Cause (UFC) coverage -- is sufficiently rigorous to ensure test suites satisfying the metric outperform randomly generated test suites of similar size and (2) that test suites satisfying more rigorous coverage metrics provide better fault finding than test suites satisfying less rigorous coverage metrics.

Staats, Matt

Crew Exploration Vehicle (CEV) (Orion) Occupant Protection

The purpose of this study was to determine the similarity between the response of the THUMS model and the Hybrid III Anthropometric Test Device (ATD) given existing Wright-Patterson (WP) sled tests. There were four tests selected for this comparison with frontal, spinal, rear, and lateral loading. The THUMS was placed in a sled configuration that replicated the WP configuration and the recorded seat acceleration for each test was applied to model seat. Once the modeling simulations were complete, they were compared to the WP results using two methods. The first was a visual inspection of the sled test videos compared to the THUMS d3plot files. This comparison resulted in an assessment of the overall kinematics of the two results. The other comparison was a comparison of the plotted data recorded for both tests. The metrics selected for comparison were seat acceleration, belt forces, head acceleration and chest acceleration. These metrics were recorded in all WP tests and were outputs of the THUMS model. Once the comparison of the THUMS to the WP tests was complete, the THUMS model output was also examined for possible injuries in these scenarios. These outputs included metrics for injury risk to the head, neck, thorax, lumbar spine and lower extremities. The metrics to evaluate head response were peak head acceleration, HIC15, and HIC36. For the neck, N (sub ij) was calculated. The thorax response was evaluated with peak chest acceleration, the Combined Thoracic Index (CTI), sternal deflection, chest deflection, and chest acceleration- 3 ms clip. The lumbar spine response was evaluated with lumbar spine force. Finally the lower extremity response was evaluated by femur and tibia force. The results of the simulation comparisons indicate the THUMS model had a similar response to the Hybrid III dummy given the same input. The primary difference seen between the two was a more flexible response of the THUMS compared to the Hybrid III. This flexibility was most pronounced in the neck flexion, shoulder deflection and chest deflection. Due to the flexibility of the THUMS, the resulting head and chest accelerations tended to lag the Hybrid III acceleration trace and have a lower peak value. The results of the injury metric comparison identified possible injury trends between simulations. Risk of head injury was highest for the lateral simulations. The risk of chest injury was highest for the rear impact. However, neck injury risk was approximately the same for all simulations. The injury metric value for lumbar spine force was highest for the spinal impact. The leg forces were highest for the rear and lateral impacts. The results of this comparison indicate the THUMS model performs in a similar manner as the Hybrid III ATD. The differences in the responses of model and the ATD are primarily due to the flexibility of the THUMS. This flexibility of the THUMS would be a more human like response. Based on the similarity between the two models, the THUMS should be used in further testing to assess risk of injury to the occupant.

Currie-Gregg, Nancy J.

Using Trajectory Smoothness Metrics to Identify Drones in Radar Track Data

The identification of unmanned aircraft systems (UAS) using trajectory data is considered. Specifically, a number of smoothness metrics are proposed, which can be used to distinguish UAS from other aerial objects even when they are engaged in accelerative maneuvers (non-constant-velocity flight). The metrics are evaluated on a data set from a UAS sense-and-avoid field test, which contains track data of aerial objects recorded by a vehicle-board radar system during a flight test. The metrics are found to effectively differentiate UAS from other objects such as birds for this data set. In addition, an initial statistical performance analysis of one of the smoothness metrics is undertaken, using 15 data sets deriving from multiple flight tests. The smoothness metric is shown to identify the target UAS with 95% accuracy (95% true positive rate), while achieving a false positive rate of less than 9%.

Sandip Roy

Dynamics and control of multipayload platforms - The Middeck Active Control Experiment (MACE)

A flight experiment entitled the Middeck Active Control Experiment (MACE) proposed by the Space Engineering Research Center (SERC) at the Massachusetts Institute of Technology is described. The objective of this program is to investigate and validate the modeling of the dynamics of an actively controlled flexible, articulating, multibody platform free floating in zero gravity. A rationale and experimental approach for the program are presented. The rationale shows that on-orbit testing, coupled with ground testing and a strong analytical program, is necessary in order to fully understand both how flexibility of the platform affects the pointing problem, as well as how gravity perturbs this structural flexibility causing deviations between 1-and 0-gravity behavior. The experimental approach captures the essential physics of multibody platforms, by identifying the appropriate attributes, tests, and performance metrics of the test article, and defines the tests required to successfully validate the analytical framework.

Miller, David W.

The MODE family of on-orbit experiments: The Middeck Active Control Experiment (MACE)

A flight experiment entitled the Middeck Active Control Experiment (MACE), proposed by the Space Engineering Research Center (SERC) at the Massachusetts Institute of Technology, is described. This is the second in a family of flight experiments being developed at MIT. The first is the Middeck 0-Gravity Dynamics Experiment (MODE) which investigates the nonlinear behavior of contained fluids and truss structures in zero gravity. The objective of the MACE program is to investigate and validate the modeling of the dynamics of an actively controlled flexible, articulating, multibody platform free floating in zero gravity. A rationale and experimental approach for the program are presented. The rationale shows that on-orbit testing, coupled with ground testing and a strong analytical program, is necessary in order to fully understand both how flexibility of the platform affects the pointing problem, as well as how gravity perturbs this structural flexibility causing deviations between 1- and 0-gravity behavior. The experimental approach captures the essential physics of multibody platforms, by identifying the appropriate attributes, tests, and performance metrics of the test article and defines the tests required to successfully validate the analytical framework.

Crawley, Edward F.

JPL/NASA/IEEE Test Effectiveness Workshop

(none given)From OBJECTIVES: Specific objectives of the working group are to support the innovation, development, evaluation and implementation of test methods, metrics and tools based on failure engineering/physics and/or root cause evaluations. Data sources systems and tools shall be developed and implemented that: 1) provide improved preventions, controls, analyses and tests (PACT) & field failure data collection, 2) facilities data analysis, archiving, retrieval, failure physics and/or root cause evaluations and 3) enable new and existing technology suitability evaluations to be performed.

effectiveness concurrent engineering metrics test