Search NASA⌕ Search

SEARCH · Search NASA

Results for “false positive analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Reducing False Positives in Runtime Analysis of Deadlocks

This paper presents an improvement of a standard algorithm for detecting dead-lock potentials in multi-threaded programs, in that it reduces the number of false positives. The standard algorithm works as follows. The multi-threaded program under observation is executed, while lock and unlock events are observed. A graph of locks is built, with edges between locks symbolizing locking orders. Any cycle in the graph signifies a potential for a deadlock. The typical standard example is the group of dining philosophers sharing forks. The algorithm is interesting because it can catch deadlock potentials even though no deadlocks occur in the examined trace, and at the same time it scales very well in contrast t o more formal approaches to deadlock detection. The algorithm, however, can yield false positives (as well as false negatives). The extension of the algorithm described in this paper reduces the amount of false positives for three particular cases: when a gate lock protects a cycle, when a single thread introduces a cycle, and when the code segments in different threads that cause the cycle can actually not execute in parallel. The paper formalizes a theory for dynamic deadlock detection and compares it to model checking and static analysis techniques. It furthermore describes an implementation for analyzing Java programs and its application to two case studies: a planetary rover and a space craft altitude control system.

Bensalem, Saddek↗

Abort Trigger False Positive and False Negative Analysis Methodology for Threshold-Based Abort Detection

This paper describes a quantitative methodology for bounding the false positive (FP) and false negative (FN) probabilities associated with a human-rated launch vehicle abort trigger (AT) that includes sensor data qualification (SDQ). In this context, an AT is a hardware and software mechanism designed to detect the existence of a specific abort condition. Also, SDQ is an algorithmic approach used to identify sensor data suspected of being corrupt so that suspect data does not adversely affect an AT's detection capability. The FP and FN methodologies presented here were developed to support estimation of the probabilities of loss of crew and loss of mission for the Space Launch System (SLS) which is being developed by the National Aeronautics and Space Administration (NASA). The paper provides a brief overview of system health management as being an extension of control theory; and describes how ATs and the calculation of FP and FN probabilities relate to this theory. The discussion leads to a detailed presentation of the FP and FN methodology and an example showing how the FP and FN calculations are performed. This detailed presentation includes a methodology for calculating the change in FP and FN probabilities that result from including SDQ in the AT architecture. To avoid proprietary and sensitive data issues, the example incorporates a mixture of open literature and fictitious reliability data. Results presented in the paper demonstrate the effectiveness of the approach in providing quantitative estimates that bound the probability of a FP or FN abort determination.

risk assessment↗

SafeAeroBERT: Towards a Safety-Informed Aerospace-Specific Language Model

As aviation systems continue to operate with high traffic, large amounts of documents containing safety-relevant data continue to be generated via reporting systems such as the ASRS. Advanced natural language processing techniques, specifically pre-trained language models, have shown great success in domain-specific applications; however, the text in aviation safety reports is inundated with jargon and thus not fully utilized by general pre-trained models. In this research, we work towards developing a safety-informed aerospace-specific language model by pre-training a Bidirectional Encoder Representations from Transformer (BERT) model on reports from the Aviation Safety Reporting System and the National Transportation Safety Board. The resulting model, called SafeAeroBERT, is fine-tuned for the specific task of document classification, and can be further tuned for named-entity recognition, relation detection, information retrieval, and summarization. Results from the classification task are compared between SafeAeroBERT, the base BERT, and SciBERT models and show SafeAeroBERT outperforms the general BERT and SciBERT on classifying reports about human factors, aircraft, and procedure. SafeAeroBERT can be used on custom tasks, not limited to document classification, and is intended to aid an intelligent knowledge manager for safety report repositories.

Aviation↗

Assessing Reliability of NDE Flaw Detection Using Smaller Number of Demonstration Data Points

The paper provides an engineering analysis approach for assessing reliability of NDE flaw detection using smaller number of demonstration data points. It explores dependence of probability of detection (POD), probability of false positive (POF), on contrast-to-noise ratio, and net decision threshold-to-noise ratio in a simulated data; and draws some generically applicable inferences to devise the approach. ASTM nondestructive evaluation standards provide requirements on signal-to-noise ratio and/or contrast-to-noise ratio in order to provide reliable flaw detection and limit false positive calls. POD analysis of inspection test data results in an estimated flaw size, denoted by 𝑎90/95. This flaw size has 90% POD and minimum 95% confidence. POF is also estimated in the analysis. POD demonstration requires specimens with flaws of known size. In many situations, it is very expensive to produce the large number of flaws required for the POD analysis. In some situations, only real flaws can truly represent the flaws for demonstration. Real flaws of correct size and location in part configuration specimen may be difficult to produce, if not impossible. Here, an engineering analysis approach is devised using simulation to assess reliability of NDE technique when a limited number of flaws are available for demonstration. In this simulation, a technique is considered reliable, if it provides flaw detectability size equal to or better than the theoretical 𝑎90𝑡ℎ used in simulation and also provides a POF less than or equal to a chosen value. The paper uses simulated signal response versus flaw size data to devise the approach. Linear correlation is used between the signal response data and flaw size. POD software mh1823 uses generalized linear model (GLM) in POD analysis after transforming the flaw size and signal response, if needed, using logarithm. Therefore, this approach is in agreement with the linear signal correlation used in mh1823. Using the POD analysis of data, generic conditions on contrast-to-noise ratio and net decision threshold-to-noise ratio are derived for reliable flaw detection. In order to assess technique reliability using the engineering approach, signal response-to-flaw size correlation about the flaw size of concern is needed. In addition, measurement of noise is also needed. If the technique meets the above requirements, assumption of linear signal-to-flaw size correlation and conditions on noise, then the technique can be assessed using this analysis as it fits the underlying POD model used here. The approach is conservative and is designed to provide a larger flaw size compared to the POD approach. Such NDE technique assessment approach, although, not as rigorous as POD, can be cost effective if the larger flaw size can be tolerated. Typically, this is a situation for all quality control NDE inspections. Here, an NDE technique needs to be reliable and 𝑎90/95 is not estimated, but the assessed flaw size is assumed to be larger than the unknown a90 due to conservative factors or margins. Applicability of the approach for assessing reliability of flaw detection in x-ray radiography and 2D imaging in general is also explored.

Koshti, Ajay M.↗

Applying Jlint to Space Exploration Software

Java is a very successful programming language which is also becoming widespread in embedded systems, where software correctness is critical. Jlint is a simple but highly efficient static analyzer that checks a Java program for several common errors, such as null pointer exceptions, and overflow errors. It also includes checks for multi-threading problems, such as deadlocks and data races. The case study described here shows the effectiveness of Jlint in find-false positives in the multi-threading warnings gives an insight into design patterns commonly used in multi-threaded code. The results show that a few analysis techniques are sufficient to avoid almost all false positives. These techniques include investigating all possible callers and a few code idioms. Verifying the correct application of these patterns is still crucial, because their correct usage is not trivial.

Artho, Cyrille↗

Reducing V&V Cost of Flight Critical Systems: Myth or Reality?

This paper presents an overview of NASA research program on the V&V of flight critical systems. Five years ago, NASA started an effort to reduce the cost and possibly increase the effectiveness of V&V for flight critical systems. It is the right time to take a look back and realize what progress has been made. This paper describes our overall approach and the tools introduced to address different phases of the software lifecycle. For example, we have improved testing by developing a statistical learning approach tor defining test cases. The tool automatically identifies possible unsafe conditions by analyzing outliers in output data; using an iterative learning process, it can then generate more test cases that represent potentially unsafe regions of operation. At the code level, we have developed and made available as open source a static analyzer for C and C++ programs called IKOS. We have shown that IKOS is very precise in the analysis of embedded C programs (very few false positives) and a bit less for regular C and C++ code. At the design level, in collaboration with our NRA partners, we have developed a suite of analysis tools for Simulink models. The analysis is done in a compositional framework for scalability.

Brat, Guillaume P.↗

Investigating the Use of Machine Learning (ML) to Assess Tropospheric Doppler Radar Wind Profiler (TDRWP) Data Quality

Manual Quality Control (MQC) of Tropospheric Doppler Radar Wind Profiler (TDRWP) data is essential for defining an accurate climatology for downstream aerospace vehicle assessments. MQC traditionally takes around 30.5 hours per year of radar data. The Marshall Space Flight Center Natural Environments Branch (MSFC NE) used machine learning (ML) to test the feasibility of automating the MQC process, showing a potential to reduce labor by 300%. However, analysis of the model showed some false positives. We compared a neural network to the model to validate it and develop a process for assessing comparable solutions in the future.

Corey Walker↗

Multi-Stage System for Automatic Target Recognition

A multi-stage automated target recognition (ATR) system has been designed to perform computer vision tasks with adequate proficiency in mimicking human vision. The system is able to detect, identify, and track targets of interest. Potential regions of interest (ROIs) are first identified by the detection stage using an Optimum Trade-off Maximum Average Correlation Height (OT-MACH) filter combined with a wavelet transform. False positives are then eliminated by the verification stage using feature extraction methods in conjunction with neural networks. Feature extraction transforms the ROIs using filtering and binning algorithms to create feature vectors. A feedforward back-propagation neural network (NN) is then trained to classify each feature vector and to remove false positives. The system parameter optimizations process has been developed to adapt to various targets and datasets. The objective was to design an efficient computer vision system that can learn to detect multiple targets in large images with unknown backgrounds. Because the target size is small relative to the image size in this problem, there are many regions of the image that could potentially contain the target. A cursory analysis of every region can be computationally efficient, but may yield too many false positives. On the other hand, a detailed analysis of every region can yield better results, but may be computationally inefficient. The multi-stage ATR system was designed to achieve an optimal balance between accuracy and computational efficiency by incorporating both models. The detection stage first identifies potential ROIs where the target may be present by performing a fast Fourier domain OT-MACH filter-based correlation. Because threshold for this stage is chosen with the goal of detecting all true positives, a number of false positives are also detected as ROIs. The verification stage then transforms the regions of interest into feature space, and eliminates false positives using an artificial neural network classifier. The multi-stage system allows tuning the detection sensitivity and the identification specificity individually in each stage. It is easier to achieve optimized ATR operation based on its specific goal. The test results show that the system was successful in substantially reducing the false positive rate when tested on a sonar and video image datasets.

Chao, Tien-Hsin↗

Ares I-X Ground Diagnostic Prototype

Automating prelaunch diagnostics for launch vehicles offers three potential benefits. First, it potentially improves safety by detecting faults that might otherwise have been missed so that they can be corrected before launch. Second, it potentially reduces launch delays by more quickly diagnosing the cause of anomalies that occur during prelaunch processing. Reducing launch delays will be critical to the success of NASA's planned future missions that require in-orbit rendezvous. Third, it potentially reduces costs by reducing both launch delays and the number of people needed to monitor the prelaunch process. NASA is currently developing the Ares I launch vehicle to bring the Orion capsule and its crew of four astronauts to low-earth orbit on their way to the moon. Ares I-X will be the first unmanned test flight of Ares I. It is scheduled to launch on October 27, 2009. The Ares I-X Ground Diagnostic Prototype is a prototype ground diagnostic system that will provide anomaly detection, fault detection, fault isolation, and diagnostics for the Ares I-X first-stage thrust vector control (TVC) and for the associated ground hydraulics while it is in the Vehicle Assembly Building (VAB) at John F. Kennedy Space Center (KSC) and on the launch pad. It will serve as a prototype for a future operational ground diagnostic system for Ares I. The prototype combines three existing diagnostic tools. The first tool, TEAMS (Testability Engineering and Maintenance System), is a model-based tool that is commercially produced by Qualtech Systems, Inc. It uses a qualitative model of failure propagation to perform fault isolation and diagnostics. We adapted an existing TEAMS model of the TVC to use for diagnostics and developed a TEAMS model of the ground hydraulics. The second tool, Spacecraft Health Inference Engine (SHINE), is a rule-based expert system developed at the NASA Jet Propulsion Laboratory. We developed SHINE rules for fault detection and mode identification. The prototype uses the outputs of SHINE as inputs to TEAMS. The third tool, the Inductive Monitoring System (IMS), is an anomaly detection tool developed at NASA Ames Research Center and is currently used to monitor the International Space Station Control Moment Gyroscopes. IMS automatically "learns" a model of historical nominal data in the form of a set of clusters and signals an alarm when new data fails to match this model. IMS offers the potential to detect faults that have not been modeled. The three tools have been integrated and deployed to Hangar AE at KSC where they interface with live data from the Ares I-X vehicle and from the ground hydraulics. The outputs of the tools are displayed on a console in Hangar AE, one of the locations from which the Ares I-X launch will be monitored. The full paper will describe how the prototype performed before the launch. It will include an analysis of the prototype's accuracy, including false-positive rates, false-negative rates, and receiver operating characteristics (ROC) curves. It will also include a description of the prototype's computational requirements, including CPU usage, main memory usage, and disk usage. If the prototype detects any faults during the prelaunch period then the paper will include a description of those faults. Similarly, if the prototype has any false alarms then the paper will describe them and will attempt to explain their causes.

Schwabacher, Mark↗

Evaluating Alerting and Guidance Performance of a UAS Detect-And-Avoid System

A key challenge to the routine, safe operation of unmanned aircraft systems (UAS) is the development of detect-and-avoid (DAA) systems to aid the UAS pilot in remaining "well clear" of nearby aircraft. The goal of this study is to investigate the effect of alerting criteria and pilot response delay on the safety and performance of UAS DAA systems in the context of routine civil UAS operations in the National Airspace System (NAS). A NAS-wide fast-time simulation study was conducted to assess UAS DAA system performance with a large number of encounters and a broad set of DAA alerting and guidance system parameters. Three attributes of the DAA system were controlled as independent variables in the study to conduct trade-off analyses: UAS trajectory prediction method (dead-reckoning vs. intent-based), alerting time threshold (related to predicted time to LoWC), and alerting distance threshold (related to predicted Horizontal Miss Distance, or HMD). A set of metrics, such as the percentage of true positive, false positive, and missed alerts, based on signal detection theory and analysis methods utilizing the Receiver Operating Characteristic (ROC) curves were proposed to evaluate the safety and performance of DAA alerting and guidance systems and aid development of DAA system performance standards. The effect of pilot response delay on the performance of DAA systems was evaluated using a DAA alerting and guidance model and a pilot model developed to support this study. A total of 18 fast-time simulations were conducted with nine different DAA alerting threshold settings and two different trajectory prediction methods, using recorded radar traffic from current Visual Flight Rules (VFR) operations, and supplemented with DAA-equipped UAS traffic based on mission profiles modeling future UAS operations. Results indicate DAA alerting distance threshold has a greater effect on DAA system performance than DAA alerting time threshold or ownship trajectory prediction method. Further analysis on the alert lead time (time in advance of predicted loss of well clear at which a DAA alert is first issued) indicated a strong positive correlation between alert lead time and DAA system performance (i.e. the ability of the UAS pilot to maneuver the unmanned aircraft to remain well clear). While bigger distance thresholds had beneficial effects on alert lead time and missed alert rate, it also generated a higher rate of false alerts. In the design and development of DAA alerting and guidance systems, therefore, the positive and negative effects of false alerts and missed alerts should be carefully considered to achieve acceptable alerting system performance by balancing false and missed alerts. The results and methodology presented in this study are expected to help stakeholders, policymakers and standards committees define the appropriate setting of DAA system parameter thresholds for UAS that ensure safety while minimizing operational impacts to the NAS and equipage requirements for its users before DAA operational performance standards can be finalized.

UAS Detect-and-Avoid (DAA) System↗

Comparison of epifluorescent viable bacterial count methods

Two methods, the 2-(4-Iodophenyl) 3-(4-nitrophenyl) 5-phenyltetrazolium chloride (INT) method and the direct viable count (DVC), were tested and compared for their efficiency for the determination of the viability of bacterial populations. Use of the INT method results in the formation of a dark spot within each respiring cell. The DVC method results in elongation or swelling of growing cells that are rendered incapable of cell division. Although both methods are subjective and can result in false positive results, the DVC method is best suited to analysis of waters in which the number of different types of organisms present in the same sample is assumed to be small, such as processed waters. The advantages and disadvantages of each method are discussed.

Rodgers, E. B.↗

Discovery and Vetting of Exoplanets. I. Benchmarking K2 Vetting Tools

We have adapted the algorithmic tools developed during the Kepler mission to vet the quality of transit-like signals for use on the K2 mission data. Using the four sets of publicly available light curves at MAST, we produced a uniformly vetted catalog of 772 transiting planet candidates from K2 as listed at the NASA Exoplanet Archive in the K2 Table of Candidates. Our analysis marks 676 of these as planet candidates and 96 as false positives. All confirmed planets pass our vetting tests. Sixty of our false positives are new identifications, effectively doubling the overall number of astrophysical signals mimicking planetary transits in K2 data. Most of the targets listed as false positives in our catalog show either prominent secondary eclipses, transit depths suggesting a stellar companion instead of a planet, or significant photocenter shifts during transit. We packaged our tools into the open-source, automated vetting pipeline Discovery and Vetting of Exoplanets (DAVE), designed to streamline follow-up efforts by reducing the time and resources wasted observing targets that are likely false positives. DAVE will also be a valuable tool for analyzing planet candidates from NASA's TESS mission, where several guest-investigator programs will provide independent light-curve sets—and likely many more from the community. We are currently testing DAVE on recently released TESS planet candidates and will present our results in a follow-up paper.

Kostov, Veselin B.↗

Human versus automation in responding to failures: an expected-value analysis

A simple analytical criterion is provided for deciding whether a human or automation is best for a failure detection task. The method is based on expected-value decision theory in much the same way as is signal detection. It requires specification of the probabilities of misses (false negatives) and false alarms (false positives) for both human and automation being considered, as well as factors independent of the choice--namely, costs and benefits of incorrect and correct decisions as well as the prior probability of failure. The method can also serve as a basis for comparing different modes of automation. Some limiting cases of application are discussed, as are some decision criteria other than expected value. Actual or potential applications include the design and evaluation of any system in which either humans or automation are being considered.

NASA Discipline Space Human Factors↗

Optimization of Second Fault Detection Thresholds to Maximize Mission Probability of Success

In order to support manned spaceflight safety requirements, the Space Launch System (SLS) has defined program-level requirements for key systems to ensure successful operation under single fault conditions. The SLS program has also levied requirements relating to the capability of the Inertial Navigation System to detect a second fault. This detection functionality is required in order to feed abort analysis and ensure crew safety. Increases in navigation state error due to sensor faults in a purely inertial system can drive the vehicle outside of its operational as-designed environmental and performance envelope. As this performance outside of first fault detections is defined and controlled at the vehicle level, it allows for the use of system level margins to increase probability of mission success on the operational edges of the design. A top-down approach is utilized to assess vehicle sensitivity to second sensor faults. A wide range of failure scenarios in terms of both fault magnitude and time is used for assessment. The approach also utilizes a schedule to change fault detection thresholds autonomously. These individual values are optimized along a nominal trajectory in order to maximize probability of mission success in terms of system-level insertion requirements while minimizing the probability of false positives. This paper will describe an approach integrating Genetic Algorithms and Monte Carlo analysis to tune the threshold parameters to maximize vehicle resilience to second fault events over an ascent mission profile. The analysis approach and performance assessment and verification will be presented to demonstrate the applicability of this approach to second fault detection optimization to maximize mission probability of success through taking advantage of existing margin.

Anzalone, Evan J.↗

Launch Vehicle Failure Dynamics and Abort Triggering Analysis

Launch vehicle ascent is a time of high risk for an on-board crew. There are many types of failures that can kill the crew if the crew is still on-board when the failure becomes catastrophic. For some failure scenarios, there is plenty of time for the crew to be warned and to depart, whereas in some there is insufficient time for the crew to escape. There is a large fraction of possible failures for which time is of the essence and a successful abort is possible if the detection and action happens quickly enough. This paper focuses on abort determination based primarily on data already available from the GN&C system. This work is the result of failure analysis efforts performed during the Ares I launch vehicle development program. Derivation of attitude and attitude rate abort triggers to ensure that abort occurs as quickly as possible when needed, but that false positives are avoided, forms a major portion of the paper. Some of the potential failure modes requiring use of these triggers are described, along with analysis used to determine the success rate of getting the crew off prior to vehicle demise.

Hanson, John M.↗

A Discovery of a Candidate Companion to a Transiting System KOI-94: A Direct Imaging Study for a Possibility of a False Positive

We report a discovery of a companion candidate around one of Kepler Objects of Interest (KOIs), KOI-94, and results of our quantitative investigation of the possibility that planetary candidates around KOI-94 are false positives. KOI-94 has a planetary system in which four planetary detections have been reported by Kepler, suggesting that this system is intriguing to study the dynamical evolutions of planets. However, while two of those detections (KOI-94.01 and 03) have been made robust by previous observations, the others (KOI-94.02 and 04) are marginal detections, for which future confirmations with various techniques are required. We have conducted high-contrast direct imaging observations with Subaru/HiCIAO in H band and detected a faint object located at a separation of approximately 0.6 sec from KOI-94. The object has a contrast of approximately 1 × 10(exp −3) in H band, and corresponds to an M type star on the assumption that the object is at the same distance of KOI-94. Based on our analysis, KOI-94.02 is likely to be a real planet because of its transit depth, while KOI-94.04 can be a false positive due to the companion candidate. The success in detecting the companion candidate suggests that high-contrast direct imaging observations are important keys to examine false positives of KOIs. On the other hand, our transit light curve reanalyses lead to a better period estimate of KOI-94.04 than that on the KOI catalogue and show that the planetary candidate has the same limb darkening parameter value as the other planetary candidates in the KOI-94 system, suggesting that KOI-94.04 is also a real planet in the system.

Subaru/HiCIAO↗