Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Analyzing and Predicting Effort Associated with Finding and Fixing Software Faults

Context: Software developers spend a significant amount of time fixing faults. However, not many papers have addressed the actual effort needed to fix software faults. Objective: The objective of this paper is twofold: (1) analysis of the effort needed to fix software faults and how it was affected by several factors and (2) prediction of the level of fix implementation effort based on the information provided in software change requests. Method: The work is based on data related to 1200 failures, extracted from the change tracking system of a large NASA mission. The analysis includes descriptive and inferential statistics. Predictions are made using three supervised machine learning algorithms and three sampling techniques aimed at addressing the imbalanced data problem. Results: Our results show that (1) 83% of the total fix implementation effort was associated with only 20% of failures. (2) Both safety critical failures and post-release failures required three times more effort to fix compared to non-critical and pre-release counterparts, respectively. (3) Failures with fixes spread across multiple components or across multiple types of software artifacts required more effort. The spread across artifacts was more costly than spread across components. (4) Surprisingly, some types of faults associated with later life-cycle activities did not require significant effort. (5) The level of fix implementation effort was predicted with 73% overall accuracy using the original, imbalanced data. Using oversampling techniques improved the overall accuracy up to 77%. More importantly, oversampling significantly improved the prediction of the high level effort, from 31% to around 85%. Conclusions: This paper shows the importance of tying software failures to changes made to fix all associated faults, in one or more software components and/or in one or more software artifacts, and the benefit of studying how the spread of faults and other factors affect the fix implementation effort.

software fix implementation effort↗

A hierarchical approach to reliability modeling of fault-tolerant systems

A methodology for performing fault tolerant system reliability analysis is presented. The method decomposes a system into its subsystems, evaluates vent rates derived from the subsystem's conditional state probability vector and incorporates those results into a hierarchical Markov model of the system. This is done in a manner that addresses failure sequence dependence associated with the system's redundancy management strategy. The method is derived for application to a specific system definition. Results are presented that compare the hierarchical model's unreliability prediction to that of a more complicated tandard Markov model of the system. The results for the example given indicate that the hierarchical method predicts system unreliability to a desirable level of accuracy while achieving significant computational savings relative to component level Markov model of the system.

Gossman, W. E.↗

Historical seismicity near Chagos - A complex deformation zone in the equatorial Indian Ocean

The historical seismicity of the Chagos region of the Indian Ocean is analyzed, using earthquake relocation methods and a moment variance technique to determine the focal mechanisms of quakes occurring before 1964. Moment variance analysis showed a thrust faulting mechanism associated with the earthquake of 1944 near the Chagos-Laccadive Ridge; a strike-slip mechanism was associated with a smaller 1957 event occurring west of the Chagos Bank. The location of the 1944 event, one of the largest intraplate earthquakes known (1.4 x 10 to the 27th dyne/cm), would imply that the Chagos seismicity is due to a zone of tectonic deformation stretching across the equatorial Indian Ocean. The possibility of a slow diffuse boundary extending west of the Central Indian Ridge is also discussed. This boundary is confirmed by recent plate motion studies which suggest that it separates the Australian plate from a single Indo-Arabian plate.

Wiens, D. A.↗

Application of expert systems in the Common Module electrical power system

Each Common Module (CM) of the Space Station must be capable of handling a 50 kW electricity supply, 25 kW for transmission and 25 kW for consumption. The power must be handled and managed by on-board systems, a necessity that dovetails with the objectives of Public Law 98-371, which mandates that the Space Station push the state of the art of automation and AI. Expert systems will be needed to handle the large data flow for the power system and to ensure that the system degrades gracefully. Features of the first expert systems expected for the power system, i.e., a dynamic load planner/scheduler and energy storage subsystem management, fault diagnosis/analysis, health status/trend analysis, and orbital replacement advisor expert systems, are described. Finally, growth Space Station expert systems applications are discussed.

Weeks, D. J.↗

PI-in-a-box: Intelligent onboard assistance for spaceborne experiments in vestibular physiology

In construction is a knowledge-based system that will aid astronauts in the performance of vestibular experiments in two ways: it will provide real-time monitoring and control of signals and it will optimize the quality of the data obtained, by helping the mission specialists and payload specialists make decisions that are normally the province of a principal investigator, hence the name PI-in-a-box. An important and desirable side-effect of this tool will be to make the astronauts more productive and better integrated members of the scientific team. The vestibular experiments are planned by Prof. Larry Young of MIT, whose team has already performed similar experiments in Spacelab missions SL-1 and D-1, and has experiments planned for SLS-1 and SLS-2. The knowledge-based system development work, performed in collaboration with MIT, Stanford University, and the NASA-Ames Research Center, addresses six major related functions: (1) signal quality monitoring; (2) fault diagnosis; (3) signal analysis; (4) interesting-case detection; (5) experiment replanning; and (6) integration of all of these functions within a real-time data acquisition environment. Initial prototyping work has been done in functions (1) through (4).

Colombano, Silvano↗

Test chips and ASIC qualification

A test chip set being developed to aid in the qualification of spaceborne Application Specific Integrated Circuits (ASICs) is described. The chip set consists of a process monitor for process parameter verification, a fault chip for yield analysis, a reliability chip for ASIC failure rate analysis, and total ionizing dose and single event upset chips for radiation effect analysis. The test structures contained in these chips are discussed along with representative test results.

Buehler, M. G.↗

Semi-Markov Unreliability Range Evaluator

Semi-Markov Unreliability Range Evaluator, SURE, computer program is software tool for analysis of reconfigurable, fault-tolerant systems. Traditional reliability analyses based on aggregates of fault-handling and fault-occurrence models. SURE provides efficient means for calculating accurate upper and lower bounds for probabilities of death states for large class of semi-Markov mathematical models, and not merely those reduced to critical-pair architectures.

Butler, Ricky W.↗

Automation and robotics considerations for a lunar base

An envisioned lunar outpost shares with other NASA missions many of the same criteria that have prompted the development of intelligent automation techniques with NASA. Because of increased radiation hazards, crew surface activities will probably be even more restricted than current extravehicular activity in low Earth orbit. Crew availability for routine and repetitive tasks will be at least as limited as that envisioned for the space station, particularly in the early phases of lunar development. Certain tasks are better suited to the untiring watchfulness of computers, such as the monitoring and diagnosis of multiple complex systems, and the perception and analysis of slowly developing faults in such systems. In addition, mounting costs and constrained budgets require that human resource requirements for ground control be minimized. This paper provides a glimpse of certain lunar base tasks as seen through the lens of automation and robotic (A&R) considerations. This can allow a more efficient focusing of research and development not only in A&R, but also in those technologies that will depend on A&R in the lunar environment.

Sliwa, Nancy E.↗

Design and application of electromechanical actuators for deep space missions

During the period 8/16/92 through 2/15/93, work has been focused on three major topics: (1) screw modeling and testing; (2) motor selection; and (3) health monitoring and fault diagnosis. Detailed theoretical analysis has been performed to specify a full dynamic model for the roller screw. A test stand has been designed for model parameter estimation and screw testing. In addition, the test stand is expected to be used to perform a study on transverse screw loading.

Haskew, Tim A.↗

Towards a Theory for Integration of Mathematical Verification and Empirical Testing

From the viewpoint of a project manager responsible for the V&V (verification and validation) of a software system, mathematical verification techniques provide a possibly useful orthogonal dimension to otherwise standard empirical testing. However, the value they add to an empirical testing regime both in terms of coverage and in fault detection has been difficult to quantify. Furthermore, potential cost savings from replacing testing with mathematical verification techniques cannot be realized until the tradeoffs and synergies can be formulated. Integration of formal verification with empirical testing is also difficult because the idealized view of mathematical verification providing a correctness proof with total coverage is unrealistic and does not reflect the limitations imposed by computational complexity of mathematical techniques. This paper first describes a framework based on software reliability and formalized fault models for a theory of software design fault detection - and hence the utility of various tools for debugging. It then describes a utility model for integrating mathematical and empirical techniques with respect to fault detection and coverage analysis. It then considers the optimal combination of black-box testing, white-box (structural) testing, and formal methods in V&V of a software system. Using case studies from NASA software systems, it then demonstrates how this utility model can be used in practice.

Lowry, Michael↗

IV and V Issues in Achieving High Reliability and Safety in Critical Control System Software

Risk analysis and integrated verification and validation are two important elements in a plan for ensuring the safety of critical software systems. We describe an approach we are currently developing for integrating risk analysis, and metrics analysis, and propose a fault predictor that would integrate the results of these activities. Practical difficulties associated with our approach are also discussed, as are limitations of the proposed predictor. We conclude with a discussion of what as been learned to date, and with suggestions for future work.

validation software reliability software risk asse↗

Overcoming obstacles to the exchange of information between risk tools

Our work to date in connecting risk tools hs had successes, but also has revealed there to be significant impediments to information exchange between them. These impediments stem from the well-known phenomenon of 'semantic dissonance' - mismatch between conceptual assumptions made by the separately developed tools. This issue represents a fundamental challenge that arises regardless of the mechanism of information exchange. This paper explains the issue and illustrates it with reference to our experiences to date connecting several risk tools. We motivate this work, present and discuss the solutions we have adopted to surmount these impediments, and the implications this work has for future efforts to integrate risk tools.

semantic dissonance↗

Developing Signal-Pattern-Recognition Programs

Pattern Interpretation and Recognition Application Toolkit Environment (PIRATE) is a block-oriented software system that aids the development of application programs that analyze signals in real time in order to recognize signal patterns that are indicative of conditions or events of interest. PIRATE was originally intended for use in writing application programs to recognize patterns in space-shuttle telemetry signals received at Johnson Space Center's Mission Control Center: application programs were sought to (1) monitor electric currents on shuttle ac power busses to recognize activations of specific power-consuming devices, (2) monitor various pressures and infer the states of affected systems by applying a Kalman filter to the pressure signals, (3) determine fuel-leak rates from sensor data, (4) detect faults in gyroscopes through analysis of system measurements in the frequency domain, and (5) determine drift rates in inertial measurement units by regressing measurements against time. PIRATE can also be used to develop signal-pattern-recognition software for different purposes -- for example, to monitor and control manufacturing processes.

Shelton, Robert O.↗

A Framework for Extending the Science Traceability Matrix: Application to the Planned Europa Mission

One of the most critical functions of the systems engineering requirements process for a large multi-instrument science-driven space mission is to successfully communicate customer expectations into a comprehensive and traceable science requirements flowdown. These requirements are essential to communicating the constraints on the scope of the science investigations and clarifying how multiple instruments contribute to a given science goal. They also provide insight into how the science goals of the whole mission are affected by design choices. There is little specific guidance available on best practices for developing this science-driven flowdown. A unified Science Traceability Matrix (USTM) contains a significant amount of information that can be leveraged for that purpose, but the USTM was not designed to directly produce a complete science requirements flowdown. Thus, starting with the principles codified in a USTM, the authors propose a framework that directly maps into the requirements flowdown and supports broader systems engineering processes while retaining its meaning to the science team. This Science Traceability and Alignment Framework, or STAF, defines a set of common definitions and valid relationships to structure communication across the project. In addition, STAF populates a network of information that can be useful to support complex mission analysis activities such as fault protection. This work discusses the highest-level implementation of the STAF, the project-domain or P-STAF, which describes an approach to decomposing customer requirements into science requirements. The planned Europa Mission is used as a case study for the implementation of this framework and its potential benefits to a project.

Susca, Sara↗

Lessons Learned from Recent Rapid Imaging of Earthquake Ruptures with Satellite Geodesy

Rapid determination of the location and extent of earthquake ruptures at the surface and at depth is helpful for disaster response, as it allows prediction of the likely area of major damage from the earthquake and can help with rescue and recovery planning. The Caltech-Jet Propulsion Laboratory (JPL) Advanced Rapid Imaging and Analysis (ARIA) project has responded to many recent large earthquakes to process geodetic data and help determine the location and extent of ruptures. With the increasing availability of near real-time data from the Global Positioning System (GPS) and other global navigation satellite system receivers in active tectonic regions, and with the shorter repeat times of many recent and newly launched radar satellites, geodetic data can now be obtained quickly after earthquakes or other disasters. We have been building an ARIA data system that can ingest, catalog, and process geodetic data and combine it with seismic analysis to estimate the fault rupture locations and slip distributions for large earthquakes that are on or near land.

Simons, Mark↗

NASA Space Nuclear Propulsion (SNP) MBSE Initiatives

NASA’s Space Nuclear Propulsion (SNP) program is developing several MagicDraw SysML models to support the development of high performance Nuclear Thermal Rocket Engines (NTRE). Currently, the Demonstration Rocket for Agile Cislunar Operations (DRACO) project is aiming to perform the first ever flight demonstration of an NTRE, and NASA is developing a DRACO Insight Project Model Based Systems Engineering (MBSE) model to capture, define, analyze, and report on the flight and ground test system architecture, functional behavior, requirements, risks, and lessons learned. Additional models are in work for engine component trade trees, fault detection sensor coverage analysis using a Goal Function Tree (GFT) plugin, stakeholder engagement, and technology maturation projects. The GFT plugin is the Galois, Inc. Failure Recovery Instruction Generation using Automata derived from Traditional Engineering models (FRIGATE) tool. A new capability for Jira to MagicDraw data sharing using the OpenPDM collaboration platform is under development with partner Victory Solutions, Inc. to enhance risk impact analysis.

Space Nuclear Propulsion (SNP)↗

Incipient fault detection study for advanced spacecraft systems

A feasibility study to investigate the application of vibration monitoring to the rotating machinery of planned NASA advanced spacecraft components is described. Factors investigated include: (1) special problems associated with small, high RPM machines; (2) application across multiple component types; (3) microgravity; (4) multiple fault types; (5) eight different analysis techniques including signature analysis, high frequency demodulation, cepstrum, clustering, amplitude analysis, and pattern recognition are compared; and (6) small sample statistical analysis is used to compare performance by computation of probability of detection and false alarm for an ensemble of repeated baseline and faulted tests. Both detection and classification performance are quantified. Vibration monitoring is shown to be an effective means of detecting the most important problem types for small, high RPM fans and pumps typical of those planned for the advanced spacecraft. A preliminary monitoring system design and implementation plan is presented.

Milner, G. Martin↗

Autonomous power expert system advanced development

The autonomous power expert (APEX) system is being developed at Lewis Research Center to function as a fault diagnosis advisor for a space power distribution test bed. APEX is a rule-based system capable of detecting faults and isolating the probable causes. APEX also has a justification facility to provide natural language explanations about conclusions reached during fault isolation. To help maintain the health of the power distribution system, additional capabilities were added to APEX. These capabilities allow detection and isolation of incipient faults and enable the expert system to recommend actions/procedure to correct the suspected fault conditions. New capabilities for incipient fault detection consist of storage and analysis of historical data and new user interface displays. After the cause of a fault is determined, appropriate recommended actions are selected by rule-based inferencing which provides corrective/extended test procedures. Color graphics displays and improved mouse-selectable menus were also added to provide a friendlier user interface. A discussion of APEX in general and a more detailed description of the incipient detection, recommended actions, and user interface developments during the last year are presented.

Quinn, Todd M.↗