Search NASASearch

Engineering topics

Finelli, George B.

Publications and source records attributed to Finelli, George B..

The infeasibility of quantifying the reliability of life-critical real-time software

This paper affirms that the quantification of life-critical software reliability is infeasible using statistical methods, whether these methods are applied to standard software or fault-tolerant software. The classical methods of estimating reliability are shown to lead to exorbitant amounts of testing when applied to life-critical software. Reliability growth models are examined and also shown to be incapable of overcoming the need for excessive amounts of testing. The key assumption of software fault tolerance - separately programmed versions fail independently - is shown to be problematic. This assumption cannot be justified by experimentation in the ultrareliability region, and subjective arguments in its favor are not sufficiently strong to justify it as an axiom. Also, the implications of the recent multiversion software experiments support this affirmation.

Butler, Ricky W.

The Infeasibility of Experimental Quantification of Life-Critical Software Reliability

This paper affirms that quantification of life-critical software reliability is infeasible using statistical methods whether applied to standard software or fault-tolerant software. The key assumption of software fault tolerance|separately programmed versions fail independently|is shown to be problematic. This assumption cannot be justified by experimentation in the ultra-reliability region and subjective arguments in its favor are not sufficiently strong to justify it as an axiom. Also, the implications of the recent multi-version software experiments support this affirmation.

Butler, Ricky W.

The Infeasibility of Quantifying the Reliability of Life-Critical Real-Time Software

This paper affirms that the quantification of life-critical software reliability is infeasible using statistical methods whether applied to standard software or fault-tolerant software. The classical methods of estimating reliability are shown to lead to exhorbitant amounts of testing when applied to life-critical software. Reliability growth models are examined and also shown to be incapable of overcoming the need for excessive amounts of testing. The key assumption of software fault tolerance separately programmed versions fail independently is shown to be problematic. This assumption cannot be justified by experimentation in the ultrareliability region and subjective arguments in its favor are not sufficiently strong to justify it as an axiom. Also, the implications of the recent multiversion software experiments support this affirmation.

Butler, Ricky W.

Real-time software failure characterization

A series of studies aimed at characterizing the fundamentals of the software failure process has been undertaken as part of a NASA project on the modeling of a real-time aerospace vehicle software reliability. An overview of these studies is provided, and the current study, an investigation of the reliability of aerospace vehicle guidance and control software, is examined. The study approach provides for the collection of life-cycle process data, and for the retention and evaluation of interim software life-cycle products.

Dunham, Janet R.

Results of software error-data experiments

In order to evaluate existing software reliability models and proposed modeling approaches, a search was conducted for data on the software failure process. This search revealed that the data necessary for this evaluation were not available. As a result, a research effort was initiated by NASA to generate data on which to base the development of credible methods for assessing the reliability of software targeted for flight-crucial applications. Two sets of software error-data experiments were conducted by different research groups. The results of the experiments were consistent: errors caused by different faults in a program occurred at widely varying rates; program failure rates exhibited a log-linear trend with respect to the number of faults corrected; some faults were found to interact in either concealing or revealing ways; and contiguous regions of the input space which cause a program to generate errors, called error crystals, were found and characterized for some faults. Collectively, these experiments have produced information on software failure which must be accounted for in software reliability modeling approaches.

Finelli, George B.

A technique for evaluating the application of the pin-level stuck-at fault model to VLSI circuits

Accurate fault models are required to conduct the experiments defined in validation methodologies for highly reliable fault-tolerant computers (e.g., computers with a probability of failure of 10 to the -9 for a 10-hour mission). Described is a technique by which a researcher can evaluate the capability of the pin-level stuck-at fault model to simulate true error behavior symptoms in very large scale integrated (VLSI) digital circuits. The technique is based on a statistical comparison of the error behavior resulting from faults applied at the pin-level of and internal to a VLSI circuit. As an example of an application of the technique, the error behavior of a microprocessor simulation subjected to internal stuck-at faults is compared with the error behavior which results from pin-level stuck-at faults. The error behavior is characterized by the time between errors and the duration of errors. Based on this example data, the pin-level stuck-at fault model is found to deliver less than ideal performance. However, with respect to the class of faults which cause a system crash, the pin-level, stuck-at fault model is found to provide a good modeling capability.

Palumbo, Daniel L.

Characterization of fault recovery through fault injection on FTMP

The development of fault-injection procedures and statistical analysis techniques to characterize the fault recovery of fault-tolerant systems is described. Pin-level fault-injection was conducted on a fault-tolerant microprocessor computer in order to generate data to assess the utility of current fault-injection sampling methods. The validity of common reliability-modeling assumptions concerning the statistical distribution of recovery times is investigated. A multiple comparison analysis for detecting behavior variations, and a distribution fitting for determining the best fit for the data were conducted. It is observed that the detection behavior is not homogeneous across all data sets, and that none of the factors under experimental control can account for the observed groupings of behavior. It is determined that no single distribution fits all the data sets, and that stratified random sampling and statistically robust parameter-estimation techniques are required to characterize fault detection time.

Finelli, George B.

Fault-free performance validation of avionic multiprocessors

This paper describes the application of a portion of a validation methodology to NASA's Fault-Tolerant Multiprocessor System (FTMP) and the Software Implemented Fault-Tolerance (SIFT) computer system. The methodology entails a building block approach, starting with simple baseline experiments and building to more complex experiments. The goal of the validation methodology is to thoroughly test and characterize the performance and behavior of ultrareliable computer systems. The validation methodology presented in this paper showed that the methodology is not machine specific and can be used in lieu of life testing approaches. By applying a building block approach at the systems level, the machine complexity was broken down to manageable levels independent of system implementation.

Czeck, Edward W.