Search NASASearch

Engineering topics

Trivedi, K. S.

Publications and source records attributed to Trivedi, K. S..

Performability modeling based on real data: A case study

Described is a measurement-based performability model based on error and resource usage data collected on a multiprocessor system. A method for identifying the model structure is introduced and the resulting model is validated against real data. Model development from the collection of raw data to the estimation of the expected reward is described. Both normal and error behavior of the system are characterized. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model system behavior. A reward function, based on the service rate and the error rate in each state, is then defined in order to estimate the performability of the system and to depict the cost of apparent types of errors.

Hsueh, M. C.

Performability modeling based on real data: A casestudy

Described is a measurement-based performability model based on error and resource usage data collected on a multiprocessor system. A method for identifying the model structure is introduced and the resulting model is validated against real data. Model development from the collection of raw data to the estimation of the expected reward is described. Both normal and error behavior of the system are characterized. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model the system behavior. A reward function, based on the service rate and the error rate in each state, is then defined in order to estimate the performability of the system and to depict the cost of different types of errors.

Hsueh, M. C.

A measurement-based performability model for a multiprocessor system

A measurement-based performability model based on real error-data collected on a multiprocessor system is described. Model development from the raw errror-data to the estimation of cumulative reward is described. Both normal and failure behavior of the system are characterized. The measured data show that the holding times in key operational and failure states are not simple exponential and that semi-Markov process is necessary to model the system behavior. A reward function, based on the service rate and the error rate in each state, is then defined in order to estimate the performability of the system and to depict the cost of different failure types and recovery procedures.

Ilsueh, M. C.

Provably conservative approximations to complex reliability models

Complex models can be the bases for derivation of provably conservative and optimistic reliability models that incorporate a reduced state space and fewer transitions; they accordingly possess solutions that are more cost-effective than those of the original complex models. Design space can thereby be extensively explored without incurring the expense of multiple complex model solutions. A conservative-optimistic pair of derived models produces a band that includes the solution to the complex model. Sensitivity analysis can be performed on this pair of models to determine those parameters of the original model that are most sensitive to change and therefore require further expense in obtaining tighter specifications.

Smotherman, M.

The conservativeness of reliability estimates based on instantaneous coverage

Reliability modeling must take into account two different types of phenomena, including the fault-occurrence behavior and the fault/error-handling behavior of a system. The effectiveness of the fault/error-handling behavior can be captured by instantaneous coverage probabilities. This paper has the objective to show that the assumption of instantaneous coverage leads to conservative predictions of system reliability for systems characterized by relatively long interevent times for fault occurrences and relatively short interevent times for fault/error-handling actions. The importance of this result is related to the fact that it can now be shown that model predictions based on instantaneous coverage are lower bounds on the true system reliability. Attention is given to a semi-Markov reliability model, instantaneous coverage approximations, the proof of conservative prediction, and the computation of coverage probabilities.

Mcgough, J.

Ultrahigh reliability prediction for fault-tolerant computer systems

A review and a critical evaluation of a representative class of state-of-the-art models for ultrahigh reliability prediction is presented. This evaluation naturally leads to a new model for ultrahigh reliability prediction now under development. The new model combines the flexibility and accuracy of simulation with the speed of analytic models.

Geist, R. M.

Decomposition in reliability analysis of fault-tolerant systems

The existing approaches to reliability modeling are briefly reviewed. An examination of the limitations of the existing approaches in modeling ultrareliable fault-tolerant systems illustrates the need to use decomposition techniques. The notion of behavioral decomposition is introduced for dealing with reliability models with a large number of states, and a series of examples is presented. The CARE (computer-aided reliability estimation) and HARP (hybrid automated reliability predictor) approaches to reliability are discussed.

Trivedi, K. S.

A tutorial on the CARE III approach to reliability modeling

The CARE 3 reliability model for aircraft avionics and control systems is described by utilizing a number of examples which frequently use state-of-the-art mathematical modeling techniques as a basis for their exposition. Behavioral decomposition followed by aggregration were used in an attempt to deal with reliability models with a large number of states. A comprehensive set of models of the fault-handling processes in a typical fault-tolerant system was used. These models were semi-Markov in nature, thus removing the usual restrictions of exponential holding times within the coverage model. The aggregate model is a non-homogeneous Markov chain, thus allowing the times to failure to posses Weibull-like distributions. Because of the departures from traditional models, the solution method employed is that of Kolmogorov integral equations, which are evaluated numerically.

Trivedi, K. S.

Validation Methods Research for Fault-Tolerant Avionics and Control Systems Sub-Working Group Meeting. CARE 3 peer review

A computer aided reliability estimation procedure (CARE 3), developed to model the behavior of ultrareliable systems required by flight-critical avionics and control systems, is evaluated. The mathematical models, numerical method, and fault-tolerant architecture modeling requirements are examined, and the testing and characterization procedures are discussed. Recommendations aimed at enhancing CARE 3 are presented; in particular, the need for a better exposition of the method and the user interface is emphasized.

Trivedi, K. S.

Validation Methods Research for Fault-Tolerant Avionics and Control Systems: Working Group Meeting, 2

The validation process comprises the activities required to insure the agreement of system realization with system specification. A preliminary validation methodology for fault tolerant systems documented. A general framework for a validation methodology is presented along with a set of specific tasks intended for the validation of two specimen system, SIFT and FTMP. Two major areas of research are identified. First, are those activities required to support the ongoing development of the validation process itself, and second, are those activities required to support the design, development, and understanding of fault tolerant systems.

Gault, J. W.

Reliability validation of systems for life-critical applications

A framework is proposed which addresses traditional reliability validation approaches consisting of life testing techniques which are inapplicable for digital flight control systems. A specific validation methodology is identified based on logical proofs, analytical modeling, and experimental testing. Research activities required to support continued development of validation technology are identified, and the validation procedure is driven by the reliability model obtained from the system description. The analytical reliability model is shown to be a proper abstraction of the system under consideration, and a proof of correctness of system design and system scheduler performance is proposed.

Trivedi, K. S.