Search NASASearch

Engineering topics

Hsueh, Mei-Chen

Publications and source records attributed to Hsueh, Mei-Chen.

Fault Injection Techniques and Tools

Dependability evaluation involves the study of failures and errors. The destructive nature of a crash and long error latency make it difficult to identify the causes of failures in the operational environment. It is particularly hard to recreate a failure scenario for a large, complex system. To identify and understand potential failures, we use an experiment-based approach for studying the dependability of a system. Such an approach is applied not only during the conception and design phases, but also during the prototype and operational phases. To take an experiment-based approach, we must first understand a system's architecture, structure, and behavior. Specifically, we need to know its tolerance for faults and failures, including its built-in detection and recovery mechanisms, and we need specific instruments and tools to inject faults, create failures or errors, and monitor their effects.

Hsueh, Mei-Chen

Analysis of field data on computer failures

This paper emphasizes the importance of making field measurements for effective and realistic dependability evaluations. Two examples are given, both based on real data from IBM mainframes. The first evaluates the impact of the operating environment on system failure characteristics and the second shows how an accurate model depicting this interaction can be extracted from real data.

Iyer, Ravishankar K.

Measurement-based reliability/performability models

Measurement-based models based on real error-data collected on a multiprocessor system are described. Model development from the raw error-data to the estimation of cumulative reward is also described. A workload/reliability model is developed based on low-level error and resource usage data collected on an IBM 3081 system during its normal operation in order to evaluate the resource usage/error/recovery process in a large mainframe system. Thus, both normal and erroneous behavior of the system are modeled. The results provide an understanding of the different types of errors and recovery processes. The measured data show that the holding times in key operational and error states are not simple exponentials and that a semi-Markov process is necessary to model the system behavior. A sensitivity analysis is performed to investigate the significance of using a semi-Markov process, as opposed to a Markov process, to model the measured system.

Hsueh, Mei-Chen