Search NASA⌕ Search

Engineering topics

Chillarege, Ram

Publications and source records attributed to Chillarege, Ram.

An experimental study of memory fault latency

The difficulty with the measurement of fault latency is due to the lack of observability of the fault occurrence and error generation instants in a production environment. The authors describe an experiment, using data from a VAX 11/780 under real workload, to study fault latency in the memory subsystem accurately. Fault latency distributions are generated for stuck-at-zero (s-a-0) and stuck-at-one (s-a-1) permanent fault models. The results show that the mean fault latency of an s-a-0 fault is nearly five times that of the s-a-1 fault. An analysis of variance is performed to quantify the relative influence of different workload measures on the evaluated latency.

Chillarege, Ram↗

Measurement-based analysis of error latency

This paper demonstrates a practical methodology for the study of error latency under a real workload. The method is illustrated with sampled data on the physical memory activity, gathered by hardware instrumentation on a VAX 11/780 during the normal workload cycle of the installation. These data are used to simulate fault occurrence and to reconstruct the error discovery process in the system. The technique provides a means to study the system under different workloads and for multiple days. An approach to determine the percentage of undiscovered errors is also developed and a verification of the entire methodology is performed. This study finds that the mean error latency, in the memory containing the operating system, varies by a factor of 10 to 1 (in hours) between the low and high workloads. It is found that of all errors occurring within a day, 70 percent are detected in the same day, 82 percent within the following day, and 91 percent within the third day. The increase in failure rate due to latency is not so much a function of remaining errors but is dependent on whether or not there is a latent error.

Chillarege, Ram↗

Fault and Error Latency Under Real Workload: an Experimental Study

A practical methodology for the study of fault and error latency is demonstrated under a real workload. This is the first study that measures and quantifies the latency under real workload and fills a major gap in the current understanding of workload-failure relationships. The methodology is based on low level data gathered on a VAX 11/780 during the normal workload conditions of the installation. Fault occurrence is simulated on the data, and the error generation and discovery process is reconstructed to determine latency. The analysis proceeds to combine the low level activity data with high level machine performance data to yield a better understanding of the phenomena. A strong relationship exists between latency and workload and that relationship is quantified. The sampling and reconstruction techniques used are also validated. Error latency in the memory where the operating system resides was studied using data on the physical memory access. Fault latency in the paged section of memory was determined using data from physical memory scans. Error latency in the microcontrol store was studied using data on the microcode access and usage.

Chillarege, Ram↗

Fault latency in the memory - An experimental study on VAX 11/780

Fault latency is the time between the physical occurrence of a fault and its corruption of data, causing an error. The measure of this time is difficult to obtain because the time of occurrence of a fault and the exact moment of generation of an error are not known. This paper describes an experiment to accurately study the fault latency in the memory subsystem. The experiment employs real memory data from a VAX 11/780 at the University of Illinois. Fault latency distributions are generated for s-a-0 and s-a-1 permanent fault models. Results show that the mean fault latency of a s-a-0 fault is nearly 5 times that of the s-a-1 fault. Large variations in fault latency are found for different regions in memory. An analysis of a variance model to quantify the relative influence of various workload measures on the evaluated latency is also given.

Chillarege, Ram↗