Search NASASearch

SEARCH · Search NASA

Results for “reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Software reliability: Application of a reliability model to requirements error analysis

The application of a software reliability model having a well defined correspondence of computer program properties to requirements error analysis is described. Requirements error categories which can be related to program structural elements are identified and their effect on program execution considered. The model is applied to a hypothetical B-5 requirement specification for a program module.

Logan, J.

Software reliability: A comparison of results obtained from established software reliability models

Two models of the software error detection process are compared, the Jelinski-Moranda model and a Bayes inference model. Simulation techniques are used to generate software related system failure data which is analyzed by both models. Point estimates and confidence limits are compared. It is demonstrated that uncertainty may be considerable for reasonable samples sizes and should be considered in any application of these techniques. The Jelinski-Moranda model is sensitive to the failure of data to follow internal assumptions of the model, often not providing any point estimates, a factor which may limit its usefulness in many real world situations. The Bayes model is shown to respond to the introduction of additional errors in the software correction process, a condition where error counting models such as the Jelinski-Moranda generally fail to converge.

Horn, M. H.

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability

Reliability Growth Modeling and Testing

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling

Modeling Reliability Growth

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling

Developing Reliable Life Support for Mars

A human mission to Mars will require highly reliable life support systems. Mars life support systems may recycle water and oxygen using systems similar to those on the International Space Station (ISS). However, achieving sufficient reliability is less difficult for ISS than it will be for Mars. If an ISS system has a serious failure, it is possible to provide spare parts, or directly supply water or oxygen, or if necessary bring the crew back to Earth. Life support for Mars must be designed, tested, and improved as needed to achieve high demonstrated reliability. A quantitative reliability goal should be established and used to guide development t. The designers should select reliable components and minimize interface and integration problems. In theory a system can achieve the component-limited reliability, but testing often reveal unexpected failures due to design mistakes or flawed components. Testing should extend long enough to detect any unexpected failure modes and to verify the expected reliability. Iterated redesign and retest may be required to achieve the reliability goal. If the reliability is less than required, it may be improved by providing spare components or redundant systems. The number of spares required to achieve a given reliability goal depends on the component failure rate. If the failure rate is under estimated, the number of spares will be insufficient and the system may fail. If the design is likely to have undiscovered design or component problems, it is advisable to use dissimilar redundancy, even though this multiplies the design and development cost. In the ideal case, a human tended closed system operational test should be conducted to gain confidence in operations, maintenance, and repair. The difficulty in achieving high reliability in unproven complex systems may require the use of simpler, more mature, intrinsically higher reliability systems. The limitations of budget, schedule, and technology may suggest accepting lower and less certain expected reliability. A plan to develop reliable life support is needed to achieve the best possible reliability.

life support

Uncertainties in obtaining high reliability from stress-strength models

There has been a recent interest in determining high statistical reliability in risk assessment of aircraft components. The potential consequences are identified of incorrectly assuming a particular statistical distribution for stress or strength data used in obtaining the high reliability values. The computation of the reliability is defined as the probability of the strength being greater than the stress over the range of stress values. This method is often referred to as the stress-strength model. A sensitivity analysis was performed involving a comparison of reliability results in order to evaluate the effects of assuming specific statistical distributions. Both known population distributions, and those that differed slightly from the known, were considered. Results showed substantial differences in reliability estimates even for almost nondetectable differences in the assumed distributions. These differences represent a potential problem in using the stress-strength model for high reliability computations, since in practice it is impossible to ever know the exact (population) distribution. An alternative reliability computation procedure is examined involving determination of a lower bound on the reliability values using extreme value distributions. This procedure reduces the possibility of obtaining nonconservative reliability estimates. Results indicated the method can provide conservative bounds when computing high reliability. An alternative reliability computation procedure is examined involving determination of a lower bound on the reliability values using extreme value distributions. This procedure reduces the possibility of obtaining nonconservative reliability estimates. Results indicated the method can provide conservative bounds when computing high reliability.

Neal, Donald M.

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence on a Time Limited Mission?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach applies to the case where the unit failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, the design then uses N redundant units, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is FN = F1N, and the needed N = LN(FN)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. The needed redundancy, N, to achieve the specified N unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - achieves the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some uncertainty, the confidence that the rate is not lower than the actual failure rate and the required reliability is not overestimated is about 50%. Adding more redundant units increases the confidence that the required reliability, FN, will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals have higher confidence in being achieved. Both the required reliability and confidence can be specified initially and the needed number of redundant units computed using the measured failure rate. The unit failure rate is determined by initial reliability growth testing to remove design errors and to better estimate the final constant failure rate. Reducing the failure rate and reducing its variance both reduce the number of redundant units needed for the required reliability and confidence. Since the total cost is the sum of the costs of the units and of the testing, there is an optimum test time that produces minimum cost.

Harry W. Jones

Reliability growth models for NASA applications

The objective of any reliability growth study is prediction of reliability at some future instant. Another objective is statistical inference, estimation of reliability for reliability demonstration. A cause of concern for the development engineer and management is that reliability demands an excessive number of tests for reliability demonstration. For example, the Space Transportation Main Engine (STME) program requirements call for .99 reliability at 90 pct. confidence for demonstration. This requires running 230 tests with zero failure if a classical binomial model is used. It is therefore also an objective to explore the reliability growth models for reliability demonstration and tracking and their applicability to NASA programs. A reliability growth model is an analytical tool used to monitor the reliability progress during the development program and to establish a test plan to demonstrate an acceptable system reliability.

Taneja, Vidya S.

FY12 End of Year Report for NEPP DDR2 Reliability

This document reports the status of the NASA Electronic Parts and Packaging (NEPP) Double Data Rate 2 (DDR2) Reliability effort for FY2012. The task expanded the focus of evaluating reliability effects targeted for device examination. FY11 work highlighted the need to test many more parts and to examine more operating conditions, in order to provide useful recommendations for NASA users of these devices. This year's efforts focused on development of test capabilities, particularly focusing on those that can be used to determine overall lot quality and identify outlier devices, and test methods that can be employed on components for flight use. Flight acceptance of components potentially includes considerable time for up-screening (though this time may not currently be used for much reliability testing). Manufacturers are much more knowledgeable about the relevant reliability mechanisms for each of their devices. We are not in a position to know what the appropriate reliability tests are for any given device, so although reliability testing could be focused for a given device, we are forced to perform a large campaign of reliability tests to identify devices with degraded reliability. With the available up-screening time for NASA parts, it is possible to run many device performance studies. This includes verification of basic datasheet characteristics. Furthermore, it is possible to perform significant pattern sensitivity studies. By doing these studies we can establish higher reliability of flight components. In order to develop these approaches, it is necessary to develop test capability that can identify reliability outliers. To do this we must test many devices to ensure outliers are in the sample, and we must develop characterization capability to measure many different parameters. For FY12 we increased capability for reliability characterization and sample size. We increased sample size this year by moving from loose devices to dual inline memory modules (DIMMs) with an approximate reduction of 20 to 50 times in terms of per device under test (DUT) cost. By increasing sample size we have improved our ability to characterize devices that may be considered reliability outliers. This report provides an update on the effort to improve DDR2 testing capability. Although focused on DDR2, the methods being used can be extended to DDR and DDR3 with relative ease.

Guertin, Steven M.

Addressing Unison and Uniqueness of Reliability and Safety for Better Integration

For a long time, both in theory and in practice, safety and reliability have not been clearly differentiated, which leads to confusion, inefficiency, and sometime counter-productive practices in executing each of these two disciplines. It is imperative to address the uniqueness and the unison of these two disciplines to help both disciplines become more effective and to promote a better integration of the two for enhancing safety and reliability in our products as an overall objective. There are two purposes of this paper. First, it will investigate the uniqueness and unison of each discipline and discuss the interrelationship between the two for awareness and clarification. Second, after clearly understanding the unique roles and interrelationship between the two in a product design and development life cycle, we offer suggestions to enhance the disciplines with distinguished and focused roles, to better integrate the two, and to improve unique sets of skills and tools of reliability and safety processes. From the uniqueness aspect, the paper identifies and discusses the respective uniqueness of reliability and safety from their roles, accountability, nature of requirements, technical scopes, detailed technical approaches, and analysis boundaries. It is misleading to equate unreliable to unsafe, since a safety hazard may or may not be related to the component, sub-system, or system functions, which are primarily what reliability addresses. Similarly, failing-to-function may or may not lead to hazard events. Examples will be given in the paper from aerospace, defense, and consumer products to illustrate the uniqueness and differences between reliability and safety. From the unison aspect, the paper discusses what the commonalities between reliability and safety are, and how these two disciplines are linked, integrated, and supplemented with each other to accomplish the customer requirements and product goals. In addition to understanding the uniqueness in reliability and safety, a better understanding of unison and commonalities will further help in understanding the interaction between reliability and safety. This paper discusses the unison and uniqueness of reliability and safety. It presents some suggestions for better integration of the two disciplines in terms of technical approaches, tools, techniques, and skills to enhance the role of reliability and safety in supporting a product design and development life cycle. The paper also discusses eliminating the redundant effort and minimizing the overlap of reliability and safety analyses for an efficient implementation of the two disciplines.

Huang, Zhaofeng