Methods of design stage reliability analysis.
Reliability methodology in design and development stage, discussing component stresses, performance variation, etc
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Reliability methodology in design and development stage, discussing component stresses, performance variation, etc
This report describes the preliminary results of an investigation on component reliability analysis and reliability-based design optimization of thin-walled circular composite cylinders with average diameter and average length of 15 inches. Structural reliability is based on axial buckling strength of the cylinder. Both Monte Carlo simulation and First Order Reliability Method are considered for reliability analysis with the latter incorporated into the reliability-based structural optimization problem. To improve the efficiency of reliability sensitivity analysis and design optimization solution, the buckling strength of the cylinder is estimated using a second-order response surface model. The sensitivity of the reliability index with respect to the mean and standard deviation of each random variable is calculated and compared. The reliability index is found to be extremely sensitive to the applied load and elastic modulus of the material in the fiber direction. The cylinder diameter was found to have the third highest impact on the reliability index. Also the uncertainty in the applied load, captured by examining different values for its coefficient of variation, is found to have a large influence on cylinder reliability. The optimization problem for minimum weight is solved subject to a design constraint on element reliability index. The methodology, solution procedure and optimization results are included in this report.
The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.
A reliability analysis method for computing systems is considered in which the underlying criteria for 'success' are based on the computations the system must perform in the use environment. Beginning with a general model of a 'computer with faults', intermediate concepts of a 'tolerance relation' and an 'environment space' are introduced which account for the computational needs of the user and the probabilistic nature of the use environment. These concepts are then incorporated to obtain a precisely defined class of computation-based reliability measures. Formulation of a particular measure is illustrated and results, applying this measure, are compared with those of a typical structure-based analysis.
The SCARE (Structural Ceramics Analysis and Reliability Evaluation) computer program on statistical fast fracture reliability analysis with quadratic elements for volume distributed imperfections is enhanced to include the use of linear finite elements and the capability of designing against concurrent surface flaw induced ceramic component failure. The SCARE code is presently coupled as a postprocessor to the MSC/NASTRAN general purpose, finite element analysis program. The improved version now includes the Weibull and Batdorf statistical failure theories for both surface and volume flaw based reliability analysis. The program uses the two-parameter Weibull fracture strength cumulative failure probability distribution model with the principle of independent action for poly-axial stress states, and Batdorf's shear-sensitive as well as shear-insensitive statistical theories. The shear-sensitive surface crack configurations include the Griffith crack and Griffith notch geometries, using the total critical coplanar strain energy release rate criterion to predict mixed-mode fracture. Weibull material parameters based on both surface and volume flaw induced fracture can also be calculated from modulus of rupture bar tests, using the least squares method with known specimen geometry and grouped fracture data. The statistical fast fracture theories for surface flaw induced failure, along with selected input and output formats and options, are summarized. An example problem to demonstrate various features of the program is included.
The SCARE (Structural Ceramics Analysis and Reliability Evaluation) computer program on statistical fast fracture reliability analysis with quadratic elements for volume distributed imperfections is enhanced to include the use of linear finite elements and the capability of designing against concurrent surface flaw induced ceramic component failure. The SCARE code is presently coupled as a postprocessor to the MSC/NASTRAN general purpose, finite element analysis program. The improved version now includes the Weibull and Batdorf statistical failure theories for both surface and volume flaw based reliability analysis. The program uses the two-parameter Weibull fracture strength cumulative failure probability distribution model with the principle of independent action for poly-axial stress states, and Batdorf's shear-sensitive as well as shear-insensitive statistical theories. The shear-sensitive surface crack configurations include the Griffith crack and Griffith notch geometries, using the total critical coplanar strain energy release rate criterion to predict mixed-mode fracture. Weibull material parameters based on both surface and volume flaw induced fracture can also be calculated from modulus of rupture bar tests, using the least squares method with known specimen geometry and grouped fracture data. The statistical fast fracture theories for surface flaw induced failure, along with selected input and output formats and options, are summarized. An example problem to demonstrate various features of the program is included.
Some basic methods used in reliability analysis are problematic because they produce incorrect and overoptimistic predictions. Initially gratifying forecasts are often invalidated by testing and operational experience. The problematic methods in reliability analysis include estimating the system failure rate as the sum of component failure rates, assuming that reliability growth continues indefinitely during testing, overestimating the benefits of redundancy, and using the fault tolerance count instead of a detailed reliability analysis. Reliability analysis can produce more optimism than accuracy. This bug may now be a feature. The optimistic bias inevitable in project planning should be corrected by realistic reliability analysis that reflects relevant experience. That the repeated poor performance of reliability analysis is found to be surprising suggests willful blindness. Rigorous methods and impartial critical review are necessary to improve reliability analysis.
Some basic methods used in reliability analysis are problematic because they produce incorrect and overoptimistic predictions. Initially gratifying forecasts are often invalidated by testing and operational experience. The problematic methods in reliability analysis include estimating the system failure rate as the sum of component failure rates, assuming that reliability growth continues indefinitely during testing, overestimating the benefits of redundancy, and using the fault tolerance count instead of a detailed reliability analysis. Reliability analysis can produce more optimism than accuracy. This bug may now be a feature. The optimistic bias inevitable in project planning should be corrected by realistic reliability analysis that reflects relevant experience. That the repeated poor performance of reliability analysis is found to be surprising suggests willful blindness. Rigorous methods and impartial critical review are necessary to improve reliability analysis.
To support data collection for dynamic human reliability analysis (HRA), this study investigates time distributions for task primitives defined in the Goals, Operators, Methods, and Selection rules (GOMS)–Human Reliability Analysis (HRA) method and Human Reliability data EXtraction (HuREX). GOMS-HRA was developed to provide cognition-based time and human error probability (HEP) information for dynamic HRA calculations within the Human Unimodel for Nuclear Technology to Enhance Reliability (HUNTER) framework, while HuREX is a comprehensive HRA data collection method developed by the Korea Atomic Energy Research Institute (KAERI). In this paper, we examine time distributions by using experimental data collected from the Simplified Human Error Experimental Program (SHEEP) study, which proposes an HRA data collection framework to complement full-scope simulator research and gather input data for dynamic HRA by using simplified simulators such as the Rancor Microworld simulator. This paper investigates whether the time required for GOMS-HRA and HuREX task primitives fits 13 statistical distributions. Additionally, we compare and discuss the time distributions obtained from both student operators and professional operators. The result was that this study identified several time distributions for five GOMS-HRA and four HuREX task primitives. In the future, the results of this study are expected to provide objective reference data on the elapsed time for task primitives and aid in realistically simulating scenarios within dynamic HRA.
Traditional human reliability analysis (HRA) methods have difficulty dealing with the dynamic nature of factors such as time and rely on static and expert-judgment-based assessments of performance-shaping factors (PSFs) across limited levels. In this study, we introduce a mathematical approach for dynamically evaluating the experience and training PSF. Our proposed method integrates the psychological concept of the “forgetting curve” to evaluate how PSFs are impacted by the number of trainings and the time elapsed since training. To confirm the validity of the model, we provide experimental data fitted by identifying the quantitative relationship between training and human performance. This research enables dynamic and objective assessments, thus reducing reliance on subjective expert judgment and improving the accuracy of HRA.
System reliability equation, an exact function of component reliabilities, for a system with a finite number of points is derived from the minimal states which are found by logical analysis of the configuration. The numerical value is obtained by substituting the component reliabilities or unreliabilities.
Emulation techniques applied to the analysis of the reliability of highly reliable computer systems for future commercial aircraft are described. The lack of credible precision in reliability estimates obtained by analytical modeling techniques is first established. The difficulty is shown to be an unavoidable consequence of: (1) a high reliability requirement so demanding as to make system evaluation by use testing infeasible; (2) a complex system design technique, fault tolerance; (3) system reliability dominated by errors due to flaws in the system definition; and (4) elaborate analytical modeling techniques whose precision outputs are quite sensitive to errors of approximation in their input data. Next, the technique of emulation is described, indicating how its input is a simple description of the logical structure of a system and its output is the consequent behavior. Use of emulation techniques is discussed for pseudo-testing systems to evaluate bounds on the parameter values needed for the analytical techniques. Finally an illustrative example is presented to demonstrate from actual use the promise of the proposed application of emulation.
Apollo spacecraft reliability analysis techniques, both probabilistic and qualitative
Emulation techniques are proposed as a solution to a difficulty arising in the analysis of the reliability of highly reliable computer systems for future commercial aircraft. The difficulty, viz., the lack of credible precision in reliability estimates obtained by analytical modeling techniques are established. The difficulty is shown to be an unavoidable consequence of: (1) a high reliability requirement so demanding as to make system evaluation by use testing infeasible, (2) a complex system design technique, fault tolerance, (3) system reliability dominated by errors due to flaws in the system definition, and (4) elaborate analytical modeling techniques whose precision outputs are quite sensitive to errors of approximation in their input data. The technique of emulation is described, indicating how its input is a simple description of the logical structure of a system and its output is the consequent behavior. The use of emulation techniques is discussed for pseudo-testing systems to evaluate bounds on the parameter values needed for the analytical techniques.
The current generation of reliability analysis tools concentrates on improving the efficiency of the description and solution of the fault-handling processes and providing a solution algorithm for the full system model. The tools have improved user efficiency in these areas to the extent that the problem of constructing the fault-occurrence model is now the major analysis bottleneck. For the next generation of reliability tools, it is proposed that techniques be developed to improve the efficiency of the fault-occurrence model generation and input. Further, the goal is to provide an environment permitting a user to provide a top-down design description of the system from which a Markov reliability model is automatically constructed. Thus, the user is relieved of the tedious and error-prone process of model construction, permitting an efficient exploration of the design space, and an independent validation of the system's operation is obtained. An additional benefit of automating the model construction process is the opportunity to reduce the specialized knowledge required. Hence, the user need only be an expert in the system he is analyzing; the expertise in reliability analysis techniques is supplied.
It is noted that current reliability analysis tools differ not only in their solution techniques, but also in their approach to model abstraction. The analyst must be satisfied with the constraints that are intrinsic to any combination of solution technique and model abstraction. To get a better idea of the nature of these constraints, three reliability analysis tools (HARP, ASSIST/SURE, and CAME) were used to model portions of the Integrated Airframe/Propulsion Control System architecture. When presented with the example problem, all three tools failed to produce correct results. In all cases, either the tool or the model had to be modified. It is suggested that most of the difficulty is rooted in the large model size and long computational times which are characteristic of Markov model solutions.
An outline is presented of issues raised in verifying the accuracy of reliability analysis tools. State-of-the-art reliability analysis tools implement various decomposition, aggregation, and estimation techniques to compute the reliability of a diversity of complex fault-tolerant computer systems. However, no formal methodology has been formulated for validating the reliability estimates produced by these tools. The author presents three states of testing that can be performed on most reliability analysis tools to effectively increase confidence in a tool. These testing stages were applied to the SURE (semi-Markov Unreliability Range Evaluator) reliability analysis tool, and the results of the testing are discussed.
We consider the reliability analysis of phased-mission systems with common-cause failures in this paper. Phased-mission systems (PMS) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. System components may be subject to different stresses as well as different reliability requirements throughout the course of the mission. As a result, component behavior and relationships may need to be modeled differently from phase to phase when performing a system-level reliability analysis. This consideration poses unique challenges to existing analysis methods. The challenges increase when common-cause failures (CCF) are incorporated in the model. CCF are multiple dependent component failures within a system that are a direct result of a shared root cause, such as sabotage, flood, earthquake, power outage, or human errors. It has been shown by many reliability studies that CCF tend to increase a system's joint failure probabilities and thus contribute significantly to the overall unreliability of systems subject to CCF.We propose a separable phase-modular approach to the reliability analysis of phased-mission systems with dependent common-cause failures as one way to meet the above challenges in an efficient and elegant manner. Our methodology is twofold: first, we separate the effects of CCF from the PMS analysis using the total probability theorem and the common-cause event space developed based on the elementary common-causes; next, we apply an efficient phase-modular approach to analyze the reliability of the PMS. The phase-modular approach employs both combinatorial binary decision diagram and Markov-chain solution methods as appropriate. We provide an example of a reliability analysis of a PMS with both static and dynamic phases as well as CCF as an illustration of our proposed approach. The example is based on information extracted from a Mars orbiter project. The reliability model for this orbiter considers the various phases of Launch, Cruise, Mars Orbit Insertion, and Orbit. Some of the CCF for the orbiter in this mission include environmental effects, such as micrometeoroids, human operator errors, and software errors.