Search NASA⌕ Search

SEARCH · Search NASA

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Reliability Growth Modeling and Testing

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling↗

Modeling Reliability Growth

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling↗

Heroic Reliability Improvement in Manned Space Systems

System reliability can be significantly improved by a strong continued effort to identify and remove all the causes of actual failures. Newly designed systems often have unexpected high failure rates which can be reduced by successive design improvements until the final operational system has an acceptable failure rate. There are many causes of failures and many ways to remove them. New systems may have poor specifications, design errors, or mistaken operations concepts. Correcting unexpected problems as they occur can produce large early gains in reliability. Improved technology in materials, components, and design approaches can increase reliability. The reliability growth is achieved by repeatedly operating the system until it fails, identifying the failure cause, and fixing the problem. The failure rate reduction that can be obtained depends on the number and the failure rates of the correctable failures. Under the strong assumption that the failure causes can be removed, the decline in overall failure rate can be predicted. If a failure occurs at the rate of lambda per unit time, the expected time before the failure occurs and can be corrected is 1/lambda, the Mean Time Before Failure (MTBF). Finding and fixing a less frequent failure with the rate of lambda/2 per unit time requires twice as long, time of 1/(2 lambda). Cutting the failure rate in half requires doubling the test and redesign time and finding and eliminating the failure causes.Reducing the failure rate significantly requires a heroic reliability improvement effort.

life support↗

High Reliability at Minimum Cost

This paper investigates the minimum cost of improving the reliability of complex technical systems. The two major methods to improve reliability are redesigning the system for higher reliability or providing redundant components to replace failed elements. The costs of redesign for reliability or adding redundancy are estimated. The most cost-effective combination for high reliability can be identified. The cost of increasing the intrinsic reliability of a system can be modeled as cost proportional to 1/(system failure rate) a , where the exponent “a” measures the difficulty of increasing reliability. The “a” exponent can vary from 0.25 to about 2.5. Operational reliability can also be increased by using redundant systems. The failure rate for N parallel redundant units is (system failure rate) N . The cost of redundancy is N times the system cost. The total redundant system cost is proportional to N/(system failure rate) a . The cost of redundancy increases as N gets larger, but larger N allows a higher system failure rate, which reduces the system design cost. There is a certain N, a certain level of redundancy, that has the minimum cost to achieve the required overall redundant system failure rate. The minimum cost for the redundant system is achieved at the optimum level of redundancy. The N for minimum cost is equal to -a ln (redundant system failure rate). The minimum cost of the N redundant systems is proportional to N * (original system failure rate) a . The optimum redesigned individual system failure rate is proportional to exp (-1/a), so the greater the difficulty, the higher the optimum individual system failure rate. Increasing the intrinsic reliability of a system encounters diminishing returns and at some point it becomes more cost-effective to add redundancy. The difficulty of increasing intrinsic system reliability determines the optimum design for high reliability at minimum cost.

reliability↗

High Reliability at Minimum Cost

This paper investigates the minimum cost of improving the reliability of complex technical systems. The two major methods to improve reliability are redesigning the system for higher reliability or providing redundant components to replace failed elements. The costs of redesign for reliability or adding redundancy are estimated. The most cost-effective combination for high reliability can be identified. The cost of increasing the intrinsic reliability of a system can be modeled as cost proportional to 1/(system failure rate) a , where the exponent “a” measures the difficulty of increasing reliability. The “a” exponent can vary from 0.25 to about 2.5. Operational reliability can also be increased by using redundant systems. The failure rate for N parallel redundant units is (system failure rate) N . The cost of redundancy is N times the system cost. The total redundant system cost is proportional to N/(system failure rate) a . The cost of redundancy increases as N gets larger, but larger N allows a higher system failure rate, which reduces the system design cost. There is a certain N, a certain level of redundancy, that has the minimum cost to achieve the required overall redundant system failure rate. The minimum cost for the redundant system is achieved at the optimum level of redundancy. The N for minimum cost is equal to -a ln (redundant system failure rate). The minimum cost of the N redundant systems is proportional to N * (original system failure rate) a . The optimum redesigned individual system failure rate is proportional to exp (-1/a), so the greater the difficulty, the higher the optimum individual system failure rate. Increasing the intrinsic reliability of a system encounters diminishing returns and at some point it becomes more cost-effective to add redundancy. The difficulty of increasing intrinsic system reliability determines the optimum design for high reliability at minimum cost.

reliability↗

Service Life Extension of the Propulsion System of Long-Term Manned Orbital Stations

One of the critical non-replaceable systems of a long-term manned orbital station is the propulsion system. Since the propulsion system operates beginning with the launch of station elements into orbit, its service life determines the service life of the station overall. Weighing almost a million pounds, the International Space Station (ISS) is about four times as large as the Russian space station Mir and about five times as large as the U.S. Skylab. Constructed over a span of more than a decade with the help of over 100 space flights, elements and modules of the ISS provide more research space than any spacecraft ever built. Originally envisaged for a service life of fifteen years, this Earth orbiting laboratory has been in orbit since 1998. Some elements that have been launched later in the assembly sequence were not yet built when the first elements were placed in orbit. Hence, some of the early modules that were launched at the inception of the program were already nearing the end of their design life when the ISS was finally ready and operational. To maximize the return on global investments on ISS, it is essential for the valuable research on ISS to continue as long as the station can be sustained safely in orbit. This paper describes the work performed to extend the service life of the ISS propulsion system. A system comprises of many components with varying failure rates. Reliability of a system is the probability that it will perform its intended function under encountered operating conditions, for a specified period of time. As we are interested in finding out how reliable a system would be in the future, reliability expressed as a function of time provides valuable insight. In a hypothetical bathtub shaped failure rate curve, the failure rate, defined as the number of failures per unit time that a currently healthy component will suffer in a given future time interval, decreases during infant-mortality period, stays nearly constant during the service life and increases at the end when the design service life ends and wear-out phase begins. However, the component failure rates do not remain constant over the entire cycle life. The failure rate depends on various factors such as design complexity, current age of the component, operating conditions, severity of environmental stress factors, etc. Development, qualification and acceptance test processes provide rigorous screening of components to weed out imperfections that might otherwise cause infant mortality failures. If sufficient samples are tested to failure, the failure time versus failure quantity can be analyzed statistically to develop a failure probability distribution function (PDF), a statistical model of the probability of failure versus time. Driven by cost and schedule constraints however, spacecraft components are generally not tested in large numbers. Uncertainties in failure rate and remaining life estimates increase when fewer units are tested. To account for this, spacecraft operators prefer to limit useful operations to a period shorter than the maximum demonstrated service life of the weakest component. Running each component to its failure to determine the maximum possible service life of a system can become overly expensive and impractical. Spacecraft operators therefore, specify the required service life and an acceptable factor of safety (FOS). The designers use these requirements to limit the life test duration. Midway through the design life, when benefits justify additional investments, supplementary life test may be performed to demonstrate the capability to safely extend the service life of the system. An innovative approach is required to evaluate the entire system, without having to go through an elaborate test program of propulsion system elements. Evaluating every component through a brute force test program would be a cost prohibitive and time consuming endeavor. ISS propulsion system components were designed and built decades ago. There are no representative ground test articles for some of the components. A 'test everything' approach would require manufacturing new test articles. The paper outlines some of the techniques used for selective testing, by way of cherry picking candidate components based on failure mode effects analysis, system level impacts, hazard analysis, etc. The type of testing required for extending the service life depends on the design and criticality of the component, failure modes and failure mechanisms, life cycle margin provided by the original certification, operational and environmental stresses encountered, etc. When specific failure mechanism being considered and the underlying relationship of that mode to the stresses provided in the test can be correlated by supporting analysis, time and effort required for conducting life extension testing can be significantly reduced. Exposure to corrosive propellants over long periods of time, for instance, lead to specific failure mechanisms in several components used in the propulsion system. Using Arrhenius model, which is tied to chemically dependent failure mechanisms such as corrosion or chemical reactions, it is possible to subject carefully selected test articles to accelerated life test. Arrhenius model reflects the proportional relationship between time to failure of a component and the exponential of the inverse of absolute temperature acting on the component. The acceleration factor is used to perform tests at higher stresses that allow direct correlation between the times to failure at a high test temperature to the temperatures to be expected in actual use. As long as the temperatures are such that new failure mechanisms are not introduced, this becomes a very useful method for testing to failure a relatively small sample of items for a much shorter amount of time. In this article, based on the example of the propulsion system of the first ISS module Zarya, theoretical approaches and practical activities of extending the service life of the propulsion system are reviewed with the goal of determining the maximum duration of its safe operation.

Kamath, Ulhas↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence on a Time Limited Mission?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach applies to the case where the unit failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, the design then uses N redundant units, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is FN = F1N, and the needed N = LN(FN)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. The needed redundancy, N, to achieve the specified N unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - achieves the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some uncertainty, the confidence that the rate is not lower than the actual failure rate and the required reliability is not overestimated is about 50%. Adding more redundant units increases the confidence that the required reliability, FN, will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals have higher confidence in being achieved. Both the required reliability and confidence can be specified initially and the needed number of redundant units computed using the measured failure rate. The unit failure rate is determined by initial reliability growth testing to remove design errors and to better estimate the final constant failure rate. Reducing the failure rate and reducing its variance both reduce the number of redundant units needed for the required reliability and confidence. Since the total cost is the sum of the costs of the units and of the testing, there is an optimum test time that produces minimum cost.

Harry W. Jones↗

The abcd Reliability Growth Model

This paper presents a modification of the well-known Duane-Crow reliability growth model. In the abcd reliability growth model, the initial period of exponential decline of the failure rate in the Duane-Crow model may be followed by a period of constant failure rate. Data often show that an exponential decline in failures is followed by a constant failure rate. If a growth model including only the initial period of exponential decline is applied to increasingly longer failure rate data sets, the data will include longer periods of constant failure rate, and the estimated reliability growth rate will decline from an initially high value down toward zero. Using the Duane-Crow model without extending it to include a possible period of constant failure rate may create the mistaken impression that the initial reliability growth continues forever, but at an ever decreasing rate.

Reliability growth↗

The abcd Reliability Growth Model

This paper presents a modification of the well-known Duane-Crow reliability growth model. In the abcd reliability growth model, the initial period of exponential decline of the failure rate in the Duane-Crow model may be followed by a period of constant failure rate. Data often show that an exponential decline in failures is followed by a constant failure rate. If a growth model including only the initial period of exponential decline is applied to increasingly longer failure rate data sets, the data will include longer periods of constant failure rate, and the estimated reliability growth rate will decline from an initially high value down toward zero. Using the Duane-Crow model without extending it to include a possible period of constant failure rate may create the mistaken impression that the initial reliability growth continues forever, but at an ever decreasing rate.

Reliability growth↗

International Space Station Operational Experience and Its Impacts on Future Mission Supportability

Operational experience gained on the International Space Station (ISS) has enabled significant improvements in failure rate estimates for various Orbital Replacement Units (ORUs). These improved estimates, in turn, allow more efficient and accurate spare parts allocations for future missions, enabling significant reductions in both logistics mass and risk. This paper examines the value of ISS experience to date in terms of its impact on supportability for future missions. A supportability model is presented that assesses the spares required as a function of mission endurance and risk, taking into account uncertainty in failure rate estimates. Changes in ISS Environmental Control and Life Support (ECLSS) ORU failure rate estimates are described and discussed, both in terms of the overall population of ORUs and the evolution of failure rate estimates over time for a particular item. The value of those updated failure rate estimates is assessed by calculating the estimated spares mass requirements for two cases, using the initial, pre-ISS estimates and using the estimates informed by on-orbit experience. Hidden risk resulting from underestimated failure rates is also assessed. These results indicate that, for a 1,200-day Mars mission, ISS experience has enabled a 3.9 t to 6.0 t reduction in ECLSS spares mass required and uncovered failure rate underestimates that would have resulted in an order of magnitude increase in risk had they not been discovered and corrected. The implications of these results for system development and mission planning are discussed, including approaches to accelerate the rate of failure rate refinement and the risks associated with making changes or introducing new systems. Overall, test time is a critical factor that must be carefully considered in system development, and new systems must budget appropriate time for testing in a relevant environment or accept higher risk and logistics requirements on future missions.

Owens, Andrew C.↗

Developing Reliable Life Support for Mars

A human mission to Mars will require highly reliable life support systems. Mars life support systems may recycle water and oxygen using systems similar to those on the International Space Station (ISS). However, achieving sufficient reliability is less difficult for ISS than it will be for Mars. If an ISS system has a serious failure, it is possible to provide spare parts, or directly supply water or oxygen, or if necessary bring the crew back to Earth. Life support for Mars must be designed, tested, and improved as needed to achieve high demonstrated reliability. A quantitative reliability goal should be established and used to guide development t. The designers should select reliable components and minimize interface and integration problems. In theory a system can achieve the component-limited reliability, but testing often reveal unexpected failures due to design mistakes or flawed components. Testing should extend long enough to detect any unexpected failure modes and to verify the expected reliability. Iterated redesign and retest may be required to achieve the reliability goal. If the reliability is less than required, it may be improved by providing spare components or redundant systems. The number of spares required to achieve a given reliability goal depends on the component failure rate. If the failure rate is under estimated, the number of spares will be insufficient and the system may fail. If the design is likely to have undiscovered design or component problems, it is advisable to use dissimilar redundancy, even though this multiplies the design and development cost. In the ideal case, a human tended closed system operational test should be conducted to gain confidence in operations, maintenance, and repair. The difficulty in achieving high reliability in unproven complex systems may require the use of simpler, more mature, intrinsically higher reliability systems. The limitations of budget, schedule, and technology may suggest accepting lower and less certain expected reliability. A plan to develop reliable life support is needed to achieve the best possible reliability.

life support↗

A Misperception in Reliability Growth Modelling

The Duane reliability growth model is n(t)/t = k t^-alpha (1) The reliability growth rate is alpha, the downward slope of n(t)/t versus t. It usually varies from 0.2 to 0.6. k is a constant. Crow used a 56-failure data set to illustrate reliability growth.1 A graphical Duane model fit to this data gives n(t)/t = 0.640 t^-0.283 (2) A problem in using the Duane-Crow reliability growth model is that it assumes that reliability growth continues and the failure rate decreases throughout the test period. It is more usual that reliability growth stops when the failure rated is low enough. Growth testing is often followed by testing with a low constant failure rate due to rare or uncorrectable failure modes. As more and more low constant rate acceptable failures accumulate after the period of reliability growth, the reliability growth time exponent alpha decreases toward zero. This occurs if constant rate failures are treated as occurring during the reliability growth period. It is more accurate to model a period of initial reliability growth followed by testing without repair to more accurately determine the final constant failure rate. This is done in the abcd model. n(t)/t = a t^-b + c from t = 0 to td (3) = c + d after td, where d = a td^-b (4) The term a t^-b describes the continuous reliability growth that continues out to time td and c is the constant uncorrected failure rate. The parameter d represents an additional constant failure rate due to correctable but uncorrected failure modes. After the reliability growth process is terminated, the failure rate n(t)/t = c + d.

Harry W Jones↗

Accounting for Epistemic Uncertainty in Mission Supportability Assessment: A Necessary Step in Understanding Risk and Logistics Requirements

Future crewed missions to Mars present a maintenance logistics challenge that is unprecedented in human spaceflight. Mission endurance – defined as the time between resupply opportunities – will be significantly longer than previous missions, and therefore logistics planning horizons are longer and the impact of uncertainty is magnified. Maintenance logistics forecasting typically assumes that component failure rates are deterministically known and uses them to represent aleatory uncertainty, or uncertainty that is inherent to the process being examined. However, failure rates cannot be directly measured; rather, they are estimated based on similarity to other components or statistical analysis of observed failures. As a result, epistemic uncertainty – that is, uncertainty in knowledge of the process – exists in failure rate estimates that must be accounted for. Analyses that neglect epistemic uncertainty tend to significantly underestimate risk. Epistemic uncertainty can be reduced via operational experience; for example, the International Space Station (ISS) failure rate estimates are refined using a Bayesian update process. However, design changes may re-introduce epistemic uncertainty. Thus, there is a tradeoff between changing a design to reduce failure rates and operating a fixed design to reduce uncertainty. This paper examines the impact of epistemic uncertainty on maintenance logistics requirements for future Mars missions, using data from the ISS Environmental Control and Life Support System (ECLS) as a baseline for a case study. Sensitivity analyses are performed to investigate the impact of variations in failure rate estimates and epistemic uncertainty on spares mass. The results of these analyses and their implications for future system design and mission planning are discussed.

Owens, Andrew↗

Facilitating Data Collection of Maintenance Events to Populate the Hydrogen Component Reliability Database (HyCReD)

The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Correlation Between Weather Alerts and Grid Component Failures for Grid Alert

Weather events cause most grid failures. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Recent research at Idaho National Laboratory into electric grid risk analysis methods resulted in a tool that allows for the development of the most likely scenarios given failure probabilities of grid components. INL has a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response (CESER) program to develop a Grid Alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios uses MASTERRI and then notifies the utility if there is significant risk. Historical failure data of elements that comprise the U.S. electric grid have been compiled by utilities and organizations such as the international regulatory body North American Electric Reliability Corporation (NERC). Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates are needed for different component types given the alert type, severity, and location. Historic weather-related grid element failures are correlated with historic weather events from Integrated Public Alert & Warning System (IPAWS). These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. This discusses the Grid Alert project plan but focuses on the data gathered and process used in determining failure rates for possible grid failure scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Using the System Complexity Metric (SCM) to Compare CO2 Removal Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide removal systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide removal systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Using the System Complexity Metric (SCM) to Compare CO2 Reduction Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide reduction systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide reduction systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach is developed for the case when the system failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, N redundant units can be used, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is Ffail = F1 N , and the needed redundancy N = LN(F)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. (When F1 = f * L << 1, F1 is the probability that a unit will fail during the mission. When F1 = f * L > 1, F1 is the expected number of failures during the mission.) The needed redundancy, N, to achieve the required N redundant unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - is equal to the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some probabilistic uncertainty, the actual failure rate will be randomly higher or lower. This means that the reliability of the N redundant systems will be overestimated about half the time. Adding more redundant units increases the confidence that the required reliability will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals will be achieved with higher confidence. Both the desired reliability and confidence can be specified as initial requirements and the needed number of redundant units estimated using the measured failure rate.

Redundancy↗