Search NASA⌕ Search

Engineering topics

Harry W. Jones

Publications and source records attributed to Harry W. Jones.

The System Complexity Metric (SCM) Predicts System Costs and Failure Rates

A complex system has many parts and interactions and so is difficult to understand. Systems with higher complexity generally have higher costs and failure rates. A System Complexity Metric (SCM) is defined to be the sum of the number of nodes, N, in the system block diagram plus the number of one-way interactions, I, between the nodes. SCM = N + I. SCMs are easily determined by direct inspection of high level block diagrams of life support systems. System cost was found to be directly proportional to SCM. The system MTBF (Mean Time Before Failure) is the inverse of the system failure rate. MTBF = 1/f. The system MTBF was found to be proportional to SCM^(-2.2) for estimated preflight MTBFs. As is typical for systems that are not extensively tested and redesigned to eliminate unexpected failure modes, the life support flight failure rates were about ten times higher than the preflight estimates and the MTBFs one-tenth the preflight estimates. The system MTBF was found to be proportional to SCM^(-2.6) for observed flight MTBFs.

System compleity↗

A New Vision and A More Rational Process Will Advance Space Life Support Through Better Projects

Many space life support assumptions are obviously mistaken but they are generally accepted because they mutually support each other and because they advance a mistaken futuristic vision of closed human ecosystems in space. Many projects have been mistakenly selected to increase life support system closure or reduce launch mass in order to implement this impractical vision. Project members rarely challenge their project’s assumptions and guiding vision. Systems engineering including cost, reliability, risk, and system trade-offs has been very strongly discouraged to prevent rational criticism of these dubious assumptions, the projects, and the ecosystem vision. A more realistic vision can define better near term life support goals and future aspirations that will improve mission support and increase public interest. Given more realistic goals, it will be possible to develop effective technical alternatives, use open rational project selection methods, conduct professional project management, and make swifter progress in the inevitable human advance into space. The current unworkable human space settlement scenario involves moon visits and permanent bases, zero gravity trips to Mars, Mars visits and bases, all leading to extensive human settlements on the moon and Mars. This is unrealistic because the moon and Mars have less than Earth normal gravity and high radiation that can seriously damage human health. A better alternative would use rotating space colonies that produce Earth gravity and are radiation shielded. The attractive fantasy of reproducing terrestrial ecology in space should be replaced by the idea of building a closed system based on advancing technology such as artificial photosynthesis. Past mistaken projects include growing plants for food to replicate terrestrial food sources, recovering small amounts of water from difficult to process waste to achieve high system closure, and using newly designed, lighter, but unreliable hardware to reduce launch mass. The political process supporting the false assumptions, projects, and vision has deliberately prevented the usual systems engineering and cost-benefit analysis. We are going to do much better.

Harry W. Jones↗

The Effective Use of Metrics in Space Life Support System Trade-Offs

Engineering metrics are useful in space life support technology selection, but they must be carefully used. Metrics are only part of a complete system trade-off. Metrics do harm if they cause neglect of other important technical, organizational, or intuitive decision factors. Two metrics have damaged space life support, closure and Equivalent Systems Mass (ESM). Closure measures the fraction of the required system inputs that are produced by recycling system outputs. Increasing closure produces diminishing returns and becomes increasingly expensive. Increasing closure does not directly contribute to providing better life support. ESM measures the total launch mass required to provide life support. ESM includes the mass of the system hardware and of its power, cooling, pressurized volume, spares, and logistics. ESM predicts launch costs, but recently launch costs have been reduced by a factor of 20 or more. System development cost for space hardware is often much greater than launch cost. The past nearly exclusive use of ESM has led to the neglect of Life Cycle Cost (LCC), reliability, cost, and the other engineering factors. Closure and ESM have misguided space life support technology selection for more than twenty years and have adversely affected the expenditure of 100’s of millions of dollars. Metrics can be effectively used three ways in space life support technology selection: 1. A small set of key engineering metrics for preliminary screening. 2. A full set of engineering to guide technical selection. 3. Combining engineering metrics with organizational, political, and intuitive decision factors to understand technology selection. The past emphasis on closure and ESM served to support recycling life support over resupply and built on the intuitive appeal of a human ecosystem in space.

Harry W. Jones↗

Using the System Complexity Metric (SCM) to Compare CO2 Removal Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide removal systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide removal systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Lessons Learned in Space Life Support System Testing

The earlier problems can be found and corrected, the easier and cheaper it is to fix them. Doing less testing saves cost and time but doing too little testing increases the risk of operational failures causing large costs and delays. Integrated test is necessary to determine if the subsystems work together and the overall architecture performs as intended. This report reviews the testing lessons learned from the NASA Systems Engineering Handbook, a National Research Council report, and five reviews of International Space Station (ISS) lessons learned. The five reviews all mention two important points. First, that testing should be performed on the final integrated system, one as close as possible to the intended flight system. Second, “test as you fly,” while operating as planned in an environment as close as possible to the expected flight environment. Other lessons are the need for extensive preflight ground testing, the need to establish and defend an adequate budget, the problems using protoflight hardware on ISS, and the benefit of having ISS as a zero gravity test bed. The major ISS life support systems, carbon dioxide, water recycling, and oxygen recovery, were protoflight systems with little testing before launch to ISS. The failure rates these systems have been much greater than predicted and this has caused dissatisfaction with the protoflight approach. The more costly traditional approach is building qualification and test units in addition to flight units. The test units are used to test, analyze, and fix failure modes. Other work shows that there is an optimum cost-effective intuitive appeal of a human ecosystem in space.

Life support↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence on a Time Limited Mission?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach applies to the case where the unit failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, the design then uses N redundant units, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is FN = F1N, and the needed N = LN(FN)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. The needed redundancy, N, to achieve the specified N unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - achieves the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some uncertainty, the confidence that the rate is not lower than the actual failure rate and the required reliability is not overestimated is about 50%. Adding more redundant units increases the confidence that the required reliability, FN, will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals have higher confidence in being achieved. Both the required reliability and confidence can be specified initially and the needed number of redundant units computed using the measured failure rate. The unit failure rate is determined by initial reliability growth testing to remove design errors and to better estimate the final constant failure rate. Reducing the failure rate and reducing its variance both reduce the number of redundant units needed for the required reliability and confidence. Since the total cost is the sum of the costs of the units and of the testing, there is an optimum test time that produces minimum cost.

Harry W. Jones↗

Going Beyond Reliability to Robustness and Resilience in Space Life Support Systems

The words reliability, robustness, and resilience are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environment is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to unanticipated damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes can invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging conditions. These three qualities are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from unanticipated disruptions. Increasing the resilience of space systems would greatly increase space crew safety. Improving reliability and robustness requires dealing with known problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. The need for robustness and resilience has been stated for decades but little has been done. Systems designers often assume that they understand everything they need to know. The potential failures caused by changes, failures, accidents, unknown environments, and unknown unknowns can be ignored. Such overconfidence can lead to neglect of reliability, robustness, and resilience.

Harry W. Jones↗

The Challenger tragedy was caused by an Apollo mistake, terminating risk analysis

NASA’s view of risk changed between early Apollo and the Space Shuttle. Risk was a known serious problem at the beginning of Apollo and the risk estimates were disturbingly high. To avoid public concern, risk analysis was discontinued. Risk analysis was avoided in Shuttle, leading to an unnecessarily risky design. The immediate cause of the Challenger tragedy was the mistaken decision to launch in cold weather. The fundamental cause was the high risk of the Shuttle design. Before Challenger, management thought and testified that the probability of an accident was 1 in 100,000. After Challenger, Probabilistic Risk Analysis (PRA) found a roughly 1 in 100 chance of a Shuttle failure. The recent Orion design uses the safer Apollo approach, with a hardened capsule, launch abort escape, and the crew placed above the rocket tanks and engines. During Apollo it was estimated that, “assuming all elements from propulsion to rendezvous and life support were done as well or better than ever before, that 30 astronauts would be lost before 3 were returned safely to the Earth.” The chance of astronaut survival was only 10%. After the Apollo 1 tragedy, the awareness of risk led to an intense focus on achieving safety. “The only possible explanation for the astonishing success – no losses in space and on time – was that every participant at every level in every area far exceeded the norm of human capabilities.” During Apollo, a NASA PRA found that the chance of success was “less than 5 percent.” The NASA Administrator felt that “the numbers could do irreparable harm,” and discontinued numerical risk assessment. This led to decreasing understanding of risk. The head of Apollo reliability and safety decided, “Statistics don’t count for anything,” and that risk is reduced by “attention taken in design.” The great and initially unexpected success of Apollo appeared to validate the neglect of PRA. Continuing to neglect the mathematical estimation of risk led Shuttle into a high risk design that produced tragic results. The initial design of the Shuttle emphasized increasing capability and reducing cost without analysis or even mention of risk. A retired NASA official stated, “some NASA people began to confuse desire with reality. … One result was to assess risk in terms of what was thought acceptable without regard for verifying the assessment. … Note that under such circumstances real risk management is shut out.” Not computing risk led to removing launch abort, removing crew escape, selecting less reliable Solid Rocket Boosters, placing the crew compartment next to the rocket boosters, and accepting more stressed shielding tile designs. Accepting these specific risks directly caused the shuttle disasters. The Challenger tragedy is frequently taught as a case of management failure. The focus is on the Challenger launch decision hours before, which is a dramatic example of bad management. However, the true cause of the Challenger disaster occurred decades earlier in the Apollo era. When the easily predictable failures occurred, failure investigations focused on how they might have been avoided. The Shuttle was cancelled after the space station was completed because of its high risk. The ultimate cause of the Shuttle tragedies was the choice by the Apollo-era NASA administrator to avoid a negative public reaction to realistic risk analysis.

Harry W. Jones↗

Using the System Complexity Metric (SCM) to Compare CO2 Reduction Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide reduction systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide reduction systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Lessons Learned in Space Life Support System Testing

The earlier problems can be found and corrected, the easier and cheaper it is to fix them. Doing less testing saves cost and time but doing too little testing increases the risk of operational failures causing large costs and delays. Integrated test is necessary to determine if the subsystems work together and the overall architecture performs as intended. This report reviews the testing lessons learned from the NASA Systems Engineering Handbook, a National Research Council report, and five reviews of International Space Station (ISS) lessons learned. The five reviews all mention two important points. First, that testing should be performed on the final integrated system, one as close as possible to the intended flight system. Second, “test as you fly,” while operating as planned in an environment as close as possible to the expected flight environment. Other lessons are the need for extensive preflight ground testing, the need to establish and defend an adequate budget, the problems using protoflight hardware on ISS, and the benefit of having ISS as a zero gravity test bed. The major ISS life support systems, carbon dioxide removal, water recycling, and oxygen recovery, were protoflight systems with little testing before launch to ISS. The failure rates of these systems have been much greater than predicted and this has caused dissatisfaction with the protoflight approach. The more costly traditional approach builds qualification and test units in addition to flight units. The test units are used to find, analyze, and fix failure modes. Other work shows that there is an optimum cost-effective amount of testing when redundant systems must have a specified reliability and confidence.

Harry W. Jones↗

Going Beyond Reliability to Robustness and Resilience in Space Life Support Systems

The words reliability, robustness, and resilience are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environment is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to unanticipated damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes can invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging conditions. These three qualities are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from unanticipated disruptions. Increasing the resilience of space systems would greatly increase space crew safety. Improving reliability and robustness requires dealing with known problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. The need for robustness and resilience has been stated for decades but little has been done. Systems designers often assume that they understand everything they need to know. The potential failures caused by changes, failures, accidents, unknown environments, and unknown unknowns can be ignored. Such overconfidence can lead to neglect of reliability, robustness, and resilience.

Harry W. Jones↗

The Partial Gravity of the Moon and Mars Appears Insufficient to Maintain Human Health

Astronauts who spend many weeks or months in space in microgravity suffer serious health problems including muscle atrophy, cardiovascular deconditioning, bone calcium loss, impaired vision, and immune system changes. The debilitating effects of weightlessness were first demonstrated on the early Skylab, Salyut, and Mir missions, but it was then hoped that countermeasures including in-flight exercise and resistance training could reduce most of these problems. Similar effects are anticipated in the partial gravity of the Moon and Mars. Direct evidence of the long-term effects of partial gravity on humans is not yet available, but indirect evidence suggests that partial gravity exposure below 0.4 g will be insufficient to maintain musculoskeletal and cardiopulmonary conditioning over the long term. Some studies show a strong correlation between heart rate, oxygen consumption, net metabolic rate, and simulated gravity from 0 to 1 g. Exposure to moon and Mars gravities will probably cause less severe physiological deconditioning than microgravity, but the benefit of partial gravity seems likely to be roughly proportional to the level of gravity experienced. As in microgravity, exercise countermeasures seem useful but insufficient to preserve all physiological systems as they would be in Earth gravity.

Harry W. Jones↗

The System Complexity Metric (SCM) Explains Systems Design and is Correlated with Cost and Failure Rate

The human short term memory span and working capacity is limited to three to five items, especially if they are organized complex “chunks” of information. The impression of complexity occurs when a system is simply difficult to understand, where there is no apparent pattern to predict its behavior. Hierarchical systems design can reduce perceived complexity and increase the amount of information that can be managed. The SCM was developed to measure complexity and help compare proposed overall system architectures before detailed design information is available. The SCM is defined as the sum of the number of major nodes, N, in the system block diagram plus the number of one-way interactions, I, between the nodes. SCM = N + I. SCM’s are easily determined by direct inspection of high-level block diagrams of life support systems. Axiomatic design develops a hierarchy of subsystem requirements and designs together in a top-down, back-and-forth process. A coupling matrix is used to control the relationships between the subsystem functions and design concepts. Axiomatic design can improve system design by decoupling requirements and designs. Axiomatic design was applied to the planning of a closed life support system, similar to that used on the International Space Station. A materially open as opposed to a closed system design was created by removing the interconnections required to close the system. The open system had the same number of designed subsystems as the closed system, but it had many fewer interconnections and its SCM was lower by about half. The costs were estimated and the MTBF (Mean Time Before Failure) tabulated for open and closed space life support systems. The estimated costs were linearly proportional to SCM for the wide variations of SCM in life support, but small differences may not be significant. The flight and preflight MTBF’s both declined exponentially with increasing MTBF, faster than MTBF-2, even though the preflight estimated MTBF’s were about ten times higher than the flight MTBF’s.

System Complexity Metric (SCM)↗

Long Term Human Presence in Space Requires Artificial Gravity and Radiation Shielding

Astronauts who spend many months in microgravity suffer serious health problems including muscle atrophy, cardiovascular deconditioning, bone calcium loss, impaired vision, and immune system changes. Exercise countermeasures have been insufficient to maintain normal human performance. Similar problems can be expected in the partial gravity of the Moon and Mars. Achieving the long-term presence of healthy humans in space requires providing artificial Earth level gravity. This can be done on the Moon and Mars by using horizontally rotating habitats with angled floors, but it is easier in space habitats. Astronauts travelling beyond the protection of the Earth’s magnetic field can suffer harm from cosmic background radiation and occasional strong solar flares. Supporting healthy long-term human lives will require radiation shielding on the Moon and Mars as well as in space. Human space settlement will probably begin with artificial rotating space habitats in Low Earth Orbit (LEO) where they will be shielded from radiation. The earlier anticipated human communities in pressurized domes on the Moon or Mars appear unrealistic because of the now known problems of partial gravity and radiation.

Astronauts↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗

Take Material to Space or Make It There?

Most human missions in space have been brief, lasting only days or weeks, and they have taken all the materials they need. Using the alternate method, the International Space Station recycles water and oxygen, and it is often assumed that future Moon and Mars missions should also recycle their life support materials. The “take or make” decision is primarily based on cost, and making or recycling material on longer space missions can sometimes be less expensive than taking it. Longer missions favor recycling over resupply when the initial cost to provide recycling equipment is less than the cost to provide resupply. Recycling was clearly cheaper than taking material to the space station in the space shuttle era when launch costs were very high. Launch costs have decreased and the take or make decision point has changed. The recent reduction in launch cost by a factor of about twenty-five to fifty makes taking material cost less than making or recycling it for much longer missions. The cost breakeven point when making rather than taking material is less expensive is now much farther out in time than before. Material recycling or in situ production no longer saves cost except for very large or very long missions.

Space manufacturing↗