Search NASA⌕ Search

SEARCH · Search NASA

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A Methodology for Quantifying Certain Design Requirements During the Design Phase

A methodology for developing and balancing quantitative design requirements for safety, reliability, and maintainability has been proposed. Conceived as the basis of a more rational approach to the design of spacecraft, the methodology would also be applicable to the design of automobiles, washing machines, television receivers, or almost any other commercial product. Heretofore, it has been common practice to start by determining the requirements for reliability of elements of a spacecraft or other system to ensure a given design life for the system. Next, safety requirements are determined by assessing the total reliability of the system and adding redundant components and subsystems necessary to attain safety goals. As thus described, common practice leaves the maintainability burden to fall to chance; therefore, there is no control of recurring costs or of the responsiveness of the system. The means that have been used in assessing maintainability have been oriented toward determining the logistical sparing of components so that the components are available when needed. The process established for developing and balancing quantitative requirements for safety (S), reliability (R), and maintainability (M) derives and integrates NASA s top-level safety requirements and the controls needed to obtain program key objectives for safety and recurring cost (see figure). Being quantitative, the process conveniently uses common mathematical models. Even though the process is shown as being worked from the top down, it can also be worked from the bottom up. This process uses three math models: (1) the binomial distribution (greaterthan- or-equal-to case), (2) reliability for a series system, and (3) the Poisson distribution (less-than-or-equal-to case). The zero-fail case for the binomial distribution approximates the commonly known exponential distribution or "constant failure rate" distribution. Either model can be used. The binomial distribution was selected for modeling flexibility because it conveniently addresses both the zero-fail and failure cases. The failure case is typically used for unmanned spacecraft as with missiles.

Adams, Timothy↗

Reliability of High-Voltage Tantalum Capacitors. Parts 3 and 4)

Weibull grading test is a powerful technique that allows selection and reliability rating of solid tantalum capacitors for military and space applications. However, inaccuracies in the existing method and non-adequate acceleration factors can result in significant, up to three orders of magnitude, errors in the calculated failure rate of capacitors. This paper analyzes deficiencies of the existing technique and recommends more accurate method of calculations. A physical model presenting failures of tantalum capacitors as time-dependent-dielectric-breakdown is used to determine voltage and temperature acceleration factors and select adequate Weibull grading test conditions. This model is verified by highly accelerated life testing (HALT) at different temperature and voltage conditions for three types of solid chip tantalum capacitors. It is shown that parameters of the model and acceleration factors can be calculated using a general log-linear relationship for the characteristic life with two stress levels.

Teverovsky, Alexander↗

Analysis of Weibull Grading Test for Solid Tantalum Capacitors

Weibull grading test is a powerful technique that allows selection and reliability rating of solid tantalum capacitors for military and space applications. However, inaccuracies in the existing method and non-adequate acceleration factors can result in significant, up to three orders of magnitude, errors in the calculated failure rate of capacitors. This paper analyzes deficiencies of the existing technique and recommends more accurate method of calculations. A physical model presenting failures of tantalum capacitors as time-dependent-dielectric-breakdown is used to determine voltage and temperature acceleration factors and select adequate Weibull grading test conditions. This, model is verified by highly accelerated life testing (HALT) at different temperature and voltage conditions for three types of solid chip tantalum capacitors. It is shown that parameters of the model and acceleration factors can be calculated using a general log-linear relationship for the characteristic life with two stress levels.

Teverovsky, Alexander↗

A Nuclear Interaction Model for Understanding Results of Single Event Testing with High Energy Protons

An internuclear cascade and evaporation model has been adapted to estimate the LET spectrum generated during testing with 200 MeV protons. The model-generated heavy ion LET spectrum is compared to the heavy ion LET spectrum seen on orbit. This comparison is the basis for predicting single event failure rates from heavy ions using results from a single proton test. Of equal importance, this spectra comparison also establishes an estimate of the risk of encountering a failure mode on orbit that was not detected during proton testing. Verification of the general results of the model is presented based on experiments, individual part test results, and flight data. Acceptance of this model and its estimate of remaining risk opens the hardware verification philosophy to the consideration of radiation testing with high energy protons at the board and box level instead of the more standard method of individual part testing with low energy heavy ions.

Culpepper, William X.↗

J-2X Abort System Development

The J-2X is an expendable liquid hydrogen (LH2)/liquid oxygen (LOX) gas generator cycle rocket engine that is currently being designed as the primary upper stage propulsion element for the new NASA Ares vehicle family. The J-2X engine will contain abort logic that functions as an integral component of the Ares vehicle abort system. This system is responsible for detecting and responding to conditions indicative of impending Loss of Mission (LOM), Loss of Vehicle (LOV), and/or catastrophic Loss of Crew (LOC) failure events. As an earth orbit ascent phase engine, the J-2X is a high power density propulsion element with non-negligible risk of fast propagation rate failures that can quickly lead to LOM, LOV, and/or LOC events. Aggressive reliability requirements for manned Ares missions and the risk of fast propagating J-2X failures dictate the need for on-engine abort condition monitoring and autonomous response capability as well as traditional abort agents such as the vehicle computer, flight crew, and ground control not located on the engine. This paper describes the baseline J-2X abort subsystem concept of operations, as well as the development process for this subsystem. A strategy that leverages heritage system experience and responds to an evolving engine design as well as J-2X specific test data to support abort system development is described. The utilization of performance and failure simulation models to support abort system sensor selection, failure detectability and discrimination studies, decision threshold definition, and abort system performance verification and validation is outlined. The basis for abort false positive and false negative performance constraints is described. Development challenges associated with information shortfalls in the design cycle, abort condition coverage and response assessment, engine-vehicle interface definition, and abort system performance verification and validation are also discussed.

Santi, Louis M.↗

A unified method for evaluating real-time computer controllers: A case study

A real time control system consists of a synergistic pair, that is, a controlled process and a controller computer. Performance measures for real time controller computers are defined on the basis of the nature of this synergistic pair. A case study of a typical critical controlled process is presented in the context of new performance measures that express the performance of both controlled processes and real time controllers (taken as a unit) on the basis of a single variable: controller response time. Controller response time is a function of current system state, system failure rate, electrical and/or magnetic interference, etc., and is therefore a random variable. Control overhead is expressed as a monotonically nondecreasing function of the response time and the system suffers catastrophic failure, or dynamic failure, if the response time for a control task exceeds the corresponding system hard deadline, if any. A rigorous probabilistic approach is used to estimate the performance measures. The controlled process chosen for study is an aircraft in the final stages of descent, just prior to landing. First, the performance measures for the controller are presented. Secondly, control algorithms for solving the landing problem are discussed and finally the impact of the performance measures on the problem is analyzed.

Shin, K. G.↗

A unified method for evaluating real-time computer controllers and its application

A real time control system consists of a synergistic pair, that is, a controlled process and a controller computer. Performance measures for real time controller computers are defined on the basis of the nature of this synergistic pair. A case study of a typical critical controlled process is presented in the context of new performance measures that express the performance of both controlled processes and real time controllers (taken as a unit) on the basis of a single variable: controller response time. Controller response time is a function of current system state, system failure rate, electrical and/or magnetic interference, etc., and is therefore a random variable. Control overhead is expressed as a monotonically nondecreasing function of the response time and the system suffers catastrophic failure, or dynamic failure, if the response time for a control task exceeds the corresponding system hard deadline, if any. A rigorous probabilistic approach is used to estimate the performance measures. The controlled process chosen for study is an aircraft in the final stages of descent, just prior to landing. First, the performance measures for the controller are presented. Secondly, control algorithms for solving the landing problem are discussed and finally the impact of the performance measures on the problem is analyzed.

Shin, K. G.↗

Synthesizing a New Launch Vehicle Failure Probability Based on Historical Flight Data

New launch vehicles have historically had significantly higher failure rates in early flights than what has been predicted using Probabilistic Risk Assessment - PRA. This is because PRAs typically model a mature vehicle where a significant portion of the early failure probability contributors have been eliminated due to testing and improvements after actual field operation. To capture a more accurate early flight failure probability estimate, this paper develops a method that estimates ascent failure probability starting with the first flight based on historical launch vehicle records. With new launch vehicles being developed, such as the Space Launch System - SLS, a PRA model must be extended to cover early flight failure probability contributions that are either not covered in the mature-vehicle PRA or are underestimated. These failure probability contributions include design errors, quality control deficiencies, installation errors, and environmental impacts. There are also failure dependencies due to systemic errors that still exist due to limited entire-system testing.

Cross, Robert B.↗

Motivating the sure bounds

Motivation is provided for a theorem that provides upper and lower bounds for the reliability of reconfigurable digital control systems. The reliability goals for these systems are too high to be established by natural life testing, which means the probability of system failure must be computed from mathematical models that capture the essential elements of fault occurence and system fault recovery. The upper and lower bound theorem shows that system recovery can be adequately described by its first two moments, provided component failure rate is low and system recovery is fast. This result greatly simplifies both the fault injection experiments that study system recovery and the numerical computations that estimate the probability of system failure from a mathematical model.

White, Allan L.↗

Modelling early failures in Space Station Freedom

A major problem encountered in planning for Space Station Freedom is the amount of maintenance that will be required. To predict the failure rates of components and systems aboard Space Station Freedom, the logical approach is to use data obtained from previously flown spacecraft. In order to determine the mechanisms that are driving the failures, models can be proposed, and then checked to see if they adequately fit the observed failure data obtained from a large variety of satellites. For this particular study, failure data and truncation times were available for satellites launched between 1976 and 1984; no data past 1984 was available. The study was limited to electrical subsystems and assemblies, which were studied to determine if they followed a model resulting from a mixture of exponential distributions.

Navard, Sharon E.↗

The Threat of Uncertainty: Why Using Traditional Approaches for Evaluating Spacecraft Reliability are Insufficient for Future Human Mars Missions

Through the Evolvable Mars Campaign (EMC) study, the National Aeronautics and Space Administration (NASA) continues to evaluate potential approaches for sending humans beyond low Earth orbit (LEO). A key aspect of these missions is the strategy that is employed to maintain and repair the spacecraft systems, ensuring that they continue to function and support the crew. Long duration missions beyond LEO present unique and severe maintainability challenges due to a variety of factors, including: limited to no opportunities for resupply, the distance from Earth, mass and volume constraints of spacecraft, high sensitivity of transportation element designs to variation in mass, the lack of abort opportunities to Earth, limited hardware heritage information, and the operation of human-rated systems in a radiation environment with little to no experience. The current approach to maintainability, as implemented on ISS, which includes a large number of spares pre-positioned on ISS, a larger supply sitting on Earth waiting to be flown to ISS, and an on demand delivery of logistics from Earth, is not feasible for future deep space human missions. For missions beyond LEO, significant modifications to the maintainability approach will be required.Through the EMC evaluations, several key findings related to the reliability and safety of the Mars spacecraft have been made. The nature of random and induced failures presents significant issues for deep space missions. Because spare parts cannot be flown as needed for Mars missions, all required spares must be flown with the mission or pre-positioned. These spares must cover all anticipated failure modes and provide a level of overall reliability and safety that is satisfactory for human missions. This will require a large amount of mass and volume be dedicated to storage and transport of spares for the mission. Further, there is, and will continue to be, a significant amount of uncertainty regarding failure rates for spacecraft components. This uncertainty makes it much more difficult to anticipate failures and will potentially require an even larger amount of spares to provide an acceptable level of safety. Ultimately, the approach to maintenance and repair applied to ISS, focusing on the supply of spare parts, may not be tenable for deep space missions. Other approaches, such as commonality of components, simplification of systems, and in-situ manufacturing will be required.

Stromgren, Chel↗

Investigation of mercury thruster isolators

Mercury ion thruster isolator lifetime tests were performed using different isolator materials and geometries. Tests were performed with and without the flow of mercury through the isolators in an oil diffusion pumped vacuum facility and cryogenically pumped bell jar. The onset of leakage current in isolators occurred in time intervals ranging from a few hours to many hundreds of hours. In all cases, surface contamination was responsible for the onset of leakage current and subsequent isolator failure. Rate of increase of leakage current and the leakage current level increased approximately exponentially with isolator temperature. Careful attention to shielding techniques and the elimination of sources of metal oxides appear to have eliminated isolator failures as a thruster life limiting mechanism.

Mantenieks, M. A.↗

Investigation of mercury thruster isolators

Mercury ion thruster isolator lifetime tests were performed using different isolator materials and geometries. Tests were performed with and without the flow of mercury through the isolators in an oil diffusion pumped vacuum facility and cryogenically pumped bell jar. The onset of leakage current in isolators tested occurred in time intervals ranging from a few hours to many hundreds of hours. In all cases, surface contamination was responsible for the onset of leakage current and subsequent isolator failure. Rate of increase of leakage current and the leakage current level increased approximately exponentially with isolator temperature. Careful attention to shielding techniques and the elimination of sources of metal oxides appear to have eliminated isolator failures as a thruster life limiting mechanism.

Mantenieks, M. A.↗

Designing to Mitigate Food Growing Failures in Space

Future space life support systems may use crop plants to grow most of the crew s food. A harvest failure can reduce the food available for future consumption. If the previously stored food is insufficient to reach the next harvest, the crew may go hungry. This paper considers how the overall food supply system should be modified to cope with food production failures. The food supply concept for a mission will use grown food, or stored food, cIr both. The optimum food supply mix depends on the costs and failure probabilities of stored and grown food. A simple food system model assumes that either we obtain the nominal harvest or a failure occurs and no food is harvested. Given the probability that any particular harvest fails, it is easy to compute the expected number of failures and the total food shortfall over a mission. If some food is grown and the probability of harvest failure is high, a non-redundant system has an unacceptable likelihood that the crew will have no food for a full harvest period. Food supply reliability must be increased either by supplying more food initially or by increasing food production capacity. We can obtain a very reliable food supply even when the harvest failure rate is high. If the cost of growing food is much less than the cost of providing stored food, it is better to provide redundant food growing capacity than to increase initial storage. A more realistic biomass production failure model allows the harvest amount or time to vary around the nominal values, using stochastic modeling with repeated Monte Carlo simulation, but such failures have minor impact compared to a complete harvest failure.

Jones, Harry↗

Reliability measurement during software development

During the development of data base software for a multi-sensor tracking system, reliability was measured. The failure ratio and failure rate were found to be consistent measures. Trend lines were established from these measurements that provided good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the individual run submission rather than with the code proper. Possible application of these findings for line management, project managers, functional management, and regulatory agencies is discussed. Steps for simplifying the measurement process and for use of these data in predicting operational software reliability are outlined.

Hecht, H.↗

The determination of measures of software reliability

Measurement of software reliability was carried out during the development of data base software for a multi-sensor tracking system. The failure ratio and failure rate were found to be consistent measures. Trend lines could be established from these measurements that provide good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the individual run submission rather than with the code proper. Possible application of these findings for line management, project managers, functional management, and regulatory agencies is discussed. Steps for simplifying the measurement process and for use of these data in predicting operational software reliability are outlined.

Maxwell, F. D.↗

Parts and Components Reliability Assessment: A Cost Effective Approach

System reliability assessment is a methodology which incorporates reliability analyses performed at parts and components level such as Reliability Prediction, Failure Modes and Effects Analysis (FMEA) and Fault Tree Analysis (FTA) to assess risks, perform design tradeoffs, and therefore, to ensure effective productivity and/or mission success. The system reliability is used to optimize the product design to accommodate today?s mandated budget, manpower, and schedule constraints. Stand ard based reliability assessment is an effective approach consisting of reliability predictions together with other reliability analyses for electronic, electrical, and electro-mechanical (EEE) complex parts and components of large systems based on failure rate estimates published by the United States (U.S.) military or commercial standards and handbooks. Many of these standards are globally accepted and recognized. The reliability assessment is especially useful during the initial stages when the system design is still in the development and hard failure data is not yet available or manufacturers are not contractually obliged by their customers to publish the reliability estimates/predictions for their parts and components. This paper presents a methodology to assess system reliability using parts and components reliability estimates to ensure effective productivity and/or mission success in an efficient manner, low cost, and tight schedule.

Lee, Lydia↗

Lessons Learned in Space Life Support System Testing

The earlier problems can be found and corrected, the easier and cheaper it is to fix them. Doing less testing saves cost and time but doing too little testing increases the risk of operational failures causing large costs and delays. Integrated test is necessary to determine if the subsystems work together and the overall architecture performs as intended. This report reviews the testing lessons learned from the NASA Systems Engineering Handbook, a National Research Council report, and five reviews of International Space Station (ISS) lessons learned. The five reviews all mention two important points. First, that testing should be performed on the final integrated system, one as close as possible to the intended flight system. Second, “test as you fly,” while operating as planned in an environment as close as possible to the expected flight environment. Other lessons are the need for extensive preflight ground testing, the need to establish and defend an adequate budget, the problems using protoflight hardware on ISS, and the benefit of having ISS as a zero gravity test bed. The major ISS life support systems, carbon dioxide, water recycling, and oxygen recovery, were protoflight systems with little testing before launch to ISS. The failure rates these systems have been much greater than predicted and this has caused dissatisfaction with the protoflight approach. The more costly traditional approach is building qualification and test units in addition to flight units. The test units are used to test, analyze, and fix failure modes. Other work shows that there is an optimum cost-effective intuitive appeal of a human ecosystem in space.

Life support↗