Search NASASearch

SEARCH · Search NASA

Results for “reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Addressing Uniqueness and Unison of Reliability and Safety for a Better Integration

Over time, it has been observed that Safety and Reliability have not been clearly differentiated, which leads to confusion, inefficiency, and, sometimes, counter-productive practices in executing each of these two disciplines. It is imperative to address this situation to help Reliability and Safety disciplines improve their effectiveness and efficiency. The paper poses an important question to address, "Safety and Reliability - Are they unique or unisonous?" To answer the question, the paper reviewed several most commonly used analyses from each of the disciplines, namely, FMEA, reliability allocation and prediction, reliability design involvement, system safety hazard analysis, Fault Tree Analysis, and Probabilistic Risk Assessment. The paper pointed out uniqueness and unison of Safety and Reliability in their respective roles, requirements, approaches, and tools, and presented some suggestions for enhancing and improving the individual disciplines, as well as promoting the integration of the two. The paper concludes that Safety and Reliability are unique, but compensating each other in many aspects, and need to be integrated. Particularly, the individual roles of Safety and Reliability need to be differentiated, that is, Safety is to ensure and assure the product meets safety requirements, goals, or desires, and Reliability is to ensure and assure maximum achievability of intended design functions. With the integration of Safety and Reliability, personnel can be shared, tools and analyses have to be integrated, and skill sets can be possessed by the same person with the purpose of providing the best value to a product development.

Huang, Zhaofeng

Reliability Models and Demonstration of a Fault-Tolerant Motor Concept for Vertical Takeoff and Landing Vehicles

This report documents the completion of the Revolutionary Vertical Lift Technology Project Annual Performance Indicator 24-3.2.4.1: “Apply and document reliability prediction for high reliability motor concept.” Two modeling tools were completed for calculation of reliability of fault-tolerant (FT) motors, and key FT operations of a modular FT motor were demonstrated experimentally. The two models are complementary tools for the stakeholder and user community. Both models employ Markov chain theory. The first model is a time-homogeneous Markov chain model, and the second is a time-inhomogeneous Markov-Weibull model. This report’s main sections are as follows: 1.0 Introduction, 2.0 Theory, 3.0 Motor Reliability Models, 4.0 Validation of FT Operation by Hardware Demonstration, and 5.0 Concluding Remarks. Novel contributions to the field include development of a modular FT motor concept for electrified vertical takeoff and landing (eVTOL) application, solution methods to solve the reliability calculations, development of figures of merit, and the introduction of “linked chains” to formulate a building-block approach for time-inhomogeneous Markov-Weibull modeling of motor reliability. Example case studies have been completed, and results are provided and discussed herein. A four-module FT motor concept was developed to a preliminary-design level of detail. This eVTOL FT motor concept was designed for galvanic, magnetic, and thermal isolation of stator winding faults. The reliability of the concept motor was calculated using a time-inhomogeneous Markov chain model. Employing average failure rate as a metric, 570 times greater reliability was achieved as compared to a baseline motor without fault tolerance. A demonstrator motor was built and tested. The testing demonstrated the key features of FT operation and validated the essential premises of the FT motor concepts presented herein. The experiments included successful demonstration of the feasibility of the following four key FT features: (1) terminal open-circuit operation, (2) thermal isolation after fault, (3) terminal short-circuit operation, and (4) internal short-circuit operation. These works indicate that FT modular motor drives offer promise for addressing the daunting reliability gap that electric aircraft propulsor drives are facing relative to the best conventional motor drive technology that is available today.

Electric Motor

Four Problematic Methods in Reliability Analysis

Some basic methods used in reliability analysis are problematic because they produce incorrect and overoptimistic predictions. Initially gratifying forecasts are often invalidated by testing and operational experience. The problematic methods in reliability analysis include estimating the system failure rate as the sum of component failure rates, assuming that reliability growth continues indefinitely during testing, overestimating the benefits of redundancy, and using the fault tolerance count instead of a detailed reliability analysis. Reliability analysis can produce more optimism than accuracy. This bug may now be a feature. The optimistic bias inevitable in project planning should be corrected by realistic reliability analysis that reflects relevant experience. That the repeated poor performance of reliability analysis is found to be surprising suggests willful blindness. Rigorous methods and impartial critical review are necessary to improve reliability analysis.

Reliability analysis

Four Problematic Methods in Reliability Analysis

Some basic methods used in reliability analysis are problematic because they produce incorrect and overoptimistic predictions. Initially gratifying forecasts are often invalidated by testing and operational experience. The problematic methods in reliability analysis include estimating the system failure rate as the sum of component failure rates, assuming that reliability growth continues indefinitely during testing, overestimating the benefits of redundancy, and using the fault tolerance count instead of a detailed reliability analysis. Reliability analysis can produce more optimism than accuracy. This bug may now be a feature. The optimistic bias inevitable in project planning should be corrected by realistic reliability analysis that reflects relevant experience. That the repeated poor performance of reliability analysis is found to be surprising suggests willful blindness. Rigorous methods and impartial critical review are necessary to improve reliability analysis.

Reliability analysis

A cost assessment of reliability requirements for shuttle-recoverable experiments

The relaunching of unsuccessful experiments or satellites will become a real option with the advent of the space shuttle. An examination was made of the cost effectiveness of relaxing reliability requirements for experiment hardware by allowing more than one flight of an experiment in the event of its failure. Any desired overall reliability or probability of mission success can be acquired by launching an experiment with less reliability two or more times if necessary. Although this procedure leads to uncertainty in total cost projections, because the number of flights is not known in advance, a considerable cost reduction can sometimes be achieved. In cases where reflight costs are low relative to the experiment's cost, three flights with overall reliability 0.9 can be made for less than half the cost of one flight with a reliability of 0.9. An example typical of shuttle payload cost projections is cited where three low reliability flights would cost less than $50 million and a single high reliability flight would cost over $100 million. The ratio of reflight cost to experiment cost is varied and its effect on the range in total cost is observed. An optimum design reliability selection criterion to minimize expected cost is proposed, and a simple graphical method of determining this reliability is demonstrated.

Campbell, J. W.

Recalibrating software reliability models

In spite of much research effort, there is no universally applicable software reliability growth model which can be trusted to give accurate predictions of reliability in all circumstances. Further, it is not even possible to decide a priori which of the many models is most suitable in a particular context. In an attempt to resolve this problem, techniques were developed whereby, for each program, the accuracy of various models can be analyzed. A user is thus enabled to select that model which is giving the most accurate reliability predictions for the particular program under examination. One of these ways of analyzing predictive accuracy, called the u-plot, in fact allows a user to estimate the relationship between the predicted reliability and the true reliability. It is shown how this can be used to improve reliability predictions in a completely general way by a process of recalibration. Simulation results show that the technique gives improved reliability predictions in a large proportion of cases. However, a user does not need to trust the efficacy of recalibration, since the new reliability estimates produced by the technique are truly predictive and so their accuracy in a particular application can be judged using the earlier methods. The generality of this approach would therefore suggest that it be applied as a matter of course whenever a software reliability model is used.

Brocklehurst, Sarah

Recalibrating software reliability models

In spite of much research effort, there is no universally applicable software reliability growth model which can be trusted to give accurate predictions of reliability in all circumstances. Further, it is not even possible to decide a priori which of the many models is most suitable in a particular context. In an attempt to resolve this problem, techniques were developed whereby, for each program, the accuracy of various models can be analyzed. A user is thus enabled to select that model which is giving the most accurate reliability predicitons for the particular program under examination. One of these ways of analyzing predictive accuracy, called the u-plot, in fact allows a user to estimate the relationship between the predicted reliability and the true reliability. It is shown how this can be used to improve reliability predictions in a completely general way by a process of recalibration. Simulation results show that the technique gives improved reliability predictions in a large proportion of cases. However, a user does not need to trust the efficacy of recalibration, since the new reliability estimates prodcued by the technique are truly predictive and so their accuracy in a particular application can be judged using the earlier methods. The generality of this approach would therefore suggest that it be applied as a matter of course whenever a software reliability model is used.

Brocklehurst, Sarah

Reliability Modeling of Microelectromechanical Systems Using Neural Networks

Microelectromechanical systems (MEMS) are a broad and rapidly expanding field that is currently receiving a great deal of attention because of the potential to significantly improve the ability to sense, analyze, and control a variety of processes, such as heating and ventilation systems, automobiles, medicine, aeronautical flight, military surveillance, weather forecasting, and space exploration. MEMS are very small and are a blend of electrical and mechanical components, with electrical and mechanical systems on one chip. This research establishes reliability estimation and prediction for MEMS devices at the conceptual design phase using neural networks. At the conceptual design phase, before devices are built and tested, traditional methods of quantifying reliability are inadequate because the device is not in existence and cannot be tested to establish the reliability distributions. A novel approach using neural networks is created to predict the overall reliability of a MEMS device based on its components and each component's attributes. The methodology begins with collecting attribute data (fabrication process, physical specifications, operating environment, property characteristics, packaging, etc.) and reliability data for many types of microengines. The data are partitioned into training data (the majority) and validation data (the remainder). A neural network is applied to the training data (both attribute and reliability); the attributes become the system inputs and reliability data (cycles to failure), the system output. After the neural network is trained with sufficient data. the validation data are used to verify the neural networks provided accurate reliability estimates. Now, the reliability of a new proposed MEMS device can be estimated by using the appropriate trained neural networks developed in this work.

Perera. J. Sebastian

Software Reliability 2002

In FY01 we learned that hardware reliability models need substantial changes to account for differences in software, thus making software reliability measurements more effective, accurate, and easier to apply. These reliability models are generally based on familiar distributions or parametric methods. An obvious question is 'What new statistical and probability models can be developed using non-parametric and distribution-free methods instead of the traditional parametric method?" Two approaches to software reliability engineering appear somewhat promising. The first study, begin in FY01, is based in hardware reliability, a very well established science that has many aspects that can be applied to software. This research effort has investigated mathematical aspects of hardware reliability and has identified those applicable to software. Currently the research effort is applying and testing these approaches to software reliability measurement, These parametric models require much project data that may be difficult to apply and interpret. Projects at GSFC are often complex in both technology and schedules. Assessing and estimating reliability of the final system is extremely difficult when various subsystems are tested and completed long before others. Parametric and distribution free techniques may offer a new and accurate way of modeling failure time and other project data to provide earlier and more accurate estimates of system reliability.

Wallace, Dolores R.

Lunar RFC Reliability Testing for Assured Mission Success

NASA's Constellation program has selected the closed cycle hydrogen oxygen Polymer Electrolyte Membrane (PEM) regenerative Fuel Cell (RFC) as its baseline solar energy storage system for the lunar outpost and manned rover vehicles. Since the outpost and manned rovers are "human-rated", these energy storage systems will have to be of proven reliability exceeding 99 percent over the length of the mission. Because of the low (TRL=5) development state of the closed cycle hydrogen oxygen PEM RFC at present, and because there is no equivalent technology base in the commercial sector from which to draw or infer reliability information from, NASA will have to spend significant resources developing this technology from TRL 5 to TRL 9, and will have to embark upon an ambitious reliability development program to make this technology ready for a manned mission. Because NASA would be the first user of this new technology, NASA will likely have to bear all the costs associated with its development. When well-known reliability estimation techniques are applied to the hydrogen oxygen RFC to determine the amount of testing that will be required to assure RFC unit reliability over life of the mission, the analysis indicates the reliability testing phase by itself will take at least 2 yr, and could take up to 6 yr depending on the number of QA units that are built and tested and the individual unit reliability that is desired. The cost and schedule impacts of reliability development need to be considered in NASA's Exploration Technology Development Program (ETDP) plans, since life cycle testing to build meaningful reliability data is the only way to assure "return to the moon, this time to stay, then on to Mars" mission success.

Bents, David J.

Lunar Regenerative Fuel Cell (RFC) Reliability Testing for Assured Mission Success

NASA's Constellation program has selected the closed cycle hydrogen oxygen Polymer Electrolyte Membrane (PEM) Regenerative Fuel Cell (RFC) as its baseline solar energy storage system for the lunar outpost and manned rover vehicles. Since the outpost and manned rovers are "human-rated," these energy storage systems will have to be of proven reliability exceeding 99 percent over the length of the mission. Because of the low (TRL=5) development state of the closed cycle hydrogen oxygen PEM RFC at present, and because there is no equivalent technology base in the commercial sector from which to draw or infer reliability information from, NASA will have to spend significant resources developing this technology from TRL 5 to TRL 9, and will have to embark upon an ambitious reliability development program to make this technology ready for a manned mission. Because NASA would be the first user of this new technology, NASA will likely have to bear all the costs associated with its development.When well-known reliability estimation techniques are applied to the hydrogen oxygen RFC to determine the amount of testing that will be required to assure RFC unit reliability over life of the mission, the analysis indicates the reliability testing phase by itself will take at least 2 yr, and could take up to 6 yr depending on the number of QA units that are built and tested and the individual unit reliability that is desired. The cost and schedule impacts of reliability development need to be considered in NASA's Exploration Technology Development Program (ETDP) plans, since life cycle testing to build meaningful reliability data is the only way to assure "return to the moon, this time to stay, then on to Mars" mission success.

Bents, David J.

Developing Ultra Reliable Life Support for the Moon and Mars

Recycling life support systems can achieve ultra reliability by using spares to replace failed components. The added mass for spares is approximately equal to the original system mass, provided the original system reliability is not very low. Acceptable reliability can be achieved for the space shuttle and space station by preventive maintenance and by replacing failed units, However, this maintenance and repair depends on a logistics supply chain that provides the needed spares. The Mars mission must take all the needed spares at launch. The Mars mission also must achieve ultra reliability, a very low failure rate per hour, since it requires years rather than weeks and cannot be cut short if a failure occurs. Also, the Mars mission has a much higher mass launch cost per kilogram than shuttle or station. Achieving ultra reliable space life support with acceptable mass will require a well-planned and extensive development effort. Analysis must define the reliability requirement and allocate it to subsystems and components. Technologies, components, and materials must be designed and selected for high reliability. Extensive testing is needed to ascertain very low failure rates. Systems design should segregate the failure causes in the smallest, most easily replaceable parts. The systems must be designed, produced, integrated, and tested without impairing system reliability. Maintenance and failed unit replacement should not introduce any additional probability of failure. The overall system must be tested sufficiently to identify any design errors. A program to develop ultra reliable space life support systems with acceptable mass must start soon if it is to produce timely results for the moon and Mars.

Jones, Harry W.

ECLSS Reliability for Long Duration Missions Beyond Lower Earth Orbit

Reliability has been highlighted by NASA as critical to future human space exploration particularly in the area of environmental controls and life support systems. The Advanced Exploration Systems (AES) projects have been encouraged to pursue higher reliability components and systems as part of technology development plans. However there is no consensus on what is meant by improving on reliability; nor on how to assess reliability within the AES projects. This became apparent when trying to assess reliability as one of several figures of merit for a regenerable water architecture trade study. In the spring of 2013, the AES Water Recovery Project (WRP) hosted a series of events at the NASA Johnson Space Center (JSC) with the intended goal of establishing a common language and understanding of our reliability goals, and equipping the projects with acceptable means of assessing our respective systems. This campaign included an educational series in which experts from across the agency and academia provided information on terminology, tools and techniques associated with evalauating and designing for system reliability. The campaign culminated in a workshop at JSC with members of the ECLSS and AES communities with the goal of developing a consensus on what reliability means to AES and identifying methods for assessing our low to mid-technology readiness level (TRL) technologies for reliability. This paper details the results of the workshop.

Sargusingh, Miriam J.

Environmental Control and Life Support System Reliability for Long-Duration Missions Beyond Lower Earth Orbit

NASA has highlighted reliability as critical to future human space exploration, particularly in the area of environmental controls and life support systems. The Advanced Exploration Systems (AES) projects have been encouraged to pursue higher reliability components and systems as part of technology development plans. However, no consensus has been reached on what is meant by improving on reliability, or on how to assess reliability within the AES projects. This became apparent when trying to assess reliability as one of several figures of merit for a regenerable water architecture trade study. In the spring of 2013, the AES Water Recovery Project hosted a series of events at Johnson Space Center with the intended goal of establishing a common language and understanding of NASA's reliability goals, and equipping the projects with acceptable means of assessing the respective systems. This campaign included an educational series in which experts from across the agency and academia provided information on terminology, tools, and techniques associated with evaluating and designing for system reliability. The campaign culminated in a workshop that included members of the Environmental Control and Life Support System and AES communities. The goal of this workshop was to develop a consensus on what reliability means to AES and identify methods for assessing low- to mid-technology readiness level technologies for reliability. This paper details the results of that workshop.

Sargusingh, Miriam J.

ECLSS Reliability for Long Duration Missions Beyond Lower Earth Orbit

Reliability has been highlighted by NASA as critical to future human space exploration particularly in the area of environmental controls and life support systems. The Advanced Exploration Systems (AES) projects have been encouraged to pursue higher reliability components and systems as part of technology development plans. However, there is no consensus on what is meant by improving on reliability; nor on how to assess reliability within the AES projects. This became apparent when trying to assess reliability as one of several figures of merit for a regenerable water architecture trade study. In the Spring of 2013, the AES Water Recovery Project (WRP) hosted a series of events at the NASA Johnson Space Center (JSC) with the intended goal of establishing a common language and understanding of our reliability goals and equipping the projects with acceptable means of assessing our respective systems. This campaign included an educational series in which experts from across the agency and academia provided information on terminology, tools and techniques associated with evaluating and designing for system reliability. The campaign culminated in a workshop at JSC with members of the ECLSS and AES communities with the goal of developing a consensus on what reliability means to AES and identifying methods for assessing our low to mid-technology readiness level (TRL) technologies for reliability. This paper details the results of the workshop.

Sargusingh, Miriam J.

Estimating Software Reliability for Space Launch Vehicles in Probabilistic Risk Assessment (PRA)

It is acutely recognized in the Probabilistic Risk assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless if the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the NASA PRA and present a future dynamic option for software and PRA. Space Launch Vehicle Software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data, respectively. Software uncertainty will also be discussed. We determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.

Novack, Steven

Estimating Software Reliability for Space Launch Vehicles in Probabilistic Risk Assessment (PRA)

It is acutely recognized in the Probabilistic Risk Assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless of whether the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the National Aeronautics and Space Administration (NASA) PRA and present a future dynamic option for software and PRA. Space Launch Vehicle software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data. Software uncertainty will also be discussed. It is determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.

Steven D. Novack

A Demonstration that Correcting for Completeness and Reliability Is Critical for Robust Occurrence Rates

A measurement of planetary occurrence rates based on a planet catalog should be robust against details of how initial detections were classified as planets or false positives. This is accomplished by supplying the catalog’s rate of missed planets (completeness) and rate of non-planets incorrectly called planets (reliability). The final Kepler data release (DR25) includes products that can be used with the DR25 planet candidate catalog to correct for completeness and reliability in occurrence rate estimates. This is made possible by the Kepler Robovetter, which algorithmically and uniformly selects planets based on a variety of metrics and thresholds. Completeness, reliability, and occurrence rates potentially depend on these Robovetter thresholds. We study the impact of varying these vetting thresholds using the techniques of Bryson et al. 2019 (arXiv:1906.03575). We explore sets of thresholds that result in more or fewer planets (trading off completeness for reliability), as well as thresholds tuned to pass DR25 false positives identified as possible planets by the Kepler False Positive Working Group. We find that when correcting only for completeness, and not reliability, the resulting occurrence rates have a strong dependence on these threshold sets. For example, the value of SAG13 eta-Earth varies by over a factor of 4 when not corrected for reliability. However, when correcting for both completeness and reliability, occurrence rates using our threshold sets are statistically indistinguishable, with differences being well inside 1-sigma error bars. We present occurrence rates integrated over several period-radius ranges. For example, SAG13 eta-Earth is consistent with 0.127 (+0.094)(-0.054) (from Bryson et al. 2019) for all the Robovetter threshold sets. This result emphasizes the importance of correcting occurrence rates for both completeness and reliability. This suggests that inconsistent completeness and reliability correction may be a significant contributor to the large variation of occurrence rates in recent literature. We plan to make the Robovetter results for our threshold sets available, and encourage the community to use them to examine whether other occurrence rate methods yield similarly robust results.

Bryson, S.