Search NASASearch

SEARCH · Search NASA

Results for “System reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Ultra reliability at NASA

Ultra reliable systems are critical to NASA particularly as consideration is being given to extended lunar missions and manned missions to Mars. NASA has formulated a program designed to improve the reliability of NASA systems. The long term goal for the NASA ultra reliability is to ultimately improve NASA systems by an order of magnitude. The approach outlined in this presentation involves the steps used in developing a strategic plan to achieve the long term objective of ultra reliability. Consideration is given to: complex systems, hardware (including aircraft, aerospace craft and launch vehicles), software, human interactions, long life missions, infrastructure development, and cross cutting technologies. Several NASA-wide workshops have been held, identifying issues for reliability improvement and providing mitigation strategies for these issues. In addition to representation from all of the NASA centers, experts from government (NASA and non-NASA), universities and industry participated. Highlights of a strategic plan, which is being developed using the results from these workshops, will be presented.

risk

Integrating Reliability Analysis with a Performance Tool

A large number of commercial simulation tools support performance oriented studies of complex computer and communication systems. Reliability of these systems, when desired, must be obtained by remodeling the system in a different tool. This has obvious drawbacks: (1) substantial extra effort is required to create the reliability model; (2) through modeling error the reliability model may not reflect precisely the same system as the performance model; (3) as the performance model evolves one must continuously reevaluate the validity of assumptions made in that model. In this paper we describe an approach, and a tool that implements this approach, for integrating a reliability analysis engine into a production quality simulation based performance modeling tool, and for modeling within such an integrated tool. The integrated tool allows one to use the same modeling formalisms to conduct both performance and reliability studies. We describe how the reliability analysis engine is integrated into the performance tool, describe the extensions made to the performance tool to support the reliability analysis, and consider the tool's performance.

Nicol, David M.

Reliability and Failure in NASA Missions: Blunders, Normal Accidents, High Reliability, Bad Luck

NASA emphasizes crew safety and system reliability but several unfortunate failures have occurred. The Apollo 1 fire was mistakenly unanticipated. After that tragedy, the Apollo program gave much more attention to safety. The Challenger accident revealed that NASA had neglected safety and that management underestimated the high risk of shuttle. Probabilistic Risk Assessment was adopted to provide more accurate failure probabilities for shuttle and other missions. NASA's "faster, better, cheaper" initiative and government procurement reform led to deliberately dismantling traditional reliability engineering. The Columbia tragedy and Mars mission failures followed. Failures can be attributed to blunders, normal accidents, or bad luck. Achieving high reliability is difficult but possible.

accidents

A Comparison of a Brain-Based Adaptive System and a Manual Adaptable System for Invoking Automation

Two experiments are presented that examine alternative methods for invoking automation. In each experiment, participants were asked to perform simultaneously a monitoring task and a resource management task as well as a tracking task that changed between automatic and manual modes. The monitoring task required participants to detect failures of an automated system to correct aberrant conditions under either high or low system reliability. Performance on each task was assessed as well as situation awareness and subjective workload. In the first experiment, half of the participants worked with a brain-based system that used their EEG signals to switch the tracking task between automatic and manual modes. The remaining participants were yoked to participants from the adaptive condition and received the same schedule of mode switches, but their EEG had no effect on the automation. Within each group, half of the participants were assigned to either the low or high reliability monitoring task. In addition, within each combination of automation invocation and system reliability, participants were separated into high and low complacency potential groups. The results revealed no significant effects of automation invocation on the performance measures; however, the high complacency individuals demonstrated better situation awareness when working with the adaptive automation system. The second experiment was the same as the first with one important exception. Automation was invoked manually. Thus, half of the participants pressed a button to invoke automation for 10 s. The remaining participants were yoked to participants from the adaptable condition and received the same schedule of mode switches, but they had no control over the automation. The results showed that participants who could invoke automation performed more poorly on the resource management task and reported higher levels of subjective workload. Further, those who invoked automation more frequently performed more poorly on the tracking task and reported higher levels of subjective workload. and the adaptable condition in the second experiment revealed only one significant difference: the subjective workload was higher in the adaptable condition. Overall, the results show that a brain-based, adaptive automation system may facilitate situation awareness for those individuals who are more complacent toward automation. By contrast, requiring operators to invoke automation manually may have some detrimental impact on performance but does appear to increases subjective workload relative to an adaptive system.

Bailey, Nathan R.

Reliability evaluation methodology for NASA applications

Liquid rocket engine technology has been characterized by the development of complex systems containing large number of subsystems, components, and parts. The trend to even larger and more complex system is continuing. The liquid rocket engineers have been focusing mainly on performance driven designs to increase payload delivery of a launch vehicle for a given mission. In otherwords, although the failure of a single inexpensive part or component may cause the failure of the system, reliability in general has not been considered as one of the system parameters like cost or performance. Up till now, quantification of reliability has not been a consideration during system design and development in the liquid rocket industry. Engineers and managers have long been aware of the fact that the reliability of the system increases during development, but no serious attempts have been made to quantify reliability. As a result, a method to quantify reliability during design and development is needed. This includes application of probabilistic models which utilize both engineering analysis and test data. Classical methods require the use of operating data for reliability demonstration. In contrast, the method described in this paper is based on similarity, analysis, and testing combined with Bayesian statistical analysis.

Taneja, Vidya S.

PEM-INST-001: Instructions for Plastic Encapsulated Microcircuit (PEM) Selection, Screening, and Qualification

Potential users of plastic encapsulated microcircuits (PEMs) need to be reminded that unlike the military system of producing robust high-reliability microcircuits that are designed to perform acceptably in a variety of harsh environments, PEMs are primarily designed for use in benign environments where equipment is easily accessed for repair or replacement. The methods of analysis applied to military products to demonstrate high reliability cannot always be applied to PEMs. This makes it difficult for users to characterize PEMs for two reasons: 1. Due to the major differences in design and construction, the standard test practices used to ensure that military devices are robust and have high reliability often cannot be applied to PEMs that have a smaller operating temperature range and are typically more frail and susceptible to moisture absorption. In contrast, high-reliability military microcircuits usually utilize large, robust, high-temperature packages that are hermetically sealed. 2. Unlike the military high-reliability system, users of PEMs have little visibility into commercial manufacturers proprietary design, materials, die traceability, and production processes and procedures. There is no central authority that monitors PEM commercial product for quality, and there are no controls in place that can be imposed across all commercial manufacturers to provide confidence to high-reliability users that a common acceptable level of quality exists for all PEMs manufacturers. Consequently, there is no guaranteed control over the type of reliability that is built into commercial product, and there is no guarantee that different lots from the same manufacturer are equally acceptable. And regarding application, there is no guarantee that commercial products intended for use in benign environments will provide acceptable performance and reliability in harsh space environments. The qualification and screening processes contained in this document are intended to detect poor-quality lots and screen out early random failures from use in space flight hardware. However, since it cannot be guaranteed that quality was designed and built into PEMs that are appropriate for space applications, users cannot screen in quality that may not exist. It must be understood that due to the variety of materials, processes, and technologies used to design and produce PEMs, this test process may not accelerate and detect all failure mechanisms. While the tests herein will increase user confidence that PEMs with otherwise unknown reliability can be used in space environments, such testing may not guarantee the same level of reliability offered by military microcircuits. PEMs should only be used where due to performance needs there are no alternatives in the military high-reliability market, and projects are willing to accept higher risk.

Teverovsky, Alexander

ETARA PC version 3.3 user's guide: Reliability, availability, maintainability simulation model

A user's manual describing an interactive, menu-driven, personal computer based Monte Carlo reliability, availability, and maintainability simulation program called event time availability reliability (ETARA) is discussed. Given a reliability block diagram representation of a system, ETARA simulates the behavior of the system over a specified period of time using Monte Carlo methods to generate block failure and repair intervals as a function of exponential and/or Weibull distributions. Availability parameters such as equivalent availability, state availability (percentage of time as a particular output state capability), continuous state duration and number of state occurrences can be calculated. Initial spares allotment and spares replenishment on a resupply cycle can be simulated. The number of block failures are tabulated both individually and by block type, as well as total downtime, repair time, and time waiting for spares. Also, maintenance man-hours per year and system reliability, with or without repair, at or above a particular output capability can be calculated over a cumulative period of time or at specific points in time.

Hoffman, David J.

The Plant Water Management Experiments: Soil

A simple means of watering plants in the low-g environment aboard orbiting spacecraft is not obvious. Since the beginning of spaceflight, numerous approaches have been pursued to water plants that seek to maximize plant viability and system reliability, while minimizing crew time and system complexity. We are not there yet. The Plant Water Management (PWM) Soil experiments seek to apply recent advances in low-g capillary fluidics phenomena to the challenges faced by plant growth operations aboard spacecraft. The primary challenge is to establish earth-like flows minimizing low-g specific adaptations required of the plants. This is difficult due to the ever-present fluid physics challenges of poorly-wetting multiphase inertial-visco-capillary flows in geometrically complex conduits and containers. In this paper, we present recent flight results for the PWM Soil experiments where arcillite ‘soil reservoirs’ are arranged in a non-wetting host soil that serves as an O2-breathing wetting barrier. In this way, a largely terrestrial water-soil environment is mimicked where, as liquid is evapo-transpired through the growing plant foliage, the effective water table passively ‘falls’ reducing viscous lengths and increasing water uptake for the plant. We present data from 6 days of 24-7 experiments on the ISS testing 3 different plant root models. We also present and correlate a capillary flow model which captures the primary features of the flow. Our summary is valued for the assessment of current and future low-g plant watering systems employing soil media.

microgravity

ARIES Annual Report FY25

Advanced Research on Integrated Energy Systems (ARIES) at the National Laboratory of the Rockies (NLR) is the U.S. Department of Energy's (DOE's) test bed for energy system demonstration and de-risking. ARIES comprises the largest collection of physical and digital assets in the DOE laboratory complex, supporting flexible configuration across a broad range of energy scenarios. In Fiscal Year 2025, ARIES provided a platform for system-level research to anticipate and address future energy needs in energy security, system reliability, and technology deployment.

24 POWER TRANSMISSION AND DISTRIBUTION

Orbiter Autoland reliability analysis

The Space Shuttle Orbiter is the only space reentry vehicle in which the crew is seated upright. This position presents some physiological effects requiring countermeasures to prevent a crewmember from becoming incapacitated. This also introduces a potential need for automated vehicle landing capability. Autoland is a primary procedure that was identified as a requirement for landing following and extended duration orbiter mission. This report documents the results of the reliability analysis performed on the hardware required for an automated landing. A reliability block diagram was used to evaluate system reliability. The analysis considers the manual and automated landing modes currently available on the Orbiter. (Autoland is presently a backup system only.) Results of this study indicate a +/- 36 percent probability of successfully extending a nominal mission to 30 days. Enough variations were evaluated to verify that the reliability could be altered with missions planning and procedures. If the crew is modeled as being fully capable after 30 days, the probability of a successful manual landing is comparable to that of Autoland because much of the hardware is used for both manual and automated landing modes. The analysis indicates that the reliability for the manual mode is limited by the hardware and depends greatly on crew capability. Crew capability for a successful landing after 30 days has not been determined yet.

Welch, D. Phillip

A latent fault Markov model for a highly reliable triplex computer system

A Markov model of a highly reliable triplex system was constructed to evaluate the probability of system failure as a function of the propagation of latent faults. It is found that if the propagation rate of latent faults is extremely high, they do not significantly affect the probability of system failure, while if the propagation rate is extremely low, the survivability of the system is improved. The propagation rate that is most harmful to the survivability of the system is determined as a function of the duration of the flight. A decrease in the probability of system failure due to latency is noted if the probability of any two faults giving the same output is extremely low.

Swern, Frederic L.

Going beyond reliability to robustness and resilience in space systems

The words reliability, robustness, and resilience, are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environments is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes are usually unanticipated. They often invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging events. These three are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from disruptions. Increasing the resilience of space systems, the ability to perform after unanticipated events, would greatly increase space crew safety. Improving reliability and robustness can be done by dealing with known sources of problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. Two contrasting approaches are reducing system complexity and adding supervisory control. The need for resilience has been claimed for decades but little has been accomplished. Systems designers assume that they understand requirements, technologies, designs, architectures, integration, testing, operations, and environments. The potential problems of changes, failures, accidents, unknown environments, and unknown unknowns are ignored. Systems designers are typically overconfident and ignore the need for robustness and resilience.

Harry W Jones

Choosing reliability level for Shuttle-carried payloads

The paramount importance of high payload reliability is questioned in the light of payload recoverability in Spacelab and Space Shuttle experiments, and cost effectiveness through less stringent reliability requirements backed up by recovery and reflight of the experiment if necessary is considered. Ground rules and parameters for assessing reliability and recoverability in relation to mission cost effectiveness are advanced. Mission success probability (for one or several flights), experiment hardware cost, design reliability, cost of orbiting the experiment, support system reliability, and other relevant parameters are identified and discussed, and graphs of experimental total cost envelope are plotted. It is proposed that reliability and quality assurance standards be weighed against failure costs and penalties in the case of recoverable space systems.

Campbell, J. W.

National launch system core vehicle

A new launch system is being defined to provide significant improvements in the national launch capability. The most challenging requirements are to reduce operational cost and increase system reliability, dependability, responsiveness while maintaining acceptable mission performance. System architecture studies have established that a family of vehicles is required to launch the required range of payloads. The vehicle design approach implements a design process that identifies the system design criteria and design concepts that enable the vehicle to meet the objectives of cost, reliability, safety and performance. The vehicle design features for the 50 to 100 K payload capability will utilize cost effective elements derived from existing vehicles, emerging technology and lessons learned to enable the system to be developed at a reasonable investment and to deliver payloads in an operationally efficient manner with high reliability.

Worlund, Armis L.

Electrical, Electronic, and Electromechanical (EEE) parts management and control requirements for NASA space flight programs

This document establishes electrical, electronic, and electromechanical (EEE) parts management and control requirements for contractors providing and maintaining space flight and mission-essential or critical ground support equipment for NASA space flight programs. Although the text is worded 'the contractor shall,' the requirements are also to be used by NASA Headquarters and field installations for developing program/project parts management and control requirements for in-house and contracted efforts. This document places increased emphasis on parts programs to ensure that reliability and quality are considered through adequate consideration of the selection, control, and application of parts. It is the intent of this document to identify disciplines that can be implemented to obtain reliable parts which meet mission needs. The parts management and control requirements described in this document are to be selectively applied, based on equipment class and mission needs. Individual equipment needs should be evaluated to determine the extent to which each requirement should be implemented on a procurement. Utilization of this document does not preclude the usage of other documents. The entire process of developing and implementing requirements is referred to as 'tailoring' the program for a specific project. Some factors that should be considered in this tailoring process include program phase, equipment category and criticality, equipment complexity, and mission requirements. Parts management and control requirements advocated by this document directly support the concept of 'reliability by design' and are an integral part of system reliability and maintainability. Achieving the required availability and mission success objectives during operation depends on the attention given reliability and maintainability in the design phase. Consequently, it is intended that the requirements described in this document are consistent with those of NASA publications, 'Reliability Program Requirements for Aeronautical and Space System Contractors,' NHB 5300.4(1A-l); 'Maintainability Program Requirements for Space Systems,' NHB 5300.4(1E); and 'Quality Program Provisions for Aeronautical and Space System Contractors,' NHB 5300.4(1B).

Source record