NASA'S reliability requirements.
NASA reliability program provisions for space system contractors, detailing program management, reliability engineering, testing and evaluation
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
NASA reliability program provisions for space system contractors, detailing program management, reliability engineering, testing and evaluation
NASA reliability program provisions for aeronautical and space system contractors, reviewing evolution process
This proposed presentation will review past efforts and assess what emerging technologies could be used to address the reliability requirements for a Mars Sample Return Mission.
The relaunching of unsuccessful experiments or satellites will become a real option with the advent of the space shuttle. An examination was made of the cost effectiveness of relaxing reliability requirements for experiment hardware by allowing more than one flight of an experiment in the event of its failure. Any desired overall reliability or probability of mission success can be acquired by launching an experiment with less reliability two or more times if necessary. Although this procedure leads to uncertainty in total cost projections, because the number of flights is not known in advance, a considerable cost reduction can sometimes be achieved. In cases where reflight costs are low relative to the experiment's cost, three flights with overall reliability 0.9 can be made for less than half the cost of one flight with a reliability of 0.9. An example typical of shuttle payload cost projections is cited where three low reliability flights would cost less than $50 million and a single high reliability flight would cost over $100 million. The ratio of reflight cost to experiment cost is varied and its effect on the range in total cost is observed. An optimum design reliability selection criterion to minimize expected cost is proposed, and a simple graphical method of determining this reliability is demonstrated.
Every institution has their “recipe” for success. Cell phones and personal computers seem like they are designed to be obsolete in 2-3 years. Automobiles seem like they are designed to have a power train failure within 1000 miles of the extended warranty expiration; at least that’s been my experience. The Jet Propulsion Laboratory (JPL) is not any different. JPL has a tried and proven recipe for success because in space things can’t fail. Or else. However, in today’s competitive environment and funding limitations, that recipe for success is being challenged and the resultant increased risk accepted. This paper will describe JPL’s Risk Informed Decision Making (RIDM) approach to tailoring reliability requirements based on mission classification and other project characteristics.
Nasa requirements for space systems reliability engineering, program management and test and evaluation
Objectives, effectiveness and results of reliability program at Kennedy Space Center /KSC/ referring to Saturn launches
The use and implementation of Ada in distributed environments in which reliability is the primary concern is investigated. Emphasis is placed on the possibility that a distributed system may be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processors they are executing on, and that failures may occur in the software or underlying hardware. A new linguistic construct, the colloquy, is introduced which solves the problems identified in an earlier proposal, the conversation. It was shown that the colloquy is at least as powerful as recovery blocks, but it is also as powerful as all the other language facilities proposed for other situations requiring backward error recovery: recovery blocks, deadlines, generalized exception handlers, traditional conversations, s-conversations, and exchanges. The major features that distinguish the colloquy are described. Sample programs that were written, but not executed, using the colloquy show that extensive backward error recovery can be included in these programs simply and elegantly. These ideas are being implemented in an experimental Ada test bed.
The use and implementation of Ada in distributed environments in which reliability is the primary concern is investigated. Emphasis is placed on the possibility that a distributed system may be programmed entirely in ADA so that the individual tasks of the system are unconcerned with which processors they are executing on, and that failures may occur in the software or underlying hardware. The primary activities are: (1) Continued development and testing of our fault-tolerant Ada testbed; (2) consideration of desirable language changes to allow Ada to provide useful semantics for failure; (3) analysis of the inadequacies of existing software fault tolerance strategies.
The use and implementation of Ada in distributed environments in which reliability is the primary concern were investigted. A distributed system, programmed entirely in Ada, was studied to assess the use of individual tasks without concern for the processor used. Continued development and testing of the fault tolerant Ada testbed; development of suggested changes to Ada to cope with the failures of interest; design of approaches to fault tolerant software in real time systems, and the integration of these ideas into Ada; and the preparation of various papers and presentations were discussed.
The use and implementation of Ada in distributed environments in which reliability is the primary concern were investigated. In particular, the concept that a distributed system may be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processors they are executing on, and that failures may occur in the software or underlying hardware was examined. Progress is discussed for the following areas: continued development and testing of the fault-tolerant Ada testbed; development of suggested changes to Ada so that it might more easily cope with the failure of interest; and design of new approaches to fault-tolerant software in real-time systems, and integration of these ideas into Ada.
The use and implementation of Ada were investigated in distributed environments in which reliability is the primary concern. In particular, the focus was on the possibility that a distributed system may be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processors are being executed, and that failures may occur in the software and underlying hardware. A secondary interest is in the performance of Ada systems and how that performance can be gauged reliably. Primary activities included: analysis of the original approach to recovery in distributed Ada programs using the Advanced Transport Operating System (ATOPS) example; review and assessment of the original approach which was found to be capable of improvement; development of a refined approach to recovery that was applied to the ATOPS example; and design and development of a performance assessment scheme for Ada programs based on a flexible user-driven benchmarking system.
Rockwell International is conducting an ongoing program to develop avionics architectures that provide high intrinsic value while meeting all mission objectives. Studies are being conducted to determine alternative configurations that have low life-cycle cost and minimum development risk, and that minimize launch delays while providing the reliability level to assure a successful mission. This effort is based on four decades of providing ballistic missile avionics to the United States Air Force and has focused on the requirements of the NASA Cargo Transfer Vehicle (CTV) program in 1991. During the development of architectural concepts it became apparent that rendezvous strategy issues have an impact on the architecture of the avionics system. This is in addition to the expected impact on propulsion and electrical power duration, flight profiles, and trajectory during approach.
The use and implementation of Ada in distributed environments in which the hardware components are assumed to be unreliable is investigated. The possibility that a distributed system can be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processor they are executing on, and that failures can occur in the underlying hardware is considered. The reduced cost of computer hardware and the advantages of distributed processing (for example, increased reliability through redundancy and greater flexibility) indicate that many aerospace computer systems can be distributed. The use of Ada and distributed systems is a good combination for aerospace embedded systems.
It is sometimes optimistically hoped that a space life support system can be kept working throughout a long duration mission by repairing failed components, as long as sufficient spares are flown. It is usually assumed that the components have constant known failure rates. Then the needed numbers of spares can be computed to have any particular probability that all failed components can be replaced by available spares. This approach can provide high reliability if its favorable assumptions, including constant known failure rates, are satisfied. Other favorable assumptions are that the failures are statistically independent, repair will be successful without causing further failures, and all failures are due to internal component failures. These assumptions are not usually justified. The failure rates may be estimates that are inadequately verified because of insufficient testing. Failure rates may change due to materials substitutions, manufacturing changes, redesigns to fix failures, and new failures caused by redesigns. Failures that are not statistically independent may result from one common cause, such as a design or manufacturing error or a cascade of cause and effect, possibly caused by an external event such as a power outage. Repair may be unsuccessful or cause damage. Many failures occur at component interfaces or at the overall systems level, not within isolated components. Other failures causes are completely external to the system, due to assembly, maintenance, and operational errors or to unexpected environmental challenges. Replacement with sufficient spares can compensate for expected internal component failures but may not be able to cope with unpredictable design and manufacturing flaws, human errors, and environmental impacts. Reliability estimates based on providing sufficient spares to compensate for expected failures may be far too high. They are essentially upper bounds on reliability that might be approached if many frequent but often unconsidered failure causes can be eliminated.
An emerging new mission for aeronautics is Urban Air Mobility (UAM), a concept for air transportation around metropolitan areas with passenger-carrying operations. UAM vehicles must be capable of vertical take-off and landing, and this requirement presents unique technical challenges for electric and hybrid-based vertical take-off and landing (eVTOL). A critical challenge for UAM market growth is to gain public acceptance for being as safe as - or safer than - commercial air travel and automotive transportation. There is a lack of data for propulsion systems, components, and the associated thermal management systems for UAM eVTOL propulsion systems. The new mission, new propulsion system concepts, safety criticality of propulsion component performance during vertical take-off and lift operations, and lack of data presents many research challenges and opportunities. NASA has developed and published UAM vehicle concept studies. For a subset of the said concept vehicles, NASA has contracted for a study to identify failure modes and hazards associated with the propulsion systems of the concept vehicles and to perform functional hazard analyses (FHA) and failure modes and effects criticality analyses (FMECA) for each. From the completed study results, it was recommended for NASA to support research toward developing electric/hybrid-electric propulsion components with improved reliability and to explore powertrain architectures that can take advantage of higher reliability components to achieve inherent air-vehicle safety. NASA has started a research effort for UAM propulsion with a focus toward improving safety and reliability. Recent results and research strategy will be discussed toward the goals by means of: 1) improving individual component reliability through advanced materials and design methods, 2) improving the thermal management system, and 3) designing propulsion system architectures to provide inherent UAM vehicle safety.
The issues involved in the use of the programming language Ada on distributed systems are discussed. The effects of Ada programs on hardware failures such as loss of a processor are emphasized. It is shown that many Ada language elements are not well suited to this environment. Processor failure can easily lead to difficulties on those processors which remain. As an example, the calling task in a rendezvous may be suspended forever if the processor executing the serving task fails. A mechanism for detecting failure is proposed and changes to the Ada run time support system are suggested which avoid most of the difficulties. Ada program structures are defined which allow programs to reconfigure and continue to provide service following processor failure.
The use and implementation of Ada (a trade mark of the US Dept. of Defense) in distributed environments in which the hardware are assumed to be unreliable were investigated. The possibility that a distributed system is programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processors they are executing on and failures occurring in the underlying hardware were examined.