Reliable software systems design: defect prevention, detection, and containment
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
The problem to be addressed in this paper is to explore how the use of Soft Computing Technologies (SCT) could be employed to improve overall vehicle system safety, reliability, and rocket engine performance by development of a qualitative and reliable engine control system (QRECS). Specifically, this will be addressed by enhancing rocket engine control using SCT, innovative data mining tools, and sound software engineering practices used in Marshall's Flight Software Group (FSG). The principle goals for addressing the issue of quality are to improve software management, software development time, software maintenance, processor execution, fault tolerance and mitigation, and nonlinear control in power level transitions. The intent is not to discuss any shortcomings of existing engine control methodologies, but to provide alternative design choices for control, implementation, performance, and sustaining engineering, all relative to addressing the issue of reliability. The approaches outlined in this paper will require knowledge in the fields of rocket engine propulsion (system level), software engineering for embedded flight software systems, and soft computing technologies (i.e., neural networks, fuzzy logic, data mining, and Bayesian belief networks); some of which are briefed in this paper. For this effort, the targeted demonstration rocket engine testbed is the MC-1 engine (formerly FASTRAC) which is simulated with hardware and software in the Marshall Avionics & Software Testbed (MAST) laboratory that currently resides at NASA's Marshall Space Flight Center, building 4476, and is managed by the Avionics Department. A brief plan of action for design, development, implementation, and testing a Phase One effort for QRECS is given, along with expected results. Phase One will focus on development of a Smart Start Engine Module and a Mainstage Engine Module for proper engine start and mainstage engine operations. The overall intent is to demonstrate that by employing soft computing technologies, the quality and reliability of the overall scheme to engine controller development is further improved and vehicle safety is further insured. The final product that this paper proposes is an approach to development of an alternative low cost engine controller that would be capable of performing in unique vision spacecraft vehicles requiring low cost advanced avionics architectures for autonomous operations from engine pre-start to engine shutdown.
This paper contains an overview of a theoretical framework for the design of reliable multivariable control systems, with special emphasis on actuator failures and necessary actuator redundancy levels. Using a linear model of the system, with Markovian failure probabilities and quadratic performance index, an optimal stochastic control problem is posed and solved. The solution requires the iteration of a set of highly coupled Riccati-like matrix difference equations; if these converge one has a reliable design; if they diverge, the design is unreliable, and the system design cannot be stabilized. In addition, it is shown that the existence of a stabilizing constant feedback gain and the reliability of its implementation is equivalent to the convergence properties of a set of coupled Riccati-like matrix difference equations. In summary, these results can be used for offline studies relating the open loop dynamics, required performance, actuator mean time to failure, and functional or identical actuator redundancy, with and without feedback gain reconfiguration strategies.
Apollo gyro reliability covering Guidance, Navigation and Control systems, stressing failure mode prediction
Reliability modeling must take into account two different types of phenomena, including the fault-occurrence behavior and the fault/error-handling behavior of a system. The effectiveness of the fault/error-handling behavior can be captured by instantaneous coverage probabilities. This paper has the objective to show that the assumption of instantaneous coverage leads to conservative predictions of system reliability for systems characterized by relatively long interevent times for fault occurrences and relatively short interevent times for fault/error-handling actions. The importance of this result is related to the fact that it can now be shown that model predictions based on instantaneous coverage are lower bounds on the true system reliability. Attention is given to a semi-Markov reliability model, instantaneous coverage approximations, the proof of conservative prediction, and the computation of coverage probabilities.
Systems for Space Defense Initiative (SDI) space applications typically require both high performance and very high reliability. These requirements present the systems engineer evaluating such systems with the extremely difficult problem of conducting performance and reliability trade-offs over large design spaces. A controlled development process supported by appropriate automated tools must be used to assure that the system will meet design objectives. This report describes an investigation of methods, tools, and techniques necessary to support performance and reliability modeling for SDI systems development. Models of the JPL Hypercubes, the Encore Multimax, and the C.S. Draper Lab Fault-Tolerant Parallel Processor (FTPP) parallel-computing architectures using candidate SDI weapons-to-target assignment algorithms as workloads were built and analyzed as a means of identifying the necessary system models, how the models interact, and what experiments and analyses should be performed. As a result of this effort, weaknesses in the existing methods and tools were revealed and capabilities that will be required for both individual tools and an integrated toolset were identified.
This Phase I project Summary will be formatted as a series of summary statements, with additional comments. The summary statements are intended to focus on the most significant findings and discoveries – things that the project team knows now, but we didn’t know at the beginning of the project: 1) Direct Contact Between Gases is an ECLS (Environmental Control and Life Support) System enabling capability; 2) CO2 capture using liquid sorbents in microgravity is feasible; 3) There are other ways to contact gases and liquids – but thin film capillary techniques are new, exciting, and have amazing potential; 4) ECLS system reliability is the key to exploration missions – the key to reliability is having system attributes that favor reliability; 5) The processes with favorable reliability attributes tend to be biological; 6) The single greatest impact on launch mass of an ECLS system is water. The best way to enable biological water processing is to develop a capillary based method of urine capture that doesn’t use pretreat chemicals; 7) There is good, promising, forward design and development work – but no fundamental “show stoppers” – to develop a thin film liquid sorbent CO2 capture system; 8) The most capable Ionic Liquids are not presently feasible, but other chemically active liquids can be used to make an effective thin film CO2 capture device; 9) The C9 reduced gravity flight showed feasibility – and taught us about flow instability issues; 10) Capillary fluid management has ECLS system wide implications.
A digital computer code was developed to simulate the time-dependent behavior of the 5-kwe reactor thermoelectric system. The code was used to determine lifetime sensitivity coefficients for a number of system design parameters, such as thermoelectric module efficiency and degradation rate, radiator absorptivity and emissivity, fuel element barrier defect constant, beginning-of-life reactivity, etc. A probability distribution (mean and standard deviation) was estimated for each of these design parameters. Then, error analysis was used to obtain a probability distribution for the system lifetime (mean = 7.7 years, standard deviation = 1.1 years). From this, the probability that the system will achieve the design goal of 5 years lifetime is 0.993. This value represents an estimate of the degradation reliability of the system.
Nasa reliability evaluation program for operational systems set forth in npc 250-1
This report analyzes the reliability of NASA's Ultra-reliable Fault Tolerant Control System (UFTCS) architecture as it is currently envisioned for helicopter control. The analysis is extended to air transport and spacecraft control using the same computational and voter modules applied within the UFTCS architecture. The system reliability is calculated for several points in the helicopter, air transport, and space flight missions when there are initially 4, 5, and 6 operating channels. Sensitivity analyses are used to explore the effects of sensor failure rates and different system configurations at the 10 hour point of the helicopter mission. These analyses show that the primary limitation to system reliability is the number of flux windings on each flux summer (4 are assumed for the baseline case). Tables of system reliability at the 10 hour point are provided to allow designers to choose a configuration to meet specified reliability goals.
The study of a comparative analysis of distinct multiplex and fault-tolerant configurations for a PLC-based safety system from a reliability point of view is presented. It considers simplex, duplex and fault-tolerant triple redundancy configurations. The standby unit in case of a duplex configuration has a failure rate which is k times the failure rate of the standby unit, the value of k varying from 0 to 1. For distinct values of MTTR and MTTF of the main unit, MTBF and availability for these configurations are calculated. The effect of duplexing only the PLC module or only the sensors and the actuators module, on the MTBF of the configuration, is also presented. The results are summarized and merits and demerits of various configurations under distinct environments are discussed.
It is acutely recognized in the Probabilistic Risk assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless if the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the NASA PRA and present a future dynamic option for software and PRA. Space Launch Vehicle Software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data, respectively. Software uncertainty will also be discussed. We determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.
It is acutely recognized in the Probabilistic Risk Assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless of whether the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the National Aeronautics and Space Administration (NASA) PRA and present a future dynamic option for software and PRA. Space Launch Vehicle software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data. Software uncertainty will also be discussed. It is determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.
The development of a methodology for the production of highly reliable software is one of the greatest challenges facing the computer industry. Meeting this challenge will undoubtably involve the integration of many technologies. This paper describes the use of Artificial Intelligence technologies in the automated analysis of the formal algebraic specifications of abstract data types. These technologies include symbolic execution of specifications using techniques of automated deduction and machine learning through the use of examples. On-going research into the role of knowledge representation and problem solving in the process of developing software is also discussed.
The International Space Station (ISS) Water Recovery System (WRS) includes the Water Processor Assembly (WPA) and the Urine Processor Assembly (UPA). The WRS produces potable water from a combination of crew urine (first processed through the UPA), crew latent, and Sabatier product water. Though the WRS has performed well since operations began in November 2008, several modifications have been identified to improve the overall system performance. These modifications aim to reduce resupply and improve overall system reliability, which is beneficial for the ongoing ISS mission as well as for future NASA manned missions. The following paper details efforts to improve the WPA through the use of reverse osmosis membrane technology to reduce the resupply mass of the WPA Multi-filtration Bed and improved catalyst for the WPA Catalytic Reactor to reduce the operational temperature and pressure. For the UPA, this paper discusses progress on various concepts for improving the reliability of the system, including the implementation of a more reliable drive belt, improved methods for managing condensate in the stationary bowl of the Distillation Assembly, and evaluating upgrades to the UPA vacuum pump.
Fast acting, pneumatically and centrally controlled, fire extinguisher /firex/ system is effective in freezing climates. The easy-to-operate system provides a fail-dry function which is activated by an electrical power failure.
Multibody system equations can be generated in various forms. All of these may be interpreted as results of two basic approaches, the augmentation- and the elimination-method. The former method yields the descriptor form of the system motion, a set of differential-algebraic equations (DAE), and the latter the state space representation, a minimal set of ordinary differential equations (ODE). Both of these methods are surveyed. Particular emphasis is on the discussion of recursive computational schemes, generating the equations of motion with a number of operations, which is proportional to the number N of system bodies (O(N)-formulations). For simulation purposes one would like to create that set of system equations, which can be generated most efficiently and for which the most efficient and reliable solution techniques are available. Numerical solution techniques for ODE have been studied in great detail and they are well-developed. By contrast, DAE have not been investigated for such a long time. In view of new developments in the latter field the generation of all the equations required for an efficient and reliable solution of DAE describing multibody system motion is discussed. These methods, i.e., an O(N)-formulation and new techniques for solving DAE, are implemented in the SIMPACK code. Its capabilities are illustrated by simulation of multibody robot models.
Component part and system reliability program for Relay I satellite with redundancy incorporated at all levels of development