Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software Failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A self-reorganizing digital flight control system for aircraft

This paper presents a design method for digital self-reorganizing control systems which is optimally tolerant of failures in aircraft sensors. The functions of this system are accomplished with software instead of the popular and costly technique of hardware duplication. The theoretical development, based on M-ary hypothesis testing, results in a bank of M Kalman filters operating in parallel in the failure detection logic. A moving window of the innovations of each Kalman filter drives the detection logic to decide the failure state of the system. The detection logic also selects the optimal state estimate (for control logic) from the bank of Kalman filters. The design process is applied to the design of a self-reorganizing control system for a current configuration of the space shuttle orbiter at Mach 5 and 120,000 feet. The failure detection capabilities of the system are demonstrated using a real-time simulation of the system with noisy sensors.

Montgomery, R. C.↗

Software analysis handbook: Software complexity analysis and software reliability estimation and prediction

This handbook documents the three software analysis processes the Space Station Software Analysis team uses to assess space station software, including their backgrounds, theories, tools, and analysis procedures. Potential applications of these analysis results are also presented. The first section describes how software complexity analysis provides quantitative information on code, such as code structure and risk areas, throughout the software life cycle. Software complexity analysis allows an analyst to understand the software structure, identify critical software components, assess risk areas within a software system, identify testing deficiencies, and recommend program improvements. Performing this type of analysis during the early design phases of software development can positively affect the process, and may prevent later, much larger, difficulties. The second section describes how software reliability estimation and prediction analysis, or software reliability, provides a quantitative means to measure the probability of failure-free operation of a computer program, and describes the two tools used by JSC to determine failure rates and design tradeoffs between reliability, costs, performance, and schedule.

Computer systems design↗

The cleanroom case study in the Software Engineering Laboratory: Project description and early analysis

This case study analyzes the application of the cleanroom software development methodology to the development of production software at the NASA/Goddard Space Flight Center. The cleanroom methodology emphasizes human discipline in program verification to produce reliable software products that are right the first time. Preliminary analysis of the cleanroom case study shows that the method can be applied successfully in the FDD environment and may increase staff productivity and product quality. Compared to typical Software Engineering Laboratory (SEL) activities, there is evidence of lower failure rates, a more complete and consistent set of inline code documentation, a different distribution of phase effort activity, and a different growth profile in terms of lines of code developed. The major goals of the study were to: (1) assess the process used in the SEL cleanroom model with respect to team structure, team activities, and effort distribution; (2) analyze the products of the SEL cleanroom model and determine the impact on measures of interest, including reliability, productivity, overall life-cycle cost, and software quality; and (3) analyze the residual products in the application of the SEL cleanroom model, such as fault distribution, error characteristics, system growth, and computer usage.

Green, Scott↗

Flight test results of the Strapdown hexad Inertial Reference Unit (SIRU). Volume 1: Flight test summary

Flight test results of the strapdown inertial reference unit (SIRU) navigation system are presented. The fault-tolerant SIRU navigation system features a redundant inertial sensor unit and dual computers. System software provides for detection and isolation of inertial sensor failures and continued operation in the event of failures. Flight test results include assessments of the system's navigational performance and fault tolerance.

Hruby, R. J.↗

Flight test results of the strapdown hexad inertial reference unit (SIRU). Volume 2: Test report

Results of flight tests of the Strapdown Inertial Reference Unit (SIRU) navigation system are presented. The fault tolerant SIRU navigation system features a redundant inertial sensor unit and dual computers. System software provides for detection and isolation of inertial sensor failures and continued operation in the event of failures. Flight test results include assessments of the system's navigational performance and fault tolerance. Performance shortcomings are analyzed.

Hruby, R. J.↗

Flight test results of the Strapdown hexad Inertial Reference Unit (SIRU). Volume 3: Appendices A-G

Results of flight tests of the Strapdown Inertial Reference Unit (SIRU) navigation system are presented. The fault tolerant SIRU navigation system features a redundant inertial sensor unit and dual computers. System software provides for detection and isolation of inertial sensor failures and continued operation in the event of failures. Flight test results include assessments of the system's navigational performance and fault tolerance. Selected facets of the flight tests are also described in detail and include some of the following: (1) flight test plans and ground track plots; (2) navigation residual plots; (3) effects of approximations in navigation algorithms; (4) vibration spectrum of the CV-340 aircraft; and (5) modification of the statistical FDICR algorithm parameters for the flight environment.

Hruby, R. J.↗

Data management of Shuttle radiofrequency navigation aids

It is noted that the Shuttle navigation system employs redundant tactical air navigation (tacan) and microwave scanning beam landing system (MSBLS) equipment for use in navigation during descent from altitudes of about 150,000 feet through rollout. Attention is given here to the multiple tacan and MSBLS units (three each) that were placed onboard to provide the necessary protection in the event of possible failures. The goals, features, approach, and performance of onboard software required to manage multiple tacan MSBLS units and to provide the corresponding data for navigation processing are described.

Stokes, R. E.↗

The implementation and use of Ada on distributed systems with high reliability requirements

The use and implementation of Ada in distributed environments in which reliability is the primary concern were investigted. A distributed system, programmed entirely in Ada, was studied to assess the use of individual tasks without concern for the processor used. Continued development and testing of the fault tolerant Ada testbed; development of suggested changes to Ada to cope with the failures of interest; design of approaches to fault tolerant software in real time systems, and the integration of these ideas into Ada; and the preparation of various papers and presentations were discussed.

Knight, J. C.↗

Voyager programmability - Experiences in control system adaptation

The two Voyager spacecraft were designed with reprogrammable attitude and articulation control system flight control processors to permit inflight modification of the software. These software modifications were designed to: (1) correct for hardware failures that occurred during the interstellar journey, and (2) improve spacecraft dynamic performance for better control so that, by minimizing the use of propellant, the life of the mission could be extended. It is suggested that, in the future, spacecraft be launched with additional unused memory to deal with anomalies unforeseen in the decision-making process.

Patel, Keyur↗

Growing wheat to maturity in reduced gas pressures

The main objective of this project was to determine assimilation of CO2 and efficiency of water use in wheat grown to maturity in a low pressure total gas pressure environment. A functional test of the low pressure plant growth chamber system was accomplished in February and March of 1993 wherein this objective was partially achieved. Plants were grown to maturity in the chambers. Data were actively collected during the first 29 days. The plants were allowed to maintain themselves at the CO2 compensation point until day 45 of the study at which point active atmospheric regulation was resumed. This provided data at the vegetative and reproductive stages of the life cycle of the plants. However, this information may not be representative of the performance of the plants due to the loss of low pressure on a number of days during the study, which affected the plants by changing the pressure potential of the tissues. The performance of the system will be discussed on a component by component basis. The maintenance of the plants at the CO2 compensation point was driven by the failure of the computer program operating the system. The software problems that arose during the functional test have since been corrected. Results from the functional test also indicated that the plants were not receiving adequate light and nutrients. The growth chambers have been relocated and the growth room modified to compensate for these deficiencies.

Rykiel, Edward J., Jr.↗

Design for testability and diagnosis at the system-level

The growing complexity of full-scale systems has surpassed the capabilities of most simulation software to provide detailed models or gate-level failure analyses. The process of system-level diagnosis approaches the fault-isolation problem in a manner that differs significantly from the traditional and exhaustive failure mode search. System-level diagnosis is based on a functional representation of the system. For example, one can exercise one portion of a radar algorithm (the Fast Fourier Transform (FFT) function) by injecting several standard input patterns and comparing the results to standardized output results. An anomalous output would point to one of several items (including the FFT circuit) without specifying the gate or failure mode. For system-level repair, identifying an anomalous chip is sufficient. We describe here an information theoretic and dependency modeling approach that discards much of the detailed physical knowledge about the system and analyzes its information flow and functional interrelationships. The approach relies on group and flow associations and, as such, is hierarchical. Its hierarchical nature allows the approach to be applicable to any level of complexity and to any repair level. This approach has been incorporated in a product called STAMP (System Testability and Maintenance Program) which was developed and refined through more than 10 years of field-level applications to complex system diagnosis. The results have been outstanding, even spectacular in some cases. In this paper we describe system-level testability, system-level diagnoses, and the STAMP analysis approach, as well as a few STAMP applications.

Simpson, William R.↗

An improved approach for flight readiness certification: Probabilistic models for flaw propagation and turbine blade failure. Volume 1: Methodology and applications

An improved methodology for quantitatively evaluating failure risk of spaceflight systems to assess flight readiness and identify risk control measures is presented. This methodology, called Probabilistic Failure Assessment (PFA), combines operating experience from tests and flights with analytical modeling of failure phenomena to estimate failure risk. The PFA methodology is of particular value when information on which to base an assessment of failure risk, including test experience and knowledge of parameters used in analytical modeling, is expensive or difficult to acquire. The PFA methodology is a prescribed statistical structure in which analytical models that characterize failure phenomena are used conjointly with uncertainties about analysis parameters and/or modeling accuracy to estimate failure probability distributions for specific failure modes. These distributions can then be modified, by means of statistical procedures of the PFA methodology, to reflect any test or flight experience. State-of-the-art analytical models currently employed for designs failure prediction, or performance analysis are used in this methodology. The rationale for the statistical approach taken in the PFA methodology is discussed, the PFA methodology is described, and examples of its application to structural failure modes are presented. The engineering models and computer software used in fatigue crack growth and fatigue crack initiation applications are thoroughly documented.

Moore, N. R.↗

Automated Diagnosis Of Conditions In A Plant-Growth Chamber

Biomass Production Chamber Operations Assistant software and hardware constitute expert system that diagnoses mechanical failures in controlled-environment hydroponic plant-growth chamber and recommends corrective actions to be taken by technicians. Subjects of continuing research directed toward development of highly automated closed life-support systems aboard spacecraft to process animal (including human) and plant wastes into food and oxygen. Uses Microsoft Windows interface to give technicians intuitive, efficient access to critical data. In diagnostic mode, system prompts technician for information. When expert system has enough information, it generates recovery plan.

Clinger, Barry R.↗

The Strengths and Weaknesses of Logic Formalisms to Support Mishap Analysis

The increasing complexity of many safety critical systems poses new problems for mishap analysis. Techniques developed in the sixties and seventies cannot easily scale-up to analyze incidents involving tightly integrated software and hardware components. Similarly, the realization that many failures have systemic causes has widened the scope of many mishap investigations. Organizations, including NASA and the NTSB, have responded by starting research and training initiatives to ensure that their personnel are well equipped to meet these challenges. One strand of research has identified a range of mathematically based techniques that can be used to reason about the causes of complex, adverse events. The proponents of these techniques have argued that they can be used to formally prove that certain events created the necessary and sufficient causes for a mishap to occur. Mathematical proofs can reduce the bias that is often perceived to effect the interpretation of adverse events. Others have opposed the introduction of these techniques by identifying social and political aspects to incident investigation that cannot easily be reconciled with a logic-based approach. Traditional theorem proving mechanisms cannot accurately capture the wealth of inductive, deductive and statistical forms of inference that investigators routinely use in their analysis of adverse events. This paper summarizes some of the benefits that logics provide, describes their weaknesses, and proposes a number of directions for future research.

Johnson, C. W.↗

Orion GNC Mitigation Efforts for Van Allen Radiation

The Orion Crew Module (CM) is NASA's next generation manned space vehicle, scheduled to return humans to lunar orbit in the coming decade. The Orion avionics and GN&C architectures have progressed through a number of project phases and are nearing completion of a major milestone. The first unmanned test mission, dubbed "Exploration Flight Test One" (EFT-1) is scheduled to launch from NASA Kennedy Space Center late next year and provides the first integrated test of all the vehicle systems, avionics and software. The EFT-1 mission will be an unmanned test flight that includes a high speed re-entry from an elliptical orbit, which will be launched on an expendable launch vehicle (ELV). The ELV will place CM and the ELV upper stage into a low Earth orbit (LEO) for one revolution. After the first LEO, the ELV upper stage will re-ignite and place the combined upper stage/CM into an elliptical orbit whose perigee results in a high energy entry to test CM response in a relatively high velocity, high heating environment. While not producing entry velocities as high as those experienced in returning from a lunar orbit, the trajectory was chosen to provide higher stresses on the thermal protection and guided entry systems, as compared against a lower energy LEO entry. However the required entry geometry with constraints on inclination and landing site result in a trajectory that lingers for many hours in the Van Allen radiation belts. This exposes the vehicle and avionics to much higher levels of high energy proton radiation than a typical LEO or lunar trajectory would encounter. As a result, Van Allen radiation poses a significant risk to the Orion avionics system, and particularly the Flight Control Module (FCM) computers that house the GN&C flight software. The measures taken by the Orion GN&C, Flight Software and Avionics teams to mitigate the risks associated with the Van Allen radiation on EFT-1 are covered in the paper. Background on the Orion avionics subsystem is provided, as well as an overview of the GN&C software architecture. The measures taken to handle radiation induced failure of the one or both of the FCM's are presented, and finally simulation and actual hardware-in-the-loop (HWIL) results are shown confirming the validity of the implementation. The paper presents an overview of the Orion avionics architecture describing the GNC sensors, onboard data network as well as the flight control computers and their planned restart capabilities. GN&C sensors include two Orion Inertial Measurement Units (OIMU's), a Vision Processing Unit (VPU) to process camera images, three barometric altimeters and a single GPS receiver. All of the sensors communicate to one of two Power and Data Units (PDU's). The PDU's multiplex analog and serial data from the sensors and write the data to the Orion Data Network (ODN). The OIMU s write measurement messages directly as onto the ODN, but they are routed through PDU network switches.

King, Ellis T.↗

Transient Reliability Analysis Capability Developed for CARES/Life

The CARES/Life software developed at the NASA Glenn Research Center provides a general-purpose design tool that predicts the probability of the failure of a ceramic component as a function of its time in service. This award-winning software has been widely used by U.S. industry to establish the reliability and life of a brittle material (e.g., ceramic, intermetallic, and graphite) structures in a wide variety of 21st century applications.Present capabilities of the NASA CARES/Life code include probabilistic life prediction of ceramic components subjected to fast fracture, slow crack growth (stress corrosion), and cyclic fatigue failure modes. Currently, this code can compute the time-dependent reliability of ceramic structures subjected to simple time-dependent loading. For example, in slow crack growth failure conditions CARES/Life can handle sustained and linearly increasing time-dependent loads, whereas in cyclic fatigue applications various types of repetitive constant-amplitude loads can be accounted for. However, in real applications applied loads are rarely that simple but vary with time in more complex ways such as engine startup, shutdown, and dynamic and vibrational loads. In addition, when a given component is subjected to transient environmental and or thermal conditions, the material properties also vary with time. A methodology has now been developed to allow the CARES/Life computer code to perform reliability analysis of ceramic components undergoing transient thermal and mechanical loading. This means that CARES/Life will be able to analyze finite element models of ceramic components that simulate dynamic engine operating conditions. The methodology developed is generalized to account for material property variation (on strength distribution and fatigue) as a function of temperature. This allows CARES/Life to analyze components undergoing rapid temperature change in other words, components undergoing thermal shock. In addition, the capability has been developed to perform reliability analysis for components that undergo proof testing involving transient loads. This methodology was developed for environmentally assisted crack growth (crack growth as a function of time and loading), but it will be extended to account for cyclic fatigue (crack growth as a function of load cycles) as well.

Nemeth, Noel N.↗

Process membership in asynchronous environments

The development of reliable distributed software is simplified by the ability to assume a fail-stop failure model. The emulation of such a model in an asynchronous distributed environment is discussed. The solution proposed, called Strong-GMP, can be supported through a highly efficient protocol, and was implemented as part of a distributed systems software project at Cornell University. The precise definition of the problem, the protocol, correctness proofs, and an analysis of costs are addressed.

Ricciardi, Aleta M.↗