Search NASASearch

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Automated Addressing Failure Detection for Multiplexer Systems

Multiplexers are electronic devices which select between several input signals and deliver an output signal. Also known as data selectors or mux, multiplexers are commonly used in circuit design to provide signal switching, simplify hardware and manufacturing, and decrease cost. Systems which use a digital mux as a switch or selector are vulnerable to events called addressing failures, arising in hardware or software. An addressing failure results in the system requesting, or the mux delivering, a data signal other than the one intended. The unintended use of this erroneous signal may lead to system-level failures. This paper discusses a novel method of automatic detection for addressing failures, utilizing multiplexer channelization and a monitoring algorithm.

multiplexer

Man-rated flight software for the F-8 DFBW program

The design, implementation, and verification of the flight control software used in the F-8 DFBW program are discussed. Since the DFBW utilizes an Apollo computer and hardware, the procedures, controls, and basic management techniques employed are based on those developed for the Apollo software system. Program assembly control, simulator configuration control, erasable-memory load generation, change procedures and anomaly reporting are discussed. The primary verification tools are described, as well as the program test plans and their implementation on the various simulators. Failure effects analysis and the creation of special failure generating software for testing purposes are described.

Bairnsfather, R. R.

Autonomous failure detection and correction on Landsat-4

An integrated hardware/software fault tolerant system for an earth oriented, computer controlled spacecraft is described. The design philosophy as well as the rationale behind the chosen fault tolerant system is outlined. In-flight performance of the system is included for several different instances where the Failure Detection and Correction system acted autonomously to protect the spacecraft. This system exceeded the expectations of the designers by demonstrating the capability to provide a measure of safety to the spacecraft for inadvertent and undesirable ground commands as well as satisfying its primary function of monitoring the flight hardware and software for failures.

Welch, R. V.

A nonparametric software reliability growth model

Miller and Sofer have presented a nonparametric method for estimating the failure rate of a software program. The method is based on the complete monotonicity property of the failure rate function, and uses a regression approach to obtain estimates of the current software failure rate. This completely monotone software model is extended. It is shown how it can also provide long-range predictions of future reliability growth. Preliminary testing indicates that the method is competitive with parametric approaches, while being more robust.

Miller, Douglas R.

An innovative design for autonomous backup attitude control of the Gamma Ray Observatory

The Gamma Ray Observatory is a NASA funded three-axis stabilized spacecraft which will carry four scientific instruments to observe gamma ray phenomena. The requirement to protect the scientific mission from system failures led to the attitude control and determination system design described in this paper. The design employs nine control modes with error detection, hardware substitution, and autonomous mode switching. The system architecture evolved to eliminate cross-dependence between the primary on-board computer (OBC) and the backup control processor electronics. Cross strapping of sensors and actuators and separation of the input/output electronics ensure that a reliable set of sensors and actuators will be available for backup mode operation. The OBC software includes failure detection, hardware reconfiguration, and mode switching logic which provide the ability to autonomously transfer, upon anomaly, to a reliable backup mode. Verification of this mode transition design is done in four test programs: at the unit level, by analytical simulation, by a hybrid breadboard electronics-simulation setup, and by a flight hardware-simulation test.

Tai, F.

Failure-Modes-And-Effects Analysis Of Software Logic

Rigorous analysis applied early in design effort. Method of identifying potential inadequacies and modes and effects of failures caused by inadequacies (failure-modes-and-effects analysis or "FMEA" for short) devised for application to software logic.

Garcia, Danny

Wetware, Hardware, or Software Incapacitation: Observational Methods to Determine When Autonomy Should Assume Control

Control-theoretic modeling of human operator's dynamic behavior in manual control tasks has a long, rich history. There has been significant work on techniques used to identify the pilot model of a given structure. This research attempts to go beyond pilot identification based on experimental data to develop a predictor of pilot behavior. Two methods for pre-dicting pilot stick input during changing aircraft dynamics and deducing changes in pilot behavior are presented This approach may also have the capability to detect a change in a subject due to workload, engagement, etc., or the effects of changes in vehicle dynamics on the pilot. With this ability to detect changes in piloting behavior, the possibility now exists to mediate human adverse behaviors, hardware failures, and software anomalies with autono-my that may ameliorate these undesirable effects. However, appropriate timing of when au-tonomy should assume control is dependent on criticality of actions to safety, sensitivity of methods to accurately detect these adverse changes, and effects of changes in levels of auto-mation of the system as a whole.

Trujillo, Anna C.

General test plan redundant sensor strapdown IMU evaluation program

The general test plan for a redundant sensor strapdown inertial measuring unit evaluation program is presented. The inertial unit contains six gyros and three orthogonal accelerometers. The software incorporates failure detection and correction logic and a land vehicle navigation program. The principal objective of the test is a demonstration of the practicability, reliability, and performance of the inertial measuring unit with failure detection and correction in operational environments.

Hartwell, T.

JPL's Real-Time Weather Processor project (RWP) metrics and observations at system completion

As an integral part of the overall upgraded National Airspace System (NAS), the objective of the Real-Time Weather Processor (RWP) project is to improve the quality of weather information and the timeliness of its dissemination to system users. To accomplish this, an RWP will be installed in each of the Center Weather Service Units (CWSUs), located in 21 of the 23 Air Route Traffic Control Centers (ARTCCs). The RWP System is a prototype system. It is planned that the software will be GFE and that production hardware will be acquired via industry competitive procurement. The ARTCC is a facility established to provide air traffic control service to aircraft operating on Instrument Flight Rules (IFR) flight plans within controlled airspace, principally during the en route phase of the flight. Covered here are requirement metrics, Software Problem Failure Reports (SPFRs), and Ada portability metrics and observations.

Loesh, Robert E.

Rocket Science for the Internet

Rainfinity, a company resulting from the commercialization of Reliable Array of Independent Nodes (RAIN), produces the product, Rainwall. Rainwall runs a cluster of computer workstations, creating a distributed Internet gateway. When Rainwall detects a failure in software or hardware, traffic is shifted to a healthy gateway without interruptions to Internet service. It more evenly distributes workload across servers, providing less down time.

Source record

Applying Formal Methods and Object-Oriented Design to Existing Flight Software

This paper describes a project appling formal methods to a portion of the shuttle on-orbit digital autopilot (DAP). Three objectives of the project were to: demonstrate the use of formal methods on a shuttle application, facilitate the incorporation and validation of new requirements for the system, and verify the safety-critical properties to be exhibited by the software.

Software Failures

Parachute Testing for the NASA X-38 Crew Return Vehicle

NASA's X-38 program was an in-house technology demonstration program to develop a Crew Return Vehicle (CRV) for the International Space Station capable of returning seven crewmembers to Earth when the Space Shuttle was not present at the station. The program, managed out of NASA's Johnson Space Center, was started in 1995 and was cancelled in 2003. Eight flights with a prototype atmospheric vehicle were successfully flown at Edwards Air Force Base, demonstrating the feasibility of a parachute landing system for spacecraft. The intensive testing conducted by the program included testing of large ram-air parafoils. The flight test techniques, instrumentation, and simulation models developed during the parachute test program culminated in the successful demonstration of a guided parafoil system to land a 25,000 Ib spacecraft. The test program utilized parafoils of sizes ranging from 750 to 7500 p. The guidance, navigation, and control system (GN&C) consisted of winches, laser or radar altimeter, global positioning system (GPS), magnetic compass, barometric altimeter, flight computer, and modems for uplink commands and downlink data. The winches were used to steer the parafoil and to perform the dynamic flare maneuver for a soft landing. The laser or radar altimeter was used to initiate the flare. In the event of a GPS failure, the software navigated by dead reckoning using the compass and barometric altimeter data. The GN&C test beds included platforms dropped from cargo aircraft, atmospheric vehicles released from a 8-52, and a Buckeye powered parachute. This paper will describe the test program and significant results.

Stein, Jenny M.

Introduction of Virtualization Technology to Multi-Process Model Checking

Model checkers find failures in software by exploring every possible execution schedule. Java PathFinder (JPF), a Java model checker, has been extended recently to cover networked applications by caching data transferred in a communication channel. A target process is executed by JPF, whereas its peer process runs on a regular virtual machine outside. However, non-deterministic target programs may produce different output data in each schedule, causing the cache to restart the peer process to handle the different set of data. Virtualization tools could help us restore previous states of peers, eliminating peer restart. This paper proposes the application of virtualization technology to networked model checking, concentrating on JPF.

Leungwattanakit, Watcharin

Software Risk Identification for Interplanetary Probes

The need for a systematic and effective software risk identification methodology is critical for interplanetary probes that are using increasingly complex and critical software. Several probe failures are examined that suggest more attention and resources need to be dedicated to identifying software risks. The direct causes of these failures can often be traced to systemic problems in all phases of the software engineering process. These failures have lead to the development of a practical methodology to identify risks for interplanetary probes. The proposed methodology is based upon the tailoring of the Software Engineering Institute's (SEI) method of taxonomy-based risk identification. The use of this methodology will ensure a more consistent and complete identification of software risks in these probes.

Dougherty, Robert J.

Free-Swinging Failure Tolerance for Robotic Manipulators

Under this GSRP fellowship, software-based failure-tolerance techniques were developed for robotic manipulators. The focus was on failures characterized by the loss of actuator torque at a joint, called free-swinging failures. The research results spanned many aspects of the free-swinging failure-tolerance problem, from preparing for an expected failure to discovery of postfailure capabilities to establishing efficient methods to realize those capabilities. Developed algorithms were verified using computer-based dynamic simulations, and these were further verified using hardware experiments at Johnson Space Center.

English, James

Free-Swinging Failure Tolerance for Robotic Manipulators

Under this GSRP fellowship, software-based failure-tolerance techniques were developed for robotic manipulators. The focus was on failures characterized by the loss of actuator torque at a joint, called free-swinging failures. The research results spanned many aspects of the free-swinging failure-tolerance problem, from preparing for an expected failure to discovery of postfailure capabilities to establishing efficient methods to realize those capabilities. Developed algorithms were verified using computer-based dynamic simulations, and these were further verified using hardware experiments at Johnson Space Center.

English, James

Estimating Software Reliability for Space Launch Vehicles in Probabilistic Risk Assessment (PRA)

It is acutely recognized in the Probabilistic Risk Assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless of whether the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the National Aeronautics and Space Administration (NASA) PRA and present a future dynamic option for software and PRA. Space Launch Vehicle software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data. Software uncertainty will also be discussed. It is determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.

Steven D. Novack

Estimating Software Reliability for Space Launch Vehicles in Probabilistic Risk Assessment (PRA)

It is acutely recognized in the Probabilistic Risk assessment (PRA) field that software plays a defining role in overall system reliability for all modern systems across a wide variety of industries. Regardless if the software is embedded firmware for working components or elements, part of a Human-Machine-Interface, or automated command and control logic, the success of the software to fulfill its function under nominal and off-nominal environments will be a dominant contributor to system reliability. It is also recognized that software reliability prediction and estimation is one of the more challenging and questionable aspects of any PRA or system analyses due to the nature of software and its integration with physics based systems. Irrespective of this dichotomy, any incorporation of software reliability methods requires that the contributions are accountable, quantitative, and tractable. This paper provides a brief overview of software reliability methods, establishes some minimum requirements that the methods should incorporate for completeness, and provides a logic structure for applying software reliability. Model resolution will be discussed that supports current testing plans and trade studies. We will provide initial recommendations for use in the NASA PRA and present a future dynamic option for software and PRA. Space Launch Vehicle Software is recognized to be reliable in static conditions, yet relatively vulnerable to a set of failure modes in changing environments/flight phases. Two quantitative methods were chosen to incorporate software reliability into a Space Launch Vehicle PRA accounting for phase adjustments. One method predicts latent software failure using statistical methods, and the second provides estimates of coding errors and software operating system failures based on test and historical data, respectively. Software uncertainty will also be discussed. We determined that recommendations for PRA software reliability should be modeled at the software module level where multiple software components compose a module and combinations of the software architecture can lead to a functional failure.

Novack, Steven