Search NASA⌕ Search

SEARCH · Search NASA

Results for “Software Failures”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

[A Handling Qualities Metric for Damaged Aircraft]

In recent flight tests of F-15 Intelligent Flight Control System (IFCS), software simulated aircraft control surface failures were inserted to evaluate the IFCS adaptive systems. The failure commanded the left stabilator to a fixed position. The adaptive system uses a neural network that is designed to change control law gains, in the event of damage (real or simulated), that allows the aircraft to fly as it had before the damage. The performance of the adaptive system was assessed in terms of its ability to re-establish good onboard model tracking and its ability to decouple roll and pitch response.

Cogan, Bruce↗

Noninvasive Diagnosis of Coronary Artery Disease Using 12-Lead High-Frequency Electrocardiograms

A noninvasive, sensitive method of diagnosing certain pathological conditions of the human heart involves computational processing of digitized electrocardiographic (ECG) signals acquired from a patient at all 12 conventional ECG electrode positions. In the processing, attention is focused on low-amplitude, high-frequency components of those portions of the ECG signals known in the art as QRS complexes. The unique contribution of this method lies in the utilization of signal features and combinations of signal features from various combinations of electrode positions, not reported previously, that have been found to be helpful in diagnosing coronary artery disease and such related pathological conditions as myocardial ischemia, myocardial infarction, and congestive heart failure. The electronic hardware and software used to acquire the QRS complexes and perform some preliminary analyses of their high-frequency components were summarized in Real-Time, High-Frequency QRS Electrocardiograph (MSC- 23154), NASA Tech Briefs, Vol. 27, No. 7 (July 2003), pp. 26-28. To recapitulate, signals from standard electrocardiograph electrodes are preamplified, then digitized at a sampling rate of 1,000 Hz, then analyzed by the software that detects R waves and QRS complexes and analyzes them from several perspectives. The software includes provisions for averaging signals over multiple beats and for special-purpose nonrecursive digital filters with specific low- and high-frequency cutoffs. These filters, applied to the averaged signal, effect a band-pass operation in the frequency range from 150 to 250 Hz. The output of the bandpass filter is the desired high-frequency QRS signal. Further processing is then performed in real time to obtain the beat-to-beat root mean square (RMS) voltage amplitude of the filtered signal, certain variations of the RMS voltage, and such standard measures as the heart rate and R-R interval at any given time. A key signal feature analyzed in the present method is the presence versus the absence of reduced-amplitude zones (RAZs). In terms that must be simplified for the sake of brevity, an RAZ comprises several cycles of a high-frequency QRS signal during which the amplitude of the high-frequency oscillation in a portion of the signal is abnormally low (see figure). A given signal sample exhibiting an interval of reduced amplitude may or may not be classified as an RAZ, depending on quantitative criteria regarding peaks and troughs within the reduced-amplitude portion of the high-frequency QRS signal. This analysis is performed in all 12 leads in real time.

Schlegel, Todd T.↗

Projected Impact of Compositional Verification on Current and Future Aviation Safety Risk

The projected impact of compositional verification research conducted by the National Aeronautic and Space Administration System-Wide Safety and Assurance Technologies on aviation safety risk was assessed. Software and compositional verification was described. Traditional verification techniques have two major problems: testing at the prototype stage where error discovery can be quite costly and the inability to test for all potential interactions leaving some errors undetected until used by the end user. Increasingly complex and nondeterministic aviation systems are becoming too large for these tools to check and verify. Compositional verification is a "divide and conquer" solution to addressing increasingly larger and more complex systems. A review of compositional verification research being conducted by academia, industry, and Government agencies is provided. Forty-four aviation safety risks in the Biennial NextGen Safety Issues Survey were identified that could be impacted by compositional verification and grouped into five categories: automation design; system complexity; software, flight control, or equipment failure or malfunction; new technology or operations; and verification and validation. One capability, 1 research action, 5 operational improvements, and 13 enablers within the Federal Aviation Administration Joint Planning and Development Office Integrated Work Plan that could be addressed by compositional verification were identified.

Reveley, Mary S.↗

The Raid distributed database system

Raid, a robust and adaptable distributed database system for transaction processing (TP), is described. Raid is a message-passing system, with server processes on each site to manage concurrent processing, consistent replicated copies during site failures, and atomic distributed commitment. A high-level layered communications package provides a clean location-independent interface between servers. The latest design of the package delivers messages via shared memory in a configuration with several servers linked into a single process. Raid provides the infrastructure to investigate various methods for supporting reliable distributed TP. Measurements on TP and server CPU time are presented, along with data from experiments on communications software, consistent replicated copy control during site failures, and concurrent distributed checkpointing. A software tool for evaluating the implementation of TP algorithms in an operating-system kernel is proposed.

Bhargava, Bharat↗

Pivotal-Function Assessment Of Reliability Of Software

Approach developed to establish utility of pivotal functions for estimation and prediction of reliability of software. Improved estimates of reliability with statistical confidence obtained when relatively few testing data available. Pivotal functions effective tools for determination of confidence limits for reliability of software and prediction limits for time to next failure. Provides exact confidence and prediction limits regardless of how many bugs found in software.

Hayhurst, Kelly J.↗

NASA's Software Safety Standard

NASA relies more and more on software to control, monitor, and verify its safety critical systems, facilities and operations. Since the 1960's there has hardly been a spacecraft launched that does not have a computer on board that will provide command and control services. There have been recent incidents where software has played a role in high-profile mission failures and hazardous incidents. For example, the Mars Orbiter, Mars Polar Lander, the DART (Demonstration of Autonomous Rendezvous Technology), and MER (Mars Exploration Rover) Spirit anomalies were all caused or contributed to by software. The Mission Control Centers for the Shuttle, ISS, and unmanned programs are highly dependant on software for data displays, analysis, and mission planning. Despite this growing dependence on software control and monitoring, there has been little to no consistent application of software safety practices and methodology to NASA's projects with safety critical software. Meanwhile, academia and private industry have been stepping forward with procedures and standards for safety critical systems and software, for example Dr. Nancy Leveson's book Safeware: System Safety and Computers. The NASA Software Safety Standard, originally published in 1997, was widely ignored due to its complexity and poor organization. It also focused on concepts rather than definite procedural requirements organized around a software project lifecycle. Led by NASA Headquarters Office of Safety and Mission Assurance, the NASA Software Safety Standard has recently undergone a significant update. This new standard provides the procedures and guidelines for evaluating a project for safety criticality and then lays out the minimum project lifecycle requirements to assure the software is created, operated, and maintained in the safest possible manner. This update of the standard clearly delineates the minimum set of software safety requirements for a project without detailing the implementation for those requirements. This allows the projects leeway to meet these requirements in many forms that best suit a particular project's needs and safety risk. In other words, it tells the project what to do, not how to do it. This update also incorporated advances in the state of the practice of software safety from academia and private industry. It addresses some of the more common issues now facing software developers in the NASA environment such as the use of Commercial-Off-the-Shelf Software (COTS), Modified OTS (MOTS), Government OTS (GOTS), and reused software. A team from across NASA developed the update and it has had both NASA-wide internal reviews by software engineering, quality, safety, and project management. It has also had expert external review. This presentation and paper will discuss the new NASA Software Safety Standard, its organization, and key features. It will start with a brief discussion of some NASA mission failures and incidents that had software as one of their root causes. It will then give a brief overview of the NASA Software Safety Process. This will include an overview of the key personnel responsibilities and functions that must be performed for safety-critical software.

Ramsay, Christopher M.↗

Information Extraction for System-Software Safety Analysis: Calendar Year 2008 Year-End Report

This annual report describes work to integrate a set of tools to support early model-based analysis of failures and hazards due to system-software interactions. The tools perform and assist analysts in the following tasks: 1) extract model parts from text for architecture and safety/hazard models; 2) combine the parts with library information to develop the models for visualization and analysis; 3) perform graph analysis and simulation to identify and evaluate possible paths from hazard sources to vulnerable entities and functions, in nominal and anomalous system-software configurations and scenarios; and 4) identify resulting candidate scenarios for software integration testing. There has been significant technical progress in model extraction from Orion program text sources, architecture model derivation (components and connections) and documentation of extraction sources. Models have been derived from Internal Interface Requirements Documents (IIRDs) and FMEA documents. Linguistic text processing is used to extract model parts and relationships, and the Aerospace Ontology also aids automated model development from the extracted information. Visualizations of these models assist analysts in requirements overview and in checking consistency and completeness.

Malin, Jane T.↗

Tools Ensure Reliability of Critical Software

In November 2006, after attempting to make a routine maneuver, NASA's Mars Global Surveyor (MGS) reported unexpected errors. The onboard software switched to backup resources, and a 2-day lapse in communication took place between the spacecraft and Earth. When a signal was finally received, it indicated that MGS had entered safe mode, a state of restricted activity in which the computer awaits instructions from Earth. After more than 9 years of successful operation gathering data and snapping pictures of Mars to characterize the planet's land and weather communication between MGS and Earth suddenly stopped. Months later, a report from NASA's internal review board found the spacecraft's battery failed due to an unfortunate sequence of events. Updates to the spacecraft's software, which had taken place months earlier, were written to the wrong memory address in the spacecraft's computer. In short, the mission ended because of a software defect. Over the last decade, spacecraft have become increasingly reliant on software to carry out mission operations. In fact, the next mission to Mars, the Mars Science Laboratory, will rely on more software than all earlier missions to Mars combined. According to Gerard Holzmann, manager at the Laboratory for Reliable Software (LaRS) at NASA's Jet Propulsion Laboratory (JPL), even the fault protection systems on a spacecraft are mostly software-based. For reasons like these, well-functioning software is critical for NASA. In the same year as the failure of MGS, Holzmann presented a new approach to critical software development to help reduce risk and provide consistency. He proposed The Power of 10: Rules for Developing Safety-Critical Code, which is a small set of rules that can easily be remembered, clearly relate to risk, and allow compliance to be verified. The reaction at JPL was positive, and developers in the private sector embraced Holzmann's ideas.

Source record↗

Medium Fidelity Simulation of Oxygen Tank Venting

The item to he cleared is a medium-fidelity software simulation model of a vented cryogenic tank. Such tanks are commonly used to transport cryogenic liquids such as liquid oxygen via truck, and have appeared on liquid-fueled rockets for decades. This simulation model works with the HCC simulation system that was developed by Xerox PARC and NASA Ames Research Center. HCC has been previously cleared for distribution. When used with the HCC software, the model generates simulated readings for the tank pressure and temperature as the simulated cryogenic liquid boils off and is vented. Failures (such as a broken vent valve) can be injected into the simulation to produce readings corresponding to the failure. Release of this simulation will allow researchers to test their software diagnosis systems by attempting to diagnose the simulated failure from the simulated readings. This model does not contain any encryption software nor can it perform any control tasks that might be export controlled.

Sweet, Adam↗

Orion Burn Management, Nominal and Response to Failures

An approach for managing Orion on-orbit burn execution is described for nominal and failure response scenarios. The burn management strategy for Orion takes into account per-burn variations in targeting, timing, and execution; crew and ground operator intervention and overrides; defined burn failure triggers and responses; and corresponding on-board software sequencing functionality. Burn-to- burn variations are managed through the identification of specific parameters that may be updated for each progressive burn. Failure triggers and automatic responses during the burn timeframe are defined to provide safety for the crew in the case of vehicle failures, along with override capabilities to ensure operational control of the vehicle. On-board sequencing software provides the timeline coordination for performing the required activities related to targeting, burn execution, and responding to burn failures.

Odegard, Ryan↗

Sensory redundancy management: The development of a design methodology for determining threshold values through a statistical analysis of sensor output data

Sensor redundancy management (SRM) requires a system which will detect failures and reconstruct avionics accordingly. A probability density function to determine false alarm rates, using an algorithmic approach was generated. Microcomputer software was developed which will print out tables of values for the cummulative probability of being in the domain of failure; system reliability; and false alarm probability, given a signal is in the domain of failure. The microcomputer software was applied to the sensor output data for various AFT1 F-16 flights and sensor parameters. Practical recommendations for further research were made.

Scalzo, F.↗

Optimal integral controller with sensor failure accommodation

An Optimal Integral Controller that readily accommodates Sensor Failure - without resorting to (Kalman) filter or observer generation - has been designed. The system is based on Navy-sponsored research for the control of high performance aircraft. In conjunction with a NASA developed Numerical Optimization Code, the Integral Feedback Controller will provide optimal system response even in the case of incomplete state feedback. Hence, the need for costly replication of plant sensors is avoided since failure accommodation is effected by system software reconfiguration. The control design has been applied to a particularly ill-behaved, third-order system. Dominant-root design in the classical sense produced an almost 100 percent overshoot for the third-order system response. An application of the newly-developed Optimal Integral Controller - assuming all state information available - produces a response with no overshoot. A further application of the controller design - assuming a one-third sensor failure scenario - produced a slight overshoot response that still preserved the steady state time-point of the full-state feedback response. The control design should have wide application in space systems.

Alberts, T.↗

A Testbed for Evaluating Lunar Habitat Autonomy Architectures

A lunar outpost will involve a habitat with an integrated set of hardware and software that will maintain a safe environment for human activities. There is a desire for a paradigm shift whereby crew will be the primary mission operators, not ground controllers. There will also be significant periods when the outpost is uncrewed. This will require that significant automation software be resident in the habitat to maintain all system functions and respond to faults. JSC is developing a testbed to allow for early testing and evaluation of different autonomy architectures. This will allow evaluation of different software configurations in order to: 1) understand different operational concepts; 2) assess the impact of failures and perturbations on the system; and 3) mitigate software and hardware integration risks. The testbed will provide an environment in which habitat hardware simulations can interact with autonomous control software. Faults can be injected into the simulations and different mission scenarios can be scripted. The testbed allows for logging, replaying and re-initializing mission scenarios. An initial testbed configuration has been developed by combining an existing life support simulation and an existing simulation of the space station power distribution system. Results from this initial configuration will be presented along with suggested requirements and designs for the incremental development of a more sophisticated lunar habitat testbed.

Lawler, Dennis G.↗

Electrified Aircraft Propulsion Systems: Potential Failure Modes and Failure Mitigation Strategies

Electrified aircraft propulsion (EAP) systems hold great potential for the reduction of aircraft fuel burn, emissions, and noise. Currently, NASA and other organizations are actively working to identify and mature technologies necessary to bring EAP designs to reality. A requirement for the development of any civil aircraft and its systems is to ensure that potential hazards in the design are identified and appropriately mitigated to ensure that the system is safe. During aircraft development, a system safety assessment that consists of a functional hazard assessment is conducted to identify all potential failure conditions of each function, and classify those failures according to the severity of their effects on the aircraft or its occupants. The more severe a function's failure condition classification, the greater the development assurance level required for the function to ensure that the probability of the hazard is acceptably low. Today, aircraft engines and their control systems receive type certificate approval as a stand-alone system to signify their airworthiness. However, the complex coupling and distributed nature of EAP designs are expected to place added challenges on the certification of these systems. This presentation will provide an initial high-level review of the potential failure modes and hazards posed by a generic EAP system along with potential mitigation strategies for those failures. The EAP system is assumed to be a hybrid design consisting of gas turbine engines, mechanical drives, electric machines, power electronics and distribution systems, energy storage devices, and motor driven propulsors. The functionality provided by each of these EAP subsystems will be discussed along with the potential failure modes they may encounter. This will include a discussion of coupled failure effects, where a fault in one EAP subsystem effects the operation of other subsystems in the architecture. Next, potential failure mitigation strategies are discussed including both software-based and hardware-based mitigation strategies. The presentation will conclude with an example evaluation of the potential failure modes and mitigation strategies for a concept EAP system proposed by NASA.

Simon, Donald L.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failures. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certification processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E., Jr.↗

Spacecraft Software Maintenance: An Effective Approach to Reducing Costs and Increasing Science Return

Flight software is a mission critical element of spacecraft functionality and performance. When ground operations personnel interface to a spacecraft, they are typically dealing almost entirely with the capabilities of onboard software. This software, even more than critical ground/flight communications systems, is expected to perform perfectly during all phases of spacecraft life. Due to the fact that it can be reprogrammed on-orbit to accommodate degradations or failures in flight hardware, new insights into spacecraft characteristics, new control options which permit enhanced science options, etc., the on- orbit flight software maintenance team is usually significantly responsible for the long term success of a science mission. Failure of flight software to perform as needed can result in very expensive operations work-around costs and lost science opportunities. There are three basic approaches to maintaining spacecraft software--namely using the original developers, using the mission operations personnel, or assembling a center of excellence for multi-spacecraft software maintenance. Not planning properly for flight software maintenance can lead to unnecessarily high on-orbit costs and/or unacceptably long delays, or errors, in patch installations. A common approach for flight software maintenance is to access the original development staff. The argument for utilizing the development staff is that the people who developed the software will be the best people to modify the software on-orbit. However, it can quickly becomes a challenge to obtain the services of these key people. They may no longer be available to the organization. They may have a more urgent job to perform, quite likely on another project under different project management. If they havn't worked on the software for a long time, they may need precious time for refamiliarization to the software, testbeds and tools. Further, a lack of insight into issues related to flight software in its on-orbit environment, may find the developer unprepared for the challenges. The second approach is to train a member of the flight operations team to maintain the spacecraft software. This can prove to be a costly and inflexible solution. The person assigned to this duty may not have enough work to do during a problem free period and may have too much to do when a problem arises. If the person is a talented software engineer, he/she may not enjoy the limited software opportunities available in this position; and may eventually leave for newer technology computer science opportunities. Training replacement flight software personnel can be a difficult and lengthy process. The third approach is to assemble a center of excellence for on-orbit spacecraft software maintenance. Personnel in this specialty center can be managed to support flight software of multiple missions at once. The variety of challenges among a set of on-orbit missions, can result in a dedicated, talented staff which is fully trained and available to support each mission's needs. Such staff are not software developers but are rather spacecraft software systems engineers. The cost to any one mission is extremely low because the software staff works and charges, minimally on missions with no current operations issues; and their professional insight into on-orbit software troubleshooting and maintenance methods ensures low risk, effective and minimal-cost solutions to on-orbit issues.

Shell, Elaine M.↗

A methodology for validating software reliability

A significant problem associated with fault tolerant computer system design is how to insure that there are no embedded software errors, so that an avionics computer system meets the required reliability level. To accomplish this, it is necessary to associate a 'probability of failure' with the operational flight program. It would be more correct to say that the probability of excitation of existing latent design errors within the program is required. In this sense, latent software errors are like latent hardware faults, and techniques that were previously used to measure the probability of failure of hardware due to fault latency can be used to measure the probability of failure of the software. A methodology was developed and applied to a flight control program that was known to operate in a well defined environment. The results indicated that the technique could be used to provide a final validation of the software to a specified reliability level and to evaluate the role of flight test in software validation.

Swern, Frederic L.↗