Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Case Study of Using High Performance Commercial Processors in Space

The purpose of the Space Shuttle Cockpit Avionics Upgrade project (1999 2004) was to reduce crew workload and improve situational awareness. The upgrade was to augment the Shuttle avionics system with new hardware and software. A major success of this project was the validation of the hardware architecture and software design. This was significant because the project incorporated new technology and approaches for the development of human rated space software. An early version of this system was tested at the Johnson Space Center for one month by teams of astronauts. The results were positive, but NASA eventually cancelled the project towards the end of the development cycle. The goal to reduce crew workload and improve situational awareness resulted in the need for high performance Central Processing Units (CPUs). The choice of CPU selected was the PowerPC family, which is a reduced instruction set computer (RISC) known for its high performance. However, the requirement for radiation tolerance resulted in the re-evaluation of the selected family member of the PowerPC line. Radiation testing revealed that the original selected processor (PowerPC 7400) was too soft to meet mission objectives and an effort was established to perform trade studies and performance testing to determine a feasible candidate. At that time, the PowerPC RAD750s were radiation tolerant, but did not meet the required performance needs of the project. Thus, the final solution was to select the PowerPC 7455. This processor did not have a radiation tolerant version, but had some ability to detect failures. However, its cache tags did not provide parity and thus the project incorporated a software strategy to detect radiation failures. The strategy was to incorporate dual paths for software generating commands to the legacy Space Shuttle avionics to prevent failures due to the softness of the upgraded avionics.

Ferguson, Roscoe C.↗

Case Study of Using High Performance Commercial Processors in a Space Environment

The purpose of the Space Shuttle Cockpit Avionics Upgrade project was to reduce crew workload and improve situational awareness. The upgrade was to augment the Shuttle avionics system with new hardware and software. A major success of this project was the validation of the hardware architecture and software design. This was significant because the project incorporated new technology and approaches for the development of human rated space software. An early version of this system was tested at the Johnson Space Center for one month by teams of astronauts. The results were positive, but NASA eventually cancelled the project towards the end of the development cycle. The goal to reduce crew workload and improve situational awareness resulted in the need for high performance Central Processing Units (CPUs). The choice of CPU selected was the PowerPC family, which is a reduced instruction set computer (RISC) known for its high performance. However, the requirement for radiation tolerance resulted in the reevaluation of the selected family member of the PowerPC line. Radiation testing revealed that the original selected processor (PowerPC 7400) was too soft to meet mission objectives and an effort was established to perform trade studies and performance testing to determine a feasible candidate. At that time, the PowerPC RAD750s where radiation tolerant, but did not meet the required performance needs of the project. Thus, the final solution was to select the PowerPC 7455. This processor did not have a radiation tolerant version, but faired better than the 7400 in the ability to detect failures. However, its cache tags did not provide parity and thus the project incorporated a software strategy to detect radiation failures. The strategy was to incorporate dual paths for software generating commands to the legacy Space Shuttle avionics to prevent failures due to the softness of the upgraded avionics.

Ferguson, Roscoe C.↗

Modeling and Performance Considerations for Automated Fault Isolation in Complex Systems

The purpose of this paper is to document the modeling considerations and performance metrics that were examined in the development of a large-scale Fault Detection, Isolation and Recovery (FDIR) system. The FDIR system is envisioned to perform health management functions for both a launch vehicle and the ground systems that support the vehicle during checkout and launch countdown by using suite of complimentary software tools that alert operators to anomalies and failures in real-time. The FDIR team members developed a set of operational requirements for the models that would be used for fault isolation and worked closely with the vendor of the software tools selected for fault isolation to ensure that the software was able to meet the requirements. Once the requirements were established, example models of sufficient complexity were used to test the performance of the software. The results of the performance testing demonstrated the need for enhancements to the software in order to meet the demands of the full-scale ground and vehicle FDIR system. The paper highlights the importance of the development of operational requirements and preliminary performance testing as a strategy for identifying deficiencies in highly scalable systems and rectifying those deficiencies before they imperil the success of the project

Ferrell, Bob↗

A theoretical basis for the analysis of redundant software subject to coincident errors

Fundamental to the development of redundant software techniques fault-tolerant software, is an understanding of the impact of multiple-joint occurrences of coincident errors. A theoretical basis for the study of redundant software is developed which provides a probabilistic framework for empirically evaluating the effectiveness of the general (N-Version) strategy when component versions are subject to coincident errors, and permits an analytical study of the effects of these errors. The basic assumptions of the model are: (1) independently designed software components are chosen in a random sample; and (2) in the user environment, the system is required to execute on a stationary input series. The intensity of coincident errors, has a central role in the model. This function describes the propensity to introduce design faults in such a way that software components fail together when executing in the user environment. The model is used to give conditions under which an N-Version system is a better strategy for reducing system failure probability than relying on a single version of software. A condition which limits the effectiveness of a fault-tolerant strategy is studied, and it is posted whether system failure probability varies monotonically with increasing N or whether an optimal choice of N exists.

Eckhardt, D. E., Jr.↗

Development of a Two-Wheel Contingency Mode for the MAP Spacecraft

In the event of a failure of one of MAP's three reaction wheel assemblies (RWAs), it is not possible to achieve three-axis, full-state attitude control using the remaining two wheels. Hence, two of the attitude control algorithms implemented on the MAP spacecraft will no longer be usable in their current forms: Inertial Mode, used for slewing to and holding inertial attitudes, and Observing Mode, which implements the nominal dual-spin science mode. This paper describes the effort to create a complete strategy for using software algorithms to cope with a RWA failure. The discussion of the design process will be divided into three main subtopics: performing orbit maneuvers to reach and maintain an orbit about the second Earth-Sun libration point in the event of a RWA failure, completing the mission using a momentum-bias two-wheel science mode, and developing a new thruster-based mode for adjusting the inertially fixed momentum bias. In this summary, the philosophies used in designing these changes is shown; the full paper will supplement these with algorithm descriptions and testing results.

Starin, Scott R.↗

PISCES: A Tool for Predicting Software Testability

Before a program can fail, a software fault must be executed, that execution must alter the data state, and the incorrect data state must propagate to a state that results directly in an incorrect output. This paper describes a tool called PISCES (developed by Reliable Software Technologies Corporation) for predicting the probability that faults in a particular program location will accomplish all three of these steps causing program failure. PISCES is a tool that is used during software verification and validation to predict a program's testability.

Voas, Jeffrey M.↗

Numerical Implementation of a Multiple-ISV Thermodynamically-Based Work Potential Theory for Modeling Progressive Damage and Failure in Fiber-Reinforced Laminates

A thermodynamically-based work potential theory for modeling progressive damage and failure in fiber-reinforced laminates is presented. The current, multiple-internal state variable (ISV) formulation, enhanced Schapery theory (EST), utilizes separate ISVs for modeling the effects of damage and failure. Damage is considered to be the effect of any structural changes in a material that manifest as pre-peak non-linearity in the stress versus strain response. Conversely, failure is taken to be the effect of the evolution of any mechanisms that results in post-peak strain softening. It is assumed that matrix microdamage is the dominant damage mechanism in continuous fiber-reinforced polymer matrix laminates, and its evolution is controlled with a single ISV. Three additional ISVs are introduced to account for failure due to mode I transverse cracking, mode II transverse cracking, and mode I axial failure. Typically, failure evolution (i.e., post-peak strain softening) results in pathologically mesh dependent solutions within a finite element method (FEM) setting. Therefore, consistent character element lengths are introduced into the formulation of the evolution of the three failure ISVs. Using the stationarity of the total work potential with respect to each ISV, a set of thermodynamically consistent evolution equations for the ISVs is derived. The theory is implemented into commercial FEM software. Objectivity of total energy dissipated during the failure process, with regards to refinements in the FEM mesh, is demonstrated. The model is also verified against experimental results from two laminated, T800/3900-2 panels containing a central notch and different fiber-orientation stacking sequences. Global load versus displacement, global load versus local strain gage data, and macroscopic failure paths obtained from the models are compared to the experiments.

Pineda, Evan J.↗

A Thermodynamically-Based Mesh Objective Work Potential Theory for Predicting Intralaminar Progressive Damage and Failure in Fiber-Reinforced Laminates

A thermodynamically-based work potential theory for modeling progressive damage and failure in fiber-reinforced laminates is presented. The current, multiple-internal state variable (ISV) formulation, enhanced Schapery theory (EST), utilizes separate ISVs for modeling the effects of damage and failure. Damage is considered to be the effect of any structural changes in a material that manifest as pre-peak non-linearity in the stress versus strain response. Conversely, failure is taken to be the effect of the evolution of any mechanisms that results in post-peak strain softening. It is assumed that matrix microdamage is the dominant damage mechanism in continuous fiber-reinforced polymer matrix laminates, and its evolution is controlled with a single ISV. Three additional ISVs are introduced to account for failure due to mode I transverse cracking, mode II transverse cracking, and mode I axial failure. Typically, failure evolution (i.e., post-peak strain softening) results in pathologically mesh dependent solutions within a finite element method (FEM) setting. Therefore, consistent character element lengths are introduced into the formulation of the evolution of the three failure ISVs. Using the stationarity of the total work potential with respect to each ISV, a set of thermodynamically consistent evolution equations for the ISVs is derived. The theory is implemented into commercial FEM software. Objectivity of total energy dissipated during the failure process, with regards to refinements in the FEM mesh, is demonstrated. The model is also verified against experimental results from two laminated, T800/3900-2 panels containing a central notch and different fiber-orientation stacking sequences. Global load versus displacement, global load versus local strain gage data, and macroscopic failure paths obtained from the models are compared to the experiments.

Pineda, Evan J.↗

Design Methods, Tools, and Data for Ceramic Solar Receivers

This report presents the development of tools and methods for evaluating the reliability and performance of ceramic materials in high temperature solar receivers. As Concentrating Solar Power (CSP) technologies aim for higher operating temperatures to enhance efficiency and meet industrial process heat requirements, current high temperature metallic materials face challenges in maintaining structural integrity. This report explores advanced ceramics as a promising alternative, given their superior high temperature strength and lower thermal expansion, compared to metals. To address the need for effective ceramic receiver design tools, this report integrates statistical failure models of ceramics into the existing srlife tool: an open-source software package designed to estimate the life of high temperature CSP components. These failure models account for the inherent variability and flaw distribution in ceramics, as well as the impact of subcritical crack growth under high temperature cyclic loads. The report also presents experimental data collected for a commercially available ceramic material, SiC, and details the process of estimating reliability model parameters from these data. A comparative design analysis is then performed between ceramic (SiC) and metallic (current nickel-based superalloys A740H and A282) receiver. This comparison demonstrates that SiC receivers can achieve service life exceeding 30 years under high incident heat flux conditions, compared to just a few years for metallic receivers.

14 SOLAR ENERGY↗

Predicting System Accidents with Model Analysis During Hybrid Simulation

Standard discrete event simulation is commonly used to identify system bottlenecks and starving and blocking conditions in resources and services. The CONFIG hybrid discrete/continuous simulation tool can simulate such conditions in combination with inputs external to the simulation. This provides a means for evaluating the vulnerability to system accidents of a system's design, operating procedures, and control software. System accidents are brought about by complex unexpected interactions among multiple system failures , faulty or misleading sensor data, and inappropriate responses of human operators or software. The flows of resource and product materials play a central role in the hazardous situations that may arise in fluid transport and processing systems. We describe the capabilities of CONFIG for simulation-time linear circuit analysis of fluid flows in the context of model-based hazard analysis. We focus on how CONFIG simulates the static stresses in systems of flow. Unlike other flow-related properties, static stresses (or static potentials) cannot be represented by a set of state equations. The distribution of static stresses is dependent on the specific history of operations performed on a system. We discuss the use of this type of information in hazard analysis of system designs.

Malin, Jane T.↗

Flow Testing of Corrugated Metal Flexhoses to Evaluate Flow-Induced Vibration and Stiffness

Corrugated metal flexhoses are used to supply fluid routing where straight rigid pipes cannot meet the design requirements due to vibrations, thermal expansion, or motion. These types of hoses are found in a wide range of applications at KSC (Exploration Ground Systems Program in particular) with various fluid commodities such as fuel, oxidizers, coolant, and cryogenics. However, flexibility of design comes at a price: due to increased levels of turbulence generated by the hose geometry, the necessary supply pressure must increase to achieve the same flow rate. Furthermore, the convolutes of the flexhoses interact with the flow field to generate areas of flow separation which leads to a phenomenon known as vortex shedding. When the frequency of vortex shedding couples with the natural frequency of the hose, this can be detrimental and cause premature failure. There are a limited amount of software that can correctly model the coupled fluid structure interactions(FSI) between the solid and fluid physics. This project aims to develop a computer model that can predict Flow-Induced Vibration (FIV) on flexhoses in one of the atypical configurations found in the State-of-the-Art (SOTA) standard for hoses in an angulated state. Our research seeks to extend the literature and expand NASA’s FIV standard. The computational rigor and resources available in the Apollo Era did not allow for flow coupling between fluids and structure interfaces to evaluate FIV of the flexhose. Our team believes that modern techniques and methods of evaluation should be used to reevaluate and extend the database of FIV and stiffness properties of flexhoses to assess the risk of failure. The core performance criteria is a computer model that is able to predict the FIV frequency within 10%.

Jared F. Congiardo↗

Chandra Space Flight Software: Using Software to Autonomously Operation the Largest and Most Sensitive X-Ray Telescope in the World

Chandra is the world's largest and most sensitive X-ray telescope. The Chandra X-ray Observatory is the third in NASA's family of "Great Observatories." The Chandra X-ray Observatory, launched by Space Shuttle Columbia on July 23, 1999, is NASA's newest Great Observatory. The Chandra space flight software is the operational software, which controls and directs the Chandra X-ray Observatory. The Chandra flight software has executed faultlessly for over 13,000 hours on-orbit. The Chandra flight software directly controls the Pointing, Aspect Determination, Electrical Power Subsystem, Propulsion system, and the Command, Communications, and Data Management subsystems. The software controls the spacecraft operations during all phases of the mission. The software also performs thermal control of the telescope to maintain pointing accuracy and monitors radiation levels throughout the orbit so that the Science Instruments can be safed if radiation thresholds are exceeded. The efficient operation of Chandra flight software has enabled the gathering of crucial science data. The Chandra flight software fault protection is the key to early detection and prevention of science instrument or spacecraft damage in an operating platform/environment, which is completely unforgiving. Permanently open Sun Shade Door and ACIS focal plane radiator sensitivity exposes science instruments and mirrors to damage for pointing anomalies causing an attitude excursion. The Chandra flight software must prevent these attitude excursions from occurring for ANY failure. Another example is that the power system has an unregulated bus, which imposes severe operating requirements on Chandra flight software to control array pointing and battery connection/disconnect using a unique algorithmic and logic approach. The Chandra flight software has enabled a truly autonomous vehicle with greater than 99% of all mission data collected as planned. Less than 15% of spacecraft operations are conducted in view (1 hour out of 8) leading to very extended periods without ground contact. The Chandra flight software implements the flexible mission plan during this out of view period, manages the solid state recorder capacity, controls all pointing and maneuvers, provides fault detection for all satellite subsystems, and initiates communications with the ground at the appropriate time. This paper will describe the software architecture features, key design elements and software testing techniques that have facilitated Chandra's success.

Crumbley, Tim↗

Run Time Assurance for Electric Vertical Takeoff and Landing Aircraft

NASA is conducting research to demonstrate and evaluate the application of Run Time Assurance (RTA) as a means to assure safety in Electric Vertical Takeoff and Landing (eVTOL) aircraft with highly automated or autonomous flight capability supervised by a single onboard pilot. The work described in this report demonstrates an application of RTA and examines the implications for design and analysis of aircraft functions and systems; aircraft safety hazards; safety assurance; development assurance; and pilot tasks and performance. This research effort also seeks to assess the efficacy of the combined application of traditional Functional Hazard Analysis (FHA) and the more modern System Theoretic Process Analysis (STPA) techniques to perform hazard analyses on aircraft with complex automated and autonomous systems and an onboard pilot. During the research effort we developed architectural designs of two alternate eVTOL aircraft, generally following the process characterized in the SAE standards ARP4754 and ARP4761. The design has focused on the control architectures of these aircraft, which are identical except that one incorporates RTA techniques to reduce the criticality of some key software components. Artifacts of this process include a taxonomy of aircraft-level functions, aircraft-level architecture diagrams, aircraft-level functional hazard assessments (AFHA), function allocations onto aircraft systems and subsystems, functional block diagrams for a select set of control-related functions, and system-level functional hazard assessments (SFHA) for those functions. This project has highlighted the notion that DAL D is something of a sweet spot for low-confidence controllers in an RTA-based design. Among the many activities described in DO-178C, the activities related to requirement verifiability, algorithmic accuracy, and test coverage can be the most challenging for the kinds of advanced control techniques that may be desirable in novel UAM designs, such as adaptive control, machine-learning, artificial intelligence, numerical search, and Monte Carlo based algorithms. Moreover, the standard requires that development teams demonstrate that errors leading to unacceptable failure conditions have been removed from the software. The RTA architecture, which cordons off the low-confidence function, makes it much easier to show this for these kinds of algorithms. With regard to the use of STPA and FHA as complementary hazard analysis techniques, our research effort led us to the conclusion that STPA should be used to derive requirements for hardware and software systems and/or components. Also, STPA is a natural complement to other processes in ARP4754A involving design studies and iteration.

Run-time assurance↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failires. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with the estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certifications processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E.↗

Space Shuttle Main Engine Quantitative Risk Assessment: Illustrating Modeling of a Complex System with a New QRA Software Package

During 1997, a team from Hernandez Engineering, MSFC, Rocketdyne, Thiokol, Pratt & Whitney, and USBI completed the first phase of a two year Quantitative Risk Assessment (QRA) of the Space Shuttle. The models for the Shuttle systems were entered and analyzed by a new QRA software package. This system, termed the Quantitative Risk Assessment System(QRAS), was designed by NASA and programmed by the University of Maryland. The software is a groundbreaking PC-based risk assessment package that allows the user to model complex systems in a hierarchical fashion. Features of the software include the ability to easily select quantifications of failure modes, draw Event Sequence Diagrams(ESDs) interactively, perform uncertainty and sensitivity analysis, and document the modeling. This paper illustrates both the approach used in modeling and the particular features of the software package. The software is general and can be used in a QRA of any complex engineered system. The author is the project lead for the modeling of the Space Shuttle Main Engines (SSMEs), and this paper focuses on the modeling completed for the SSMEs during 1997. In particular, the groundrules for the study, the databases used, the way in which ESDs were used to model catastrophic failure of the SSMES, the methods used to quantify the failure rates, and how QRAS was used in the modeling effort are discussed. Groundrules were necessary to limit the scope of such a complex study, especially with regard to a liquid rocket engine such as the SSME, which can be shut down after ignition either on the pad or in flight. The SSME was divided into its constituent components and subsystems. These were ranked on the basis of the possibility of being upgraded and risk of catastrophic failure. Once this was done the Shuttle program Hazard Analysis and Failure Modes and Effects Analysis (FMEA) were used to create a list of potential failure modes to be modeled. The groundrules and other criteria were used to screen out the many failure modes that did not contribute significantly to the catastrophic risk. The Hazard Analysis and FMEA for the SSME were also used to build ESDs that show the chain of events leading from the failure mode occurence to one of the following end states: catastrophic failure, engine shutdown, or siccessful operation( successful with respect to the failure mode under consideration).

Smart, Christian↗

An experimental investigation of fault tolerant software structures in an avionics application

The objective of this experimental investigation is to compare the functional performance and software reliability of competing fault tolerant software structures utilizing software diversity. In this experiment, three versions of the redundancy management software for a skewed sensor array have been developed using three diverse failure detection and isolation algorithms and incorporated into various N-version, recovery block and hybrid software structures. The empirical results show that, for maximum functional performance improvement in the selected application domain, the results of diverse algorithms should be voted before being processed by multiple versions without enforced diversity. Results also suggest that when the reliability gain with an N-version structure is modest, recovery block structures are more feasible since higher reliability can be obtained using an acceptance check with a modest reliability.

Caglayan, Alper K.↗

Hardware Aware Mitigation of Timing Side-Channel Vulnerabilities in Critical Infrastructure Software

Program runtime/timing attacks exploit variations in a program’s execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime sidechannel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware/device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. We propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate targeted software fixes. We implement DISARM to support C/C++/Java source codes and validate it across 22 standard benchmarks. DISARM outperforms state-of-the-art solutions such as PENDULUM and DifFuzzAR in terms of execution time overhead (up to −46%), code size overhead (up to −10%), and correctness (no failures) on five different embedded/edge devices.

Suha, Tasneem [University of Maine]↗

A real-time diagnostic and performance monitor for UNIX

There are now over one million UNIX sites and the pace at which new installations are added is steadily increasing. Along with this increase, comes a need to develop simple efficient, effective and adaptable ways of simultaneously collecting real-time diagnostic and performance data. This need exists because distributed systems can give rise to complex failure situations that are often un-identifiable with single-machine diagnostic software. The simultaneous collection of error and performance data is also important for research in failure prediction and error/performance studies. This paper introduces a portable method to concurrently collect real-time diagnostic and performance data on a distributed UNIX system. The combined diagnostic/performance data collection is implemented on a distributed multi-computer system using SUN4's as servers. The approach uses existing UNIX system facilities to gather system dependability information such as error and crash reports. In addition, performance data such as CPU utilization, disk usage, I/O transfer rate and network contention is also collected. In the future, the collected data will be used to identify dependability bottlenecks and to analyze the impact of failures on system performance.

Dong, Hongchao↗