Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Modelling the brittle failure of graphite induced by the controlled impact of runaway electrons in DIII-D

The thermo-mechanical response of an ATJ graphite sample to controlled runaway electron (RE) dissipation, realized in DIII-D, is modelled with a novel work-flow that features the RE orbit code KORC, the Monte Carlo particle transport code Geant4 and the finite element multiphysics software COMSOL. KORC provides the RE striking positions and momenta, Geant4 calculates the volumetric energy deposition and COMSOL simulates the thermoelastic response. Brittle failure is predicted according to the maximum normal stress criterion, which is suitable for ATJ graphite owing to its linear elastic behavior up to fracture and its isotropic mechanical properties. Measurements of the conducted energy, damage topology, explosion timing and blown-off material volume, impose a number of empirical constraints that suffice to distinguish between different RE impact scenarios and to identify RE parameters which provide the best match to the observations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Autonomous Cryogenic Load Operations: Knowledge-Based Autonomous Test Engineer

The Knowledge-Based Autonomous Test Engineer (KATE) program has a long history at KSC. Now a part of the Autonomous Cryogenic Load Operations (ACLO) mission, this software system has been sporadically developed over the past 20 years. Originally designed to provide health and status monitoring for a simple water-based fluid system, it was proven to be a capable autonomous test engineer for determining sources of failure in the system. As part of a new goal to provide this same anomaly-detection capability for a complicated cryogenic fluid system, software engineers, physicists, interns and KATE experts are working to upgrade the software capabilities and graphical user interface. Much progress was made during this effort to improve KATE. A display of the entire cryogenic system's graph, with nodes for components and edges for their connections, was added to the KATE software. A searching functionality was added to the new graph display, so that users could easily center their screen on specific components. The GUI was also modified so that it displayed information relevant to the new project goals. In addition, work began on adding new pneumatic and electronic subsystems into the KATE knowledge base, so that it could provide health and status monitoring for those systems. Finally, many fixes for bugs, memory leaks, and memory errors were implemented and the system was moved into a state in which it could be presented to stakeholders. Overall, the KATE system was improved and necessary additional features were added so that a presentation of the program and its functionality in the next few months would be a success.

Schrading, J. Nicolas↗

Lessons Learned from Seven Space Shuttle Missions

Much can be learned from well-written descriptions of the technical and organizational factors that lead to an accident. Subsequent analysis by third parties of investigation reports and associated evidence collected during the investigations can lead to additional insight. Much can also be learned from documented close calls that do not result in loss of life or a spacecraft, such as the Mars Exploration Rover Spirit software anomaly, the SOHO mission interruption, and the NEAR burn anomaly. Seven space shuttle incidents fall into the latter category: Rendezvous Target Failure On STS-41B; Rendezvous Radar Anomaly and Trajectory Dispersion-STS-32 ;Rendezvous Lambert Targeting Anomaly-STS-49; Rendezvous Lambert Targeting Anomaly-STS-51; Zero Doppler Steering Maneuver Anomaly-STS-59; Excessive Propellant Consumption During Rendezvous-STS-69; Global Positioning System Receiver and Associated Shuttle Flight Software Anomalies-STS-91 Procedural work-arounds or software changes prevented them from threatening mission success. Extensive investigations, which included the independent recreation of the anomalies by multiple Shuttle Program organizations, were the key to determining the cause, accurately assessing risk, and identifying software and software process improvements. Lessons learned from these incidents not only validated long-standing operational best practices, but serve to promote discussion and mentoring among Program personnel and are applicable to future space flight programs.

Goodman, John↗

Sensor Data Quality and Angular Rate Down-Selection Algorithms on SLS EM-1

The NASA Space Launch System Block 1 launch vehicle is equipped with an Inertial Navigation System (INS) and multiple Rate Gyro Assemblies (RGA) that are used in the Guidance, Navigation, and Control (GN&C) algorithms. The INS provides the inertial position, velocity, and attitude of the vehicle along with both angular rate and specific force measurements. Additionally, multiple sets of co-located rate gyros supply angular rate data. The collection of angular rate data, taken along the launch vehicle, is used to separate out vehicle motion from flexible body dynamics. Since the system architecture uses redundant sensors, the capability was developed to evaluate the health (or validity) of the independent measurements. A suite of Sensor Data Quality (SDQ) algorithms is responsible for assessing the angular rate data from the redundant sensors. When failures are detected, SDQ will take the appropriate action and disqualify or remove faulted sensors from forward processing. Additionally, the SDQ algorithms contain logic for down-selecting the angular rate data used by the GNC software from the set of healthy measurements. This paper explores the trades and analyses that were performed in selecting a set of robust fault-detection algorithms included in the GN&C flight software. These trades included both an assessment of hardware-provided health and status data as well as an evaluation of different algorithms based on time-to-detection, type of failures detected, and probability of detecting false positives. We then provide an overview of the algorithms used for both fault-detection and measurement down selection. We next discuss the role of trajectory design, flexible-body models, and vehicle response to off-nominal conditions in setting the detection thresholds. Lastly, we present lessons learned from software integration and hardware-in-the-loop testing.

Park, Thomas↗

Development of confidence limits by pivotal functions for estimating software reliability

The utility of pivotal functions is established for assessing software reliability. Based on the Moranda geometric de-eutrophication model of reliability growth, confidence limits for attained reliability and prediction limits for the time to the next failure are derived using a pivotal function approach. Asymptotic approximations to the confidence and prediction limits are considered and are shown to be inadequate in cases where only a few bugs are found in the software. Departures from the assumed exponentially distributed interfailure times in the model are also investigated. The effect of these departures is discussed relative to restricting the use of the Moranda model.

Dotson, Kelly J.↗

NASA Tech Briefs, April 2002

The contents include: 1) Application Briefs; 2) Sneak Preview of Sensors Expo; 3) The Complexity of the Diagnosis Problem; 4) Design Concepts for the ISS TransHab Module; 5) Characteristics of Supercritical Transitional Mixing Layers; 6) Electrometer for Triboelectric Evaluation of Materials; 7) Infrared CO2 Sensor With Built-In Calibration Chambers; 8) Solid-State Potentiometric CO Sensor; 9) Planetary Rover Absolute Heading Detection Using a Sun Sensor; 10) Concept for Utilizing Full Areas of STJ Photodetector Arrays; 11) Development of Cognitive Sensors; 12) Enabling Higher-Voltage Operation of SOl CMOS Transistors; 13) Estimating Antenna-Pointing Errors From Beam Squints; 14) Advanced-Fatigue-Crack-Growth and Fracture- Mechanics Program; 15) Software for Sequencing Spacecraft Actions; 16) Program Distributes and Tracks Organizational Memoranda; 16) Flat Membrane Device for Dehumidification of Air; 17) Inverted Hindle Mount Reduces Sag of a Large, Precise Mirror; 18) Heart-Pump-Outlet/Cannula Coupling; 19) Externally Triggered Microcapsules Release Drugs In Situ; 20) Combinatorial Drug Design Augmented by Information Theory; 21) Multiple-Path-Length Optical Absorbance Cell; 22) Model of a Fluidized Bed Containing a Mixture of Particles; 23) Refractive Secondary Concentrators for Solar Thermal Systems; 24) Cold Flow Calorimeter; 25) Methodology for Tracking Hazards and Predicting Failures; 26) Estimating Heterodyne-Interferometer Polarization Leakage; 27) An Efficient Algorithm for Propagation of Temporal- Constraint Networks; 28) Software for Continuous Replanning During Execution; 29) Surface-Launched Explorers for Reconnaissance/Scouting; 30) Firmware for a Small Motion-Control Processor; 31) Gear Bearings and Gear-Bearing Transmissions; and 32) Linear Dynamometer With Variable Stroke and Frequency.

Source record↗

Early experiences building a software quality prediction model

Early experiences building a software quality prediction model are discussed. The overall research objective is to establish a capability to project a software system's quality from an analysis of its design. The technical approach is to build multivariate models for estimating reliability and maintainability. Data from 21 Ada subsystems were analyzed to test hypotheses about various design structures leading to failure-prone or unmaintainable systems. Current design variables highlight the interconnectivity and visibility of compilation units. Other model variables provide for the effects of reusability and software changes. Reported results are preliminary because additional project data is being obtained and new hypotheses are being developed and tested. Current multivariate regression models are encouraging, explaining 60 to 80 percent of the variation in error density of the subsystems.

Agresti, W. W.↗

Markov Chains For Testing Redundant Software

Preliminary design developed for validation experiment that addresses problems unique to assuring extremely high quality of multiple-version programs in process-control software. Approach takes into account inertia of controlled system in sense it takes more than one failure of control program to cause controlled system to fail. Verification procedure consists of two steps: experimentation (numerical simulation) and computation, with Markov model for each step.

White, Allan L.↗

Tools for distributed application management

Distributed application management consists of monitoring and controlling an application as it executes in a distributed environment. It encompasses such activities as configuration, initialization, performance monitoring, resource scheduling, and failure response. The Meta system is described: a collection of tools for constructing distributed application management software. Meta provides the mechanism, while the programmer specifies the policy for application management. The policy is manifested as a control program which is a soft real time reactive program. The underlying application is instrumented with a variety of built-in and user defined sensors and actuators. These define the interface between the control program and the application. The control program also has access to a database describing the structure of the application and the characteristics of its environment. Some of the more difficult problems for application management occur when pre-existing, nondistributed programs are integrated into a distributed application for which they may not have been intended. Meta allows management functions to be retrofitted to such programs with a minimum of effort.

Marzullo, Keith↗

Contingency designs for attitude determination of TRMM

In this paper, several attitude estimation designs are developed for the Tropical Rainfall Measurement Mission (TRMM) spacecraft. A contingency attitude determination mode is required in the event of a primary sensor failure. The final design utilizes a full sixth-order Kalman filter. However, due to initial software concerns, the need to investigate simpler designs was required. The algorithms presented in this paper can be utilized in place of a full Kalman filter, and require less computational burden. These algorithms are based on filtered deterministic approaches and simplified Kalman filter approaches. Comparative performances of all designs are shown by simulating the TRMM spacecraft in mission mode. Comparisons of the simulation results indicate that comparable accuracy with respect to a full Kalman filter design is possible.

Crassidis, John L.↗

Electron Microscopy and Image Analysis for Selected Materials

This particular project was completed in collaboration with the metallurgical diagnostics facility. The objective of this research had four major components. First, we required training in the operation of the environmental scanning electron microscope (ESEM) for imaging of selected materials including biological specimens. The types of materials range from cyanobacteria and diatoms to cloth, metals, sand, composites and other materials. Second, to obtain training in surface elemental analysis technology using energy dispersive x-ray (EDX) analysis, and in the preparation of x-ray maps of these same materials. Third, to provide training for the staff of the metallurgical diagnostics and failure analysis team in the area of image processing and image analysis technology using NIH Image software. Finally, we were to assist in the sample preparation, observing, imaging, and elemental analysis for Mr. Richard Hoover, one of NASA MSFC's solar physicists and Marshall's principal scientist for the agency-wide virtual Astrobiology Institute. These materials have been collected from various places around the world including the Fox Tunnel in Alaska, Siberia, Antarctica, ice core samples from near Lake Vostoc, thermal vents in the ocean floor, hot springs and many others. We were successful in our efforts to obtain high quality, high resolution images of various materials including selected biological ones. Surface analyses (EDX) and x-ray maps were easily prepared with this technology. We also discovered and used some applications for NIH Image software in the metallurgical diagnostics facility.

Williams, George↗

Passive Superconducting Shielding: Experimental Results and Computer Models

Passive superconducting shielding for magnetic refrigerators has advantages over active shielding and passive ferromagnetic shielding in that it is lightweight and easy to construct. However, it is not as easy to model and does not fail gracefully. Failure of a passive superconducting shield may lead to persistent flwc and persistent currents. Unfortunately, modeling software for superconducting materials is not as easily available as is software for simple coils or for ferromagnetic materials. This paper will discuss ways of using available software to model passive superconducting shielding.

Warner, Brent↗

Passive Superconducting Shielding: Experimental Results and Computer Models

Passive superconducting shielding for magnetic refrigerators has advantages over active shielding and passive ferromagnetic shielding in that it is lightweight and easy to construct. However, it is not as easy to model and does not fail gracefully. Failure of a passive superconducting shield may lead to persistent flux and persistent currents. Unfortunately, modeling software for superconducting materials is not as easily available as is software for simple coils or for ferromagnetic materials. This paper will discuss ways of using available software to model passive superconducting shielding.

Warner, B. A.↗

General Purpose Data-Driven Online System Health Monitoring with Applications to Space Operations

Modern space transportation and ground support system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional parameter limit checking, or model-based or rule-based methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. System health can be monitored by comparing real-time operating data with these nominal characterizations, providing detection of anomalous data signatures indicative of system faults, failures, or precursors of significant failures. The Inductive Monitoring System (IMS) is a general purpose, data-driven system health monitoring software tool that has been successfully applied to several aerospace applications and is under evaluation for anomaly detection in vehicle and ground equipment for next generation launch systems. After an introduction to IMS application development, we discuss these NASA online monitoring applications, including the integration of IMS with complementary model-based and rule-based methods. Although the examples presented in this paper are from space operations applications, IMS is a general-purpose health-monitoring tool that is also applicable to power generation and transmission system monitoring.

Iverson, David L.↗

Fault-tolerant software

Limitations in the current capabilities for verifying programs by formal proof or by exhaustive testing have led to the investigation of fault-tolerance techniques for applications where the consequence of failure is particularly severe. Two current approaches, N-version programming and the recovery block, are described. A critical feature in the latter is the acceptance test, and a number of useful techniques for constructing these are presented. A system reliability model for the recovery block is introduced, and conclusions derived from this model that affect the design of fault-tolerant software are discussed.

Hecht, H.↗

Data diversity: An approach to software fault tolerance

Data diversity is described, and the results of a pilot study are presented. The regions of the input space that cause failure for certain experimental programs are discussed, and data reexpression, the way in which alternate input data sets can be obtained, is examined. A description is given of the retry block which is the data-diverse equivalent of the recovery block, and a model of the retry block, together with some empirical results is presented. N-copy programming which is the data-diverse equivalent of N-version programming is considered, and a simple model and some empirical results are also given.

Ammann, Paul E.↗

Analysis of faults in an N-version software experiment

The authors have conducted a large-scale experiment in N-version programming. A total of 27 versions of a program were prepared independently from the same specification at two universities. The results of executing the versions revealed that the versions were individually extremely reliable but that the number of input cases in which more than one failed was substantially more than would be expected if they were statistically independent. After the versions had been executed, the failures of each version were examined and the associated faults located. The analysis showed that in some cases the programmers made equivalent logical errors, indicating that some parts of the problem were simply more difficult than others. The authors also found cases in which apparently different logical errors yielded faults that caused statistically correlated failures, indicating that there are special cases in the input space that present difficulty in various parts of the solution. A formal model is presented to explain this phenomenon. It appears that minor differences in the software development environment would not have a major impact in reducing the incidence of faults that cause correlated failures.

Brilliant, Susan S.↗

Test and Validate Distributed Coaxial Cable Sensors for in situ Condition Monitoring of Coal-Fired Boiler Tubes

This project aims to test, validate, and advance the technology readiness level (from TRL5 to TRL7) of a novel low-cost distributed stainless-steel/ceramic coaxial cable sensing (SSC-CCS) technology for in situ monitoring of the boiler tube temperature in existing coal-fired power plants. The novel SSC-CCS sensing technology and associated condition-based monitoring (CBM) software to be demonstrated in this project will lead to an improved understanding of the boiler tube failure mechanisms and a prognostic system to improve the overall performance, reliability, and flexibility of the nation’s coal-fired power plant fleet. A boiler tube monitoring system with distributed coaxial cable temperature sensors and a sensor acquisition system was constructed. The high-temperature coaxial cable sensor with a length of 1.3m was made by using a quartz tube (1mm inner diameter (ID) and 6mm outer diameter (OD)) to concentrically separate a 304 stainless-steel (SS) rod (1mm OD) and SS tube (7.94mm OD and 6.16mm ID). The sensor acquisition system includes a vector network analyzer (VNA), a radio frequency (RF) power amplifier, multiple switches and a USB hub. The distributed stainless-steel quartz coaxial cable sensor (SSQ-CCS) had a linear response to temperature with a resolution uncertainty of σ = 0.77℃. To withstand the harsh conditions of 3,300 steam pressures and 800℃ high temperatures, the sensor was shielded by a protective tube made of the same material as the boiler tube. The protection tube had an OD of 1.5 inches and a thickness of 0.25 inches. In the laboratory tests, the sensor showed good sensitivity and fast response. The drift was bounded between +0.33% and -0.67% during a test at 600℃ for 350 hours, indicating good stability of the sensor. A field test was conducted where four sensors were welded on four superheat tubes (SH-Ts) at a coal-fired power station over 400 days. Conventional thermocouples were welded to the superheater tubes alongside the coaxial cable sensors for the purpose of comparison. Two sensors were capable of distributed sensing, with three multiplexed sensing sections. The other two sensors were single section. During the 400-day test period, the power plant experienced startups and shutdowns. At the steady state operations, the temperature of the boiler tube is about 600℃ (1112°F). The sensors recorded the entire coal-firing processes (start-up, steady state, and shut-down) and the glitch event. A GSM modem and a Watchdog were added to the system to ensure reliable data recording. The GSM modem sent daily messages to plant managers and Clemson team to inform the status of the sensor system. If the system was not normally working, the Watchdog would reboot the system automatically. The new coaxial cable based distributed sensing technology has been proven to be successful in both laboratory and field tests. A comprehensive four-stage multi-physics computational framework has been developed to assist the design, optimization, installation, and operation of SSQ-CCS. With the consideration of various operation conditions, we predict the distributions of flue gas temperatures within coal-fired boilers, the temperature correlation between the boiler tube and SSQ-CCS, and the safety of SSQ-CCS. A conditional-based monitoring system is implemented as well. The computational framework developed in this work can guide the future operation of coal-fired plants and other power plants for the safety prediction of boiler operations.

01 COAL, LIGNITE, AND PEAT↗