Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Control of Technology Transfer at JPL

Controlled Technology: 1) Design: preliminary or critical design data, schematics, technical flow charts, SNV code/diagnostics, logic flow diagrams, wirelist, ICDs, detailed specifications or requirements. 2) Development: constraints, computations, configurations, technical analyses, acceptance criteria, anomaly resolution, detailed test plans, detailed technical proposals. 3) Production: process or how-to: assemble, operated, repair, maintain, modify. 4) Manufacturing: technical instructions, specific parts, specific materials, specific qualities, specific processes, specific flow. 5) Operations: how-to operate, contingency or standard operating plans, Ops handbooks. 6) Repair: repair instructions, troubleshooting schemes, detailed schematics. 7) Test: specific procedures, data, analysis, detailed test plan and retest plans, detailed anomaly resolutions, detailed failure causes and corrective actions, troubleshooting, trended test data, flight readiness data. 8) Maintenance: maintenance schedules and plans, methods for regular upkeep, overhaul instructions. 9) Modification: modification instructions, upgrades kit parts, including software

Jet Propulsion Laboratory (JPL)↗

Rotorcraft Diagnostics

Health management (HM) in any engineering systems requires adequate understanding about the system s functioning; a sufficient amount of monitored data; the capability to extract, analyze, and collate information; and the capability to combine understanding and information for HM-related estimation and decision-making. Rotorcraft systems are, in general, highly complex. Obtaining adequate understanding about functioning of such systems is quite difficult, because of the proprietary (restricted access) nature of their designs and dynamic models. Development of an EIM (exact inverse map) solution for rotorcraft requires a process that can overcome the abovementioned difficulties and maximally utilize monitored information for HM facilitation via employing advanced analytic techniques. The goal was to develop a versatile HM solution for rotorcraft for facilitation of the Condition Based Maintenance Plus (CBM+) capabilities. The effort was geared towards developing analytic and reasoning techniques, and proving the ability to embed the required capabilities on a rotorcraft platform, paving the way for implementing the solution on an aircraft-level system for consolidation and reporting. The solution for rotorcraft can he used offboard or embedded directly onto a rotorcraft system. The envisioned solution utilizes available monitored and archived data for real-time fault detection and identification, failure precursor identification, and offline fault detection and diagnostics, health condition forecasting, optimal guided troubleshooting, and maintenance decision support. A variant of the onboard version is a self-contained hardware and software (HW+SW) package that can be embedded on rotorcraft systems. The HM solution comprises components that gather/ingest data and information, perform information/feature extraction, analyze information in conjunction with the dependency/diagnostic model of the target system, facilitate optimal guided troubleshooting, and offer decision support for optimal maintenance.

Haste, Deepak↗

Methods for Assessing Honeycomb Sandwich Panel Wrinkling Failures

Efficient closed-form methods for predicting the facesheet wrinkling failure mode in sandwich panels are assessed. Comparisons were made with finite element model predictions for facesheet wrinkling, and a validated closed-form method was implemented in the HyperSizer structure sizing software.

Zalewski, Bart F.↗

Spinoff 2013

Topics covered include: Innovative Software Tools Measure Behavioral Alertness; Miniaturized, Portable Sensors Monitor Metabolic Health; Patient Simulators Train Emergency Caregivers; Solar Refrigerators Store Life-Saving Vaccines; Monitors Enable Medication Management in Patients' Homes; Handheld Diagnostic Device Delivers Quick Medical Readings; Experiments Result in Safer, Spin-Resistant Aircraft; Interfaces Visualize Data for Airline Safety, Efficiency; Data Mining Tools Make Flights Safer, More Efficient; NASA Standards Inform Comfortable Car Seats; Heat Shield Paves the Way for Commercial Space; Air Systems Provide Life Support to Miners; Coatings Preserve Metal, Stone, Tile, and Concrete; Robots Spur Software That Lends a Hand; Cloud-Based Data Sharing Connects Emergency Managers; Catalytic Converters Maintain Air Quality in Mines; NASA-Enhanced Water Bottles Filter Water on the Go; Brainwave Monitoring Software Improves Distracted Minds; Thermal Materials Protect Priceless, Personal Keepsakes; Home Air Purifiers Eradicate Harmful Pathogens; Thermal Materials Drive Professional Apparel Line; Radiant Barriers Save Energy in Buildings; Open Source Initiative Powers Real-Time Data Streams; Shuttle Engine Designs Revolutionize Solar Power; Procedure-Authoring Tool Improves Safety on Oil Rigs; Satellite Data Aid Monitoring of Nation's Forests; Mars Technologies Spawn Durable Wind Turbines; Programs Visualize Earth and Space for Interactive Education; Processor Units Reduce Satellite Construction Costs; Software Accelerates Computing Time for Complex Math; Simulation Tools Prevent Signal Interference on Spacecraft; Software Simplifies the Sharing of Numerical Models; Virtual Machine Language Controls Remote Devices; Micro-Accelerometers Monitor Equipment Health; Reactors Save Energy, Costs for Hydrogen Production; Cameras Monitor Spacecraft Integrity to Prevent Failures; Testing Devices Garner Data on Insulation Performance; Smart Sensors Gather Information for Machine Diagnostics; Oxygen Sensors Monitor Bioreactors and Ensure Health and Safety; Vision Algorithms Catch Defects in Screen Displays; and Deformable Mirrors Capture Exoplanet Data, Reflect Lasers.

Source record↗

Exploring Cognition Using Software Defined Radios for NASA Missions

NASA missions typically operate using a communication infrastructure that requires significant schedule planning with limited flexibility when the needs of the mission change. Parameters such as modulation, coding scheme, frequency, and data rate are fixed for the life of the mission. This is due to antiquated hardware and software for both the space and ground assets and a very complex set of mission profiles. Automated techniques in place by commercial telecommunication companies are being explored by NASA to determine their usability by NASA to reduce cost and increase science return. Adding cognition the ability to learn from past decisions and adjust behavior is also being investigated. Software Defined Radios are an ideal way to implement cognitive concepts. Cognition can be considered in many different aspects of the communication system. Radio functions, such as frequency, modulation, data rate, coding and filters can be adjusted based on measurements of signal degradation. Data delivery mechanisms and route changes based on past successes and failures can be made to more efficiently deliver the data to the end user. Automated antenna pointing can be added to improve gain, coverage, or adjust the target. Scheduling improvements and automation to reduce the dependence on humans provide more flexible capabilities. The Cognitive Communications project, funded by the Space Communication and Navigation Program, is exploring these concepts and using the SCaN Testbed on board the International Space Station to implement them as they evolve. The SCaN Testbed contains three Software Defined Radios and a flight computer. These four computing platforms, along with a tracking antenna system and the supporting ground infrastructure, will be used to implement various concepts in a system similar to those used by missions. Multiple universities and SBIR companies are supporting this investigation. This paper will describe the cognitive system ideas under consideration and the plan for implementing them on platforms, including the SCaN Testbed. Discussions in the paper will include how these concepts might be used to reduce cost and improve the science return for NASA missions.

transmitters receivers↗

Command and Control Software Development Memory Management

This internship was initially meant to cover the implementation of unit test automation for a NASA ground control project. As is often the case with large development projects, the scope and breadth of the internship changed. Instead, the internship focused on finding and correcting memory leaks and errors as reported by a COTS software product meant to track such issues. Memory leaks come in many different flavors and some of them are more benign than others. On the extreme end a program might be dynamically allocating memory and not correctly deallocating it when it is no longer in use. This is called a direct memory leak and in the worst case can use all the available memory and crash the program. If the leaks are small they may simply slow the program down which, in a safety critical system (a system for which a failure or design error can cause a risk to human life), is still unacceptable. The ground control system is managed in smaller sub-teams, referred to as CSCIs. The CSCI that this internship focused on is responsible for monitoring the health and status of the system. This team's software had several methods/modules that were leaking significant amounts of memory. Since most of the code in this system is safety-critical, correcting memory leaks is a necessity.

Programmin↗

Lessons Learned from Astrobee Operations on the International Space Station

Since its launch in 2019, NASA has been operating three Astrobee free-flying robots providing an autonomous and adaptable research platform aboard the International Space Station (ISS). These robots have not only facilitated a myriad of national and international research endeavors in microgravity but have also served as a STEM outreach platform for student competitions aboard the ISS. Amidst its extensive operational tenure, spanning over five years and exceeding 1200 hours of cumulative free-flyer operation as of April 2024, the Astrobee robots have encountered software and hardware anomalies. Despite its inherent design for on-orbit repair or replacement, certain anomalies have proven to be complex, necessitating remote resolution via software and firmware updates or, in extreme cases, hardware replacements or the return of faulty units to NASA's ground facilities for repair. Such challenges underscore the delicate balance between the autonomous functionality of Astrobee and the occasional need for human intervention to maintain optimal performance. One recurring point of failure identified during Astrobee's operational lifespan has been the SD card, a critical component utilized by the different Astrobee processors and the Dock Station. The occurrence of SD card anomalies, both on orbit and within ground units, has provided invaluable insights into the improvement of Astrobee's systems and mitigation to future faults. This presentation will focus on four key areas: 1. Overview of Faults and Anomalies: A comprehensive examination of the diverse array of faults and anomalies encountered by Astrobee and its associated systems both in orbit and on the ground. From software glitches to hardware malfunctions, this section provides insights into the challenges faced during Astrobee's operational tenure. 2. Resolution Processes and Procedures: An in-depth discussion of the methodologies and procedures implemented to resolve the encountered anomalies. This includes remote troubleshooting, software patches, firmware updates, and, when necessary, the logistics involved in hardware replacements or down-massing for repair. 3. Implementation of Software Updates and Hardware Upgrades: A detailed exploration of the strategies employed to mitigate the risk of recurring anomalies through the implementation of software updates and hardware upgrades. This section highlights the iterative nature of Astrobee's development, emphasizing the continuous pursuit of robustness and reliability. 4. Lessons Learned and Future Directions: Reflecting on the insights gained from addressing anomalies, this section examines the lessons learned and outlines future directions for enhancing Astrobee's robustness and resilience. It underscores the iterative nature of space exploration and the importance of adaptability and continuous improvement in the pursuit of scientific discovery. Through a nuanced examination of Astrobee's operational challenges and the strategies employed to overcome them, this presentation sheds light on the complexities of operating autonomous robotic systems in the ISS environment. It underscores NASA's commitment to pushing the boundaries of exploration and innovation while navigating the inherent challenges of space exploration.

Astrobee↗

Lessons Learned from Astrobee Operations on the International Space Station

Since its launch in 2019, NASA has been operating three Astrobee free-flying robots providing an autonomous and adaptable research platform aboard the International Space Station (ISS). These robots have not only facilitated a myriad of national and international research endeavors in microgravity but have also served as a STEM outreach platform for student competitions aboard the ISS. Amidst its extensive operational tenure, spanning over five years and exceeding 1200 hours of cumulative free-flyer operation as of April 2024, the Astrobee robots have encountered software and hardware anomalies. Despite its inherent design for on-orbit repair or replacement, certain anomalies have proven to be complex, necessitating remote resolution via software and firmware updates or, in extreme cases, hardware replacements or the return of faulty units to NASA's ground facilities for repair. Such challenges underscore the delicate balance between the autonomous functionality of Astrobee and the occasional need for human intervention to maintain optimal performance. One recurring point of failure identified during Astrobee's operational lifespan has been the SD card, a critical component utilized by the different Astrobee processors and the Dock Station. The occurrence of SD card anomalies, both on orbit and within ground units, has provided invaluable insights into the improvement of Astrobee's systems and mitigation to future faults. This presentation will focus on four key areas: 1. Overview of Faults and Anomalies: A comprehensive examination of the diverse array of faults and anomalies encountered by Astrobee and its associated systems both in orbit and on the ground. From software glitches to hardware malfunctions, this section provides insights into the challenges faced during Astrobee's operational tenure. 2. Resolution Processes and Procedures: An in-depth discussion of the methodologies and procedures implemented to resolve the encountered anomalies. This includes remote troubleshooting, software patches, firmware updates, and, when necessary, the logistics involved in hardware replacements or down-massing for repair. 3. Implementation of Software Updates and Hardware Upgrades: A detailed exploration of the strategies employed to mitigate the risk of recurring anomalies through the implementation of software updates and hardware upgrades. This section highlights the iterative nature of Astrobee's development, emphasizing the continuous pursuit of robustness and reliability. 4. Lessons Learned and Future Directions: Reflecting on the insights gained from addressing anomalies, this section examines the lessons learned and outlines future directions for enhancing Astrobee's robustness and resilience. It underscores the iterative nature of space exploration and the importance of adaptability and continuous improvement in the pursuit of scientific discovery. Through a nuanced examination of Astrobee's operational challenges and the strategies employed to overcome them, this presentation sheds light on the complexities of operating autonomous robotic systems in the ISS environment. It underscores NASA's commitment to pushing the boundaries of exploration and innovation while navigating the inherent challenges of space exploration.

Astrobee↗

A Standardized Analysis Process Using Digital Image Correlation to Calculate In Situ Cladding Strain from Modified Burst Tests for Fuel Performance Code Validation

Historical data collection on nuclear fuel cladding materials has focused on generating a statistically significant amount of data to assess the material and its failure behavior. Furthermore, data generated to support material model and failure criteria development were previously posttest evaluations, so a large number of tests was required to gain new understanding. A way to expedite this process is to develop techniques capable of generating large, high-fidelity data sets from a single test with lower uncertainty or quantified uncertainty. One such example of this approach is Oak Ridge National Laboratory’s use of modified burst tests (MBTs) to analyze the mechanical behavior and failure conditions of cladding during a simulated reactivity-initiated accident (RIA). Each test incorporates digital image correlation (DIC) analysis techniques that are used to assess the accumulated strain in situ, as well as eventual cladding failure. This work has been fruitful in defining strain-to-failure conditions for materials like silicon carbide (SiC) fiber–reinforced/SiC matrix composite tubes (SiC/SiC), iron-chromium-aluminum (FeCrAl) alloy tubes, and chromium-coated Zircaloy-4 tubes. However, there are numerous DIC software available, including open-source and proprietary software. The different DIC software use various algorithms to process images and calculate displacement values. Using these different software and algorithms can lead to varying results, and perhaps larger-than-expected uncertainties. In the present study, previously published MBT data encompassing a variety of test conditions were reanalyzed with two different DIC software to assess the variance in the calculated strain results. The data consisted of SiC/SiC, FeCrAl, and chromium-coated Zircaloy-4 tubes. Plots of the calculated strains during the transient revealed good agreement between the two DIC software. The average root-mean-square errors between the two software was 0.20% strain, which is slightly larger than a previously reported error value for these tests. In conclusion, this variance in results is low enough that this analysis method can be used for code validation.

Reactivity-initiated accident↗

Real-time computer simulation/emulation for verification of multi-fault-tolerant control of Centaur-in-Shuttle

NASA has contracted with General Dynamics to design and develop an advanced Centaur liquid upper stage for support of the Galileo and Solar Polar interplanetary missions in 1985-86. The control of the Centaur while it resides in the Shuttle cargo bay must meet the STS safety requirements to be dual failure tolerant in all mission critical functions. The demonstration of the integrity of this control system in the event of multiple component failures and worst-case time-phase asynchroniety among the system's computers is performed by a real-time computer simulation. The simulation emulates the control hardware, subsystem interfaces, and imbedded software processes, wire-by-wire, to provide accessibility for fault insertion. Observability is provided via graphics and diagnostic software. Verification is the product of Monte Carlo simulation analysis.

Szatkowski, G. P.↗

MAC/GMC 4.0 User's Manual: Keywords Manual

This document is the second volume in the three volume set of User's Manuals for the Micromechanics Analysis Code with Generalized Method of Cells Version 4.0 (MAC/GMC 4.0). Volume 1 is the Theory Manual, this document is the Keywords Manual, and Volume 3 is the Example Problem Manual. MAC/GMC 4.0 is a composite material and laminate analysis software program developed at the NASA Glenn Research Center. It is based on the generalized method of cells (GMC) micromechanics theory, which provides access to the local stress and strain fields in the composite material. This access grants GMC the ability to accommodate arbitrary local models for inelastic material behavior and various types of damage and failure analysis. MAC/GMC 4.0 has been built around GMC to provide the theory with a user-friendly framework, along with a library of local inelastic, damage, and failure models. Further, applications of simulated thermo-mechanical loading, generation of output results, and selection of architectures to represent the composite material have been automated in MAC/GMC 4.0. Finally, classical lamination theory has been implemented within MAC/GMC 4.0 wherein GMC is used to model the composite material response of each ply. Consequently, the full range of GMC composite material capabilities is available for analysis of arbitrary laminate configurations as well. This volume describes the basic information required to use the MAC/GMC 4.0 software, including a 'Getting Started' section, and an in-depth description of each of the 22 keywords used in the input file to control the execution of the code.

Bednarcyk, Brett A.↗

Software control for large scale on-board checkout: A concept

Two level system checkout in which first level satisfies continuous monitoring requirements and second level provides fault isolation to satisfy maintenance requirements, provides self-checking capability for monitoring system and enables recovery from unexpected error or failure interruptions. System must perform operational duties of navigation, control, and experimentation.

Grounds, H. K.↗

Distributed Spacecraft Autonomy - Development of Swarm Autonomy Capability and Scalability for Spacecraft

The Distributed Spacecraft Autonomy project is developing a suite of software tools that enable an operator to command and receive data from a swarm as a single entity, enable a swarm to autonomously coordinate its actions via distributed decision making and reactive closed-loop control, and model swarm behavior in the presence of anomalies or failures. Our use case is the mapping of the electron density of the ionosphere using radio tomography by coordinating the selection of appropriate GPS channels, and by recording Total Electron Count (TEC) measurements. DSA will be demonstrated onboard the NASA Ames Starling mission – a swarm of four small, LEO spacecraft, scheduled to launch in 2021. We will also perform a ground demonstration with simulated and hardware-in-the-loop elements, to validate the tools for controlling swarms of up to 100 assets. The capability to communicate autonomously between the swarm satellites is demonstrated via a sophisticated simulation architecture. Historical Plasmasphere TEC data obtained via dual-band Novatel GPS Receivers are utilized as a representative input dataset for the swarm. The representative TEC data and GPS satellite observability information is fed to the autonomous software package in place of a true real-time ground data collection process. The swarm satellites actively share status updates amongst one another and utilize multi-agent decision making to optimally identify regions of interest in the TEC distribution. The software, aware of the bandwidth limitations of the swarm satellites, prioritizes explorative measurements, which define the range of observability for the satellites, as well as exploitative measurements, which focus on maximizing the observance potential of regions with prolonged, elevated TEC density. The science of this study can ultimately be used to determine the dynamics and coupling of Earth’s magnetosphere, ionosphere, and atmosphere and their response to solar and terrestrial inputs. The findings can be applied to the imaging of critical, transient phenomena in the magnetosphere in later missions. Meanwhile, the swarm autonomy capabilities have far reaching potential in future satellite missions. As an experimental demonstration of the autonomous capabilities of the network, a message is first printed within a core Flight Executive (cFE) application. Two cFE applications that communicate with one another within the same core Flight System (cFS) are shown. Communication between mission applications on the internal cFE bus is extended to utilize Data Distribution Service (DDS) for vehicle-to-vehicle networking. The DDS middleware provides reliable delivery, routing, and topic subscription features over User Datagram Protocol (UDP). Leveraging Linux containerization, a networked set of satellite instances are generated by script to simulate swarm behavior. Swarm commanding and synchronization through the network is demonstrated under various topologies and data-loss conditions. Finally, autonomous swarm scalability from 2 satellites to 100 satellites is shown.

Distributed Autonomy↗

Distributed Spacecraft Autonomy (DSA): Development of Swarm Autonomy Capability and Scalability for Spacecraft

The Distributed Spacecraft Autonomy project is developing a suite of software tools that enable an operator to command and receive data from a swarm as a single entity, enable a swarm to autonomously coordinate its actions via distributed decision making and reactive closed-loop control, and model swarm behavior in the presence of anomalies or failures. Our use case is the mapping of the electron density of the ionosphere using radio tomography by coordinating the selection of appropriate GPS channels, and by recording Total Electron Count (TEC)measurements. DSA will be demonstrated on board the NASA Ames Starling mission a swarm of four small, LEO spacecraft, scheduled to launch in 2021. We will also perform a ground demonstration with simulated and hardware-in-the-loop elements, to validate the tools for controlling swarms of up to 100 assets.The capability to communicate autonomously between the swarm satellites is demonstrated via a sophisticated simulation architecture. Historical Plasma sphere TEC data obtained via dual-band Novatel GPS Receivers are utilized as a representative input data set for the swarm. The representative TEC data and GPS satellite observability information is fed to the autonomous software package in place of a true real-time ground data collection process. The swarm satellites actively share status updates amongst one another and utilize multi-agent decision making to optimally identify regions of interest in the TEC distribution. The software,aware of the bandwidth limitations of the swarm satellites, prioritizes explorative measurements,which define the range of observability for the satellites, as well as exploitative measurements,which focus on maximizing the observance potential of regions with prolonged, elevated TEC density. The science of this study can ultimately be used to determine the dynamics and coupling of Earth's magnetosphere, ionosphere, and atmosphere and their response to solar and terrestrial inputs. The findings can be applied to the imaging of critical, transient phenomena in the magnetosphere in later missions. Meanwhile, the swarm autonomy capabilities have far reaching potential in future satellite missions.As an experimental demonstration of the autonomous capabilities of the network, a message is first printed within a core Flight Executive (cFE) application. Two cFE applications that communicate with one another within the same core Flight System (cFS) are shown.Communication between mission applications on the internal cFE bus is extended to utilize Data Distribution Service (DDS) for vehicle-to-vehicle networking. The DDS middle ware provides reliable delivery, routing, and topic subscription features over User Data gram Protocol (UDP).Leveraging Linux containerization, a networked set of satellite instances are generated by script to simulate swarm behavior. Swarm commanding and synchronization through the network is demonstrated under various topologies and data-loss conditions. Finally, autonomous swarms calability from 2 satellites to 100 satellites is shown.

Fugate, Jason↗

ISS Regenerative Life Support: Challenges and Success in the Quest for Long-Term Habitability in Space

This presentation will discuss the International Space Station s (ISS) Regenerative Environmental Control and Life Support System (ECLSS) operations with discussion of the on-orbit lessons learned, specifically regarding the challenges that have been faced as the system has expanded with a growing ISS crew. Over the 10 year history of the ISS, there have been numerous challenges, failures, and triumphs in the quest to keep the crew alive and comfortable. Successful operation of the ECLSS not only requires maintenance of the hardware, but also management of the station resources in case of hardware failure or missed re-supply. This involves effective communication between the primary International Partners (NASA and Roskosmos) and the secondary partners (JAXA and ESA) in order to keep a reserve of the contingency consumables and allow for re-supply of failed hardware. The ISS ECLSS utilizes consumables storage for contingency usage as well as longer-term regenerative systems, which allow for conservation of the expensive resources brought up by re-supply vehicles. This long-term hardware, and the interactions with software, was a challenge for Systems Engineers when they were designed and require multiple operational workarounds in order to function continuously. On a day-to-day basis, the ECLSS provides big challenges to the on console controllers. Main challenges involve the utilization of the resources that have been brought up by the visiting vehicles prior to undocking, balance of contributions between the International Partners for both systems and resources, and maintaining balance between the many interdependent systems, which includes providing the resources they need when they need it. The current biggest challenge for ECLSS is the Regenerative ECLSS system, which continuously recycles urine and condensate water into drinking water and oxygen. These systems were brought to full functionality on STS-126 (ULF-2) mission. Through system failures and recovery, the ECLSS console has learned how to balance the water within the systems, store and use water for contingencies, and continue to work with the International Partners for short-term failures. Through these challenges and the system failures, the most important lesson learned has been the importance of redundancy and operational workarounds. It is only because of the flexibility of the hardware and the software that flight controllers have the opportunity to continue operating the system as a whole for mission success.

Bazley, Jesse A.↗

A preliminary design for flight testing the FINDS algorithm

This report presents a preliminary design for flight testing the FINDS (Fault Inferring Nonlinear Detection System) algorithm on a target flight computer. The FINDS software was ported onto the target flight computer by reducing the code size by 65%. Several modifications were made to the computational algorithms resulting in a near real-time execution speed. Finally, a new failure detection strategy was developed resulting in a significant improvement in the detection time performance. In particular, low level MLS, IMU and IAS sensor failures are detected instantaneously with the new detection strategy, while accelerometer and the rate gyro failures are detected within the minimum time allowed by the information generated in the sensor residuals based on the point mass equations of motion. All of the results have been demonstrated by using five minutes of sensor flight data for the NASA ATOPS B-737 aircraft in a Microwave Landing System (MLS) environment.

Caglayan, A. K.↗

NASA ground terminal communication equipment automated fault isolation expert systems

The prototype expert systems are described that diagnose the Distribution and Switching System I and II (DSS1 and DSS2), Statistical Multiplexers (SM), and Multiplexer and Demultiplexer systems (MDM) at the NASA Ground Terminal (NGT). A system level fault isolation expert system monitors the activities of a selected data stream, verifies that the fault exists in the NGT and identifies the faulty equipment. Equipment level fault isolation expert systems are invoked to isolate the fault to a Line Replaceable Unit (LRU) level. Input and sometimes output data stream activities for the equipment are available. The system level fault isolation expert system compares the equipment input and output status for a data stream and performs loopback tests (if necessary) to isolate the faulty equipment. The equipment level fault isolation system utilizes the process of elimination and/or the maintenance personnel's fault isolation experience stored in its knowledge base. The DSS1, DSS2 and SM fault isolation systems, using the knowledge of the current equipment configuration and the equipment circuitry issues a set of test connections according to the predefined rules. The faulty component or board can be identified by the expert system by analyzing the test results. The MDM fault isolation system correlates the failure symptoms with the faulty component based on maintenance personnel experience. The faulty component can be determined by knowing the failure symptoms. The DSS1, DSS2, SM, and MDM equipment simulators are implemented in PASCAL. The DSS1 fault isolation expert system was converted to C language from VP-Expert and integrated into the NGT automation software for offline switch diagnoses. Potentially, the NGT fault isolation algorithms can be used for the DSS1, SM, amd MDM located at Goddard Space Flight Center (GSFC).

Tang, Y. K.↗

Multiversion software reliability through fault-avoidance and fault-tolerance

In this project we have proposed to investigate a number of experimental and theoretical issues associated with the practical use of multi-version software in providing dependable software through fault-avoidance and fault-elimination, as well as run-time tolerance of software faults. In the period reported here we have working on the following: We have continued collection of data on the relationships between software faults and reliability, and the coverage provided by the testing process as measured by different metrics (including data flow metrics). We continued work on software reliability estimation methods based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. We have continued studying back-to-back testing as an efficient mechanism for removal of uncorrelated faults, and common-cause faults of variable span. We have also been studying back-to-back testing as a tool for improvement of the software change process, including regression testing. We continued investigating existing, and worked on formulation of new fault-tolerance models. In particular, we have partly finished evaluation of Consensus Voting in the presence of correlated failures, and are in the process of finishing evaluation of Consensus Recovery Block (CRB) under failure correlation. We find both approaches far superior to commonly employed fixed agreement number voting (usually majority voting). We have also finished a cost analysis of the CRB approach.

Vouk, Mladen A.↗