Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Evaluation of a fault tolerant system for an integrated avionics sensor configuration with TSRV flight data

The performance analysis results of a fault inferring nonlinear detection system (FINDS) using sensor flight data for the NASA ATOPS B-737 aircraft in a Microwave Landing System (MLS) environment is presented. First, a statistical analysis of the flight recorded sensor data was made in order to determine the characteristics of sensor inaccuracies. Next, modifications were made to the detection and decision functions in the FINDS algorithm in order to improve false alarm and failure detection performance under real modelling errors present in the flight data. Finally, the failure detection and false alarm performance of the FINDS algorithm were analyzed by injecting bias failures into fourteen sensor outputs over six repetitive runs of the five minute flight data. In general, the detection speed, failure level estimation, and false alarm performance showed a marked improvement over the previously reported simulation runs. In agreement with earlier results, detection speed was faster for filter measurement sensors soon as MLS than for filter input sensors such as flight control accelerometers.

Caglayan, A. K.↗

Constant-Overhead Fault-Tolerant Bell-Pair Distillation Using High-Rate Codes

We present a fault-tolerant Bell-pair distillation scheme achieving constant overhead through high-rate quantum low-density parity-check (qLDPC) codes. Our approach maintains a constant distillation rate equal to the code rate while requiring no additional overhead beyond the physical qubits of the code. Full circuit-level analysis demonstrates fault-tolerance for input Bell-pair infidelities below a threshold ∼10%, readily achievable with near-term capabilities. Unlike previous proposals, our scheme keeps the output Bell pairs encoded in qLDPC codes at each node, eliminating unencoding overhead and enabling direct use in distributed quantum applications through recent advances in qLDPC computation. These results establish qLDPC-based distillation as a practical route toward resource-efficient quantum networks and distributed quantum computing.

quantum communication, protocols & technology↗

Test plan. GCPS task 7, subtask 7.1: IHM development

The overall objective of Task 7 is to identify cost-effective life cycle integrated health management (IHM) approaches for a reusable launch vehicle's primary structure. Acceptable IHM approaches must: eliminate and accommodate faults through robust designs, identify optimum inspection/maintenance periods, automate ground and on-board test and check-out, and accommodate and detect structural faults by providing wide and localized area sensor and test coverage as required. These requirements are elements of our targeted primary structure low cost operations approach using airline-like maintenance by exception philosophies. This development plan will follow an evolutionary path paving the way to the ultimate development of flight-quality production, operations, and vehicle systems. This effort will be focused on maturing the recommended sensor technologies required for localized and wide area health monitoring to a technology readiness level (TRL) of 6 and to establish flight ready system design requirements. The following is a brief list of IHM program objectives: design out faults by analyzing material properties, structural geometry, and load and environment variables and identify failure modes and damage tolerance requirements; design in system robustness while meeting performance objectives (weight limitations) of the reusable launch vehicle primary structure; establish structural integrity margins to preclude the need for test and checkout and predict optimum inspection/maintenance periods through life prediction analysis; identify optimum fault protection system concept definitions combining system robustness and integrity margins established above with cost effective health monitoring technologies; and use coupons, panels, and integrated full scale primary structure test articles to identify, evaluate, and characterize the preferred NDE/NDI/IHM sensor technologies that will be a part of the fault protection system.

Greenberg, H. S.↗

Managing Space System Faults: Coalescing NASA's Views

Managing faults and their resultant failures is a fundamental and critical part of developing and operating aerospace systems. Yet, recent studies have shown that the engineering "discipline" required to manage faults is not widely recognized nor evenly practiced within the NASA community. Attempts to simply name this discipline in recent years has been fraught with controversy among members of the Integrated Systems Health Management (ISHM), Fault Management (FM), Fault Protection (FP), Hazard Analysis (HA), and Aborts communities. Approaches to managing space system faults typically are unique to each organization, with little commonality in the architectures, processes and practices across the industry.

Muirhead, Brian↗

An Introduction to Markov Modeling: Concepts and Uses

Kharkov modeling is a modeling technique that is widely useful for dependability analysis of complex fault tolerant systems. It is very flexible in the type of systems and system behavior it can model. It is not, however, the most appropriate modeling technique for every modeling situation. The first task in obtaining a reliability or availability estimate for a system is selecting which modeling technique is most appropriate to the situation at hand. A person performing a dependability analysis must confront the question: is Kharkov modeling most appropriate to the system under consideration, or should another technique be used instead? The need to answer this gives rise to other more basic questions regarding Kharkov modeling: what are the capabilities and limitations of Kharkov modeling as a modeling technique? How does it relate to other modeling techniques? What kind of system behavior can it model? What kinds of software tools are available for performing dependability analyses with Kharkov modeling techniques? These questions and others will be addressed in this tutorial.

Boyd, Mark A.↗

Optimizing Hydronic Heating for Comfort and Performance in Multifamily Housing

Inefficient control settings in multifamily boilers often lead to substantial energy and cost penalties. To address this, a Fault Detection and Diagnostic (FDD) tool was developed to automate data analysis and identify operational faults such as suboptimal outdoor temperature sensor placement, misconfigured outdoor air reset (OAR) curves, excess boiler cycling, and domestic hot water (DHW) setpoint errors. By comparing pre- and post-implementation periods and applying engineering models, the tool quantifies energy savings and reduces manual analysis time by over 90%. Testing on over 100 monitored sites and a targeted subset of 12 buildings showed an average 11% energy savings from remote optimization; further validation across 19 OAR curve changes confirmed the tool’s accuracy, predicting actual savings within ±5% for most cases. Simple payback can be under three years for many multifamily buildings, though rising hardware, labor, and fuel costs create uncertainties, and decarbonization goals increasingly shift focus to electrification. The FDD tool remains invaluable for optimizing existing boilers, enhancing future electrification measures, and adapting to new technologies by refining building load estimates. In doing so, it supports both near-term efficiency and long-term transitions to low-carbon alternatives, ensuring buildings achieve substantial cost and energy benefits throughout their system lifecycles.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Battery test expert systems

The characteristics of NIHBES (nickel-hydrogen battery expert system) are described, with attention also given to NICBES-2 (nickel-cadmium battery expert system-2). The nickel-hydrogen battery testbed is set up almost identically to the nickel-cadmium battery testbed, with the exceptions of no battery protection and reconditioning circuits (BPRCs) and the frequency of transmission of data. The Ni-H2 testbed has no BPRCs and the data are transmitted every 30 s instead of every minute. An expert system shell was chosen to develop this particular expert system. The GoldWorks expert system shell from Gold Hill Computers was chosen for the task. NIHBES will extract the desired data and return fault diagnosis, status and advice, and decision support. Expert systems have been proven to be viable tools in the control and monitoring of space power systems. Presently, the DDAS (digital data acquisition system) monitors and controls the orbit time, and is responsible for limit checking, data acquisition, and data summaries. It is concluded that in the future control of the Hubble Space Telescope breadboard will be passed to NIHBES. NIHBES will be more beneficial to the testbed than the DDAS alone due to the limitations of the DDAS. The DDAS cannot provide long-term trend analysis, plotting capability, fault diagnosis, or advice.

Johnson, Yvette B.↗

De-risking fault leakage risk and containment integrity for subsurface storage applications

The subsurface is pivotal in the energy transition, for the sequestration of CO 2 and energy storage. It is crucial to understand to what extent geological faults may form leakage pathways that threaten the containment integrity of these projects. Fault flow behavior has been studied in the context of hydrocarbon development, supported by observations from wells drilled through faults, but such observations are rare in geoenergy projects. Focusing on mechanical behavior as early indicator of potential leakage risks, a probabilistic Coulomb Failure Stress workflow is developed and demonstrated using data from the Decatur CO 2 sequestration project to rank faults based on their containment risk. The analysis emphasizes the importance of fault throw relative to reservoir thickness and pore pressure change in assessing reactivation risks. Integrating this mechanical assessment with geological and dynamic fault analyses contributes to derisking fault containment for geoenergy applications, providing valuable insights for the successful development of subsurface storage projects.

58 GEOSCIENCES↗

Network Connectivity for Permanent, Transient, Independent, and Correlated Faults

This paper develops a method for the quantitative analysis of network connectivity in the presence of both permanent and transient faults. Even though transient noise is considered a common occurrence in networks, a survey of the literature reveals an emphasis on permanent faults. Transient faults introduce a time element into the analysis of network reliability. With permanent faults it is sufficient to consider the faults that have accumulated by the end of the operating period. With transient faults the arrival and recovery time must be included. The number and location of faults in the system is a dynamic variable. Transient faults also introduce system recovery into the analysis. The goal is the quantitative assessment of network connectivity in the presence of both permanent and transient faults. The approach is to construct a global model that includes all classes of faults: permanent, transient, independent, and correlated. A theorem is derived about this model that give distributions for (1) the number of fault occurrences, (2) the type of fault occurrence, (3) the time of the fault occurrences, and (4) the location of the fault occurrence. These results are applied to compare and contrast the connectivity of different network architectures in the presence of permanent, transient, independent, and correlated faults. The examples below use a Monte Carlo simulation, but the theorem mentioned above could be used to guide fault-injections in a laboratory.

White, Allan L.↗

A Case for Developing a Ground Based Replication of the Earth, Moon and Mars Spaceflight Infrastructure

When the systems are developed and in place to provide the services needed to operate en route and on the Lunar and Martian surfaces, an Earth based replication will need to be in place for the safety and protection of mission success. The replication will entail all aspects of the flight configuration end to end but will not include any closed loop systems. This would replicate the infrastructure from Lunar and Martian robots, manned surface excursions, through man and unmanned terrestrial bases, through the various types of communication systems and technologies, manned and un-manned space vehicles (large and small), to Earth based systems and control centers. An Earth based replicated infrastructure will enable checkout and test of new technologies, hardware, software updates and upgrades and procedures without putting humans and missions at risk. Analysis of events, what ifs and trouble resolution could be played out on the ground to remove as much risk as possible from any type of proposed change to flight operational systems. With adequate detail, it is possible that failures could be predicted with a high probability and action taken to eliminate failures. A major factor in any mission to the Moon and to Mars is the complexity of systems, interfaces, processes, their limitations, associated risks and the factor of the unknown including the development by many contractors and NASA centers. The need to be able to introduce new technologies over the life of the program requires an end to end test bed to analyze and evaluate these technologies and what will happen when they are introduced into the flight system. The ability to analyze system behaviors end to end under varying conditions would enhance safety e.g. fault tolerances. This analysis along with the ability to mine data from the development environment (e.g. test data), flight ops and modeling/simulations data would provide a level of information not currently available to operations and astronauts. In this paper we will analyze the beginnings of such a replication and what it could do in terms of reducing risk in the near term for development. We will analyze the Space Shuttle Main Engine (SSME) test lab which has to a large extent accomplished this replication for the SSME and has been highly successful in analyzing hardware and software problems and changes. The cost of replicating the flight system as proposed here could be very high if attempted as an afterthought. We will describe the initial steps for the development of a replication of this infrastructure starting with the communication infrastructure. The Constellation of Labs (CofL) under the Command, Control, Communication and Information (C3I) project for the NASA Exploration Initiative will provide the initial foundation upon which to base this replication. Simply put, there is very little margin for error in high latency situations e.g. en-route to/from Mars or in an autonomous process on the Lunar far side. Any thought out approach to reduce risk and increase safety needs to be accomplished end to end with the actual systems configuration.

Bradford, Robert N.↗

Transient upsets in microprocessor controllers

The modeling and analysis of transient faults in microprocessor based controllers are discussed. Such controllers typically consist of a microprocessor, read only memory storing and application program, random access memory for data storage, and input/output devices for external communications. The effects of transient faults on the performance of the controller are reviewed. An instruction level perspective of performance is taken which is the basis of a useful high level program state description of the microprocessor controller. A transition matrix is defined which determines the controller's response to transient fault arrivals.

Glaser, R. E.↗

Postseismic deformation due to subcrustal viscoelastic relaxation following dip-slip earthquakes

The deformation of the Earth following a dip-slip earthquake is calculated using a three layer rheological model and finite element techniques. The three layers are an elastic upper lithosphere, a standard linear solid lower lithosphere, and a Maxwell viscoelastic asthenosphere-a model previously analyzed in the strike-clip case (Cohen, 1981, 1982). Attention is focused on the magnitude of the postseismic subsidence and the width of the subsidence zone that can develop due to the viscoelastic response to coseismic reverse slip. Detailed analysis for a fault extending from the surface to 15 km with a 45 deg dip reveals that postseismic subsidence is sensitive to the depth to the asthenosphere but is only weakly dependent on lower lithosphere depth. The greatest subsidence occurs when the elastic lithosphere is about 30 km thick and the asthenosphere lies just below this layer (asthenosphere depth = 2 times the fault depth). The extremum in the subsidence pattern occurs at about 5 km from the surface trace of the fault and lies over the slip plane. In a typical case after a time t = 30 tau (tau = Maxwell time) following the earthquake the subsidence at this point is 60% of the coseismic uplift. Unlike the horizontal deformation following a strike slip earthquake, significant vertical deformation due to asthenosphere flow persists for many times tau and the magnitude of the vertical deformation is not necessarily enhanced by having a partially relaxing lower lithosphere.

Cohen, S. C.↗

Transfer-Function Simulator

Transfer function simulator constructed from analog or both analog and digital components substitute for device that has faults that confound analysis of feedback control loop. Simulator is substitute for laser and spectrophone.

Kavaya, M. J.↗

Postseismic deformation due to subcrustal viscoelastic relaxation following dip-slip earthquakes

The deformation of the earth following a dip-slip earthquake is calculated using a three layer rheological model and finite element techniques. The three layers are an elastic upper lithosphere, a standard linear solid lower lithosphere, and a Maxwell viscoelastic asthenosphere - a model previously analyzed in the strike-clip case (Cohen, 1981, 1982). Attention is focused on the magnitude of the postseismic subsidence and the width of the subsidence zone that can develop due to the viscoelastic response to coseismic reverse slip. Detailed analysis for a fault extending from the surface to 15 km with a 45 deg dip reveals that postseismic subsidence is sensitive to the depth to the asthenosphere but is only weakly dependent on lower lithosphere depth. The greatest subsidence occurs when the elastic lithosphere is about 30 km thick and the asthenosphere lies just below this layer (asthenosphere depth = 2 times the fault depth). The extremum in the subsidence pattern occurs at about 5 km from the surface trace of the fault and lies over the slip plane. In a typical case after a time t = 30 tau (tau = Maxwell time) following the earthquake, the subsidence at this point is 60 percent of the coseismic uplift. Unlike the horizontal deformation following a strike slip earthquake, significant vertical deformation due to asthenosphere flow persists for many times tau and the magnitude of the vertical deformation is not necessarily enhanced by having a partially relaxing lower lithosphere. Previously announced in STAR as N83-13683

Cohen, S. C.↗

Space platforms and autonomy

Potential applications for autonomous space platforms (SP) are discussed. The platforms are assumed to have long in-service lifetimes and therefore be flexible as to configuration modification and payload changeout. Higher degrees of autonomy, particularly from ground control, are made possible because of the rapid increase of microprocessor power and artificial intelligence advances. Functioning independently, the platforms are to rely only on periodic refurbishment visits by, e.g., the Orbiter. The Manned Space Station (MSS) will be the most complex structure, involving multifacted man-machine interfaces. The SP can be subsystems of the MSS (or other platforms), handling communications enunciation, data acquisition, analysis and telemetry, fault detection and isolation, systems monitoring and control, etc. The SP adopted will depend in all cases on costs vs benefits analyses to determine the worth of removing the function(s) from direct, regular human intervention.

Easter, R. W.↗

An astrometric facility for planetary detection on the Space Station

The preliminary system definition study for an Astrometric Telescope Facility (ATF) designed for the Space Station IOC is discussed, and a strawman system is designed which is found to meet the requirements for extrasolar planetary systems search and study. The strawman facility design, with a prime-focus 1.25-m aperture telescope and an f ratio of 13, was selected to minimize random and systematic errors. A basic operations approach is identified, including the approach to launch, initial on-orbit assembly and checkout, normal operations, and the response to anomolous conditions or failures. The preliminary system is designed to be fail-safe and single-fault tolerant. Mission analysis indicates that the basic viewing required for planetary detection can be accomplished in about 2/3 of the total viewing time.

Nishioka, Kenji↗

Mapping of Landsat satellite and gravity lineaments in west Tennessee

The analysis of earthquake fault lineament patterns within the alluvial valley of west Tennessee, which is often made difficult by the presence of unconsolidated sediments, is presently undertaken through a synergistic use of Landsat satellite images in conjunction with gravity anomaly data, which were quantitatively analyzed and compared by means of two-dimensional histograms and rose diagrams. The northeastern trend revealed for the lineaments corresponds to faults and is in keeping with reactivation of the Reelfoot rift near the Mississippi River; this suggests that deeper features, perhaps at earthquake focal depth, may extend to the land surface as Landsat-detectable lineaments.

Argialas, Demetre P.↗