Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Making the Hubble Space Telescope servicing mission safe

This paper will detail how the Hubble Space Telescope (HST) system safety program is conducted. Numerous safety analyses are conducted through the various phases of design, test, and fabrication, and results are presented to NASA management for discussion during dedicated safety reviews. This paper will then address the system safety assessment and risk analysis methodologies used (i.e. hazard analysis, fault tree analysis, and failure modes and effects analysis), and how they are coupled with enginering and test analyses for a 'synergistic picture' of the system. Some preliminary safety analysis results, showing the relationship between hazard identification, control or abatement, and finally control verification, will be presented as examples of this safety process.

Bahr, N. J.↗

Addressing Uniqueness and Unison of Reliability and Safety for a Better Integration

Over time, it has been observed that Safety and Reliability have not been clearly differentiated, which leads to confusion, inefficiency, and, sometimes, counter-productive practices in executing each of these two disciplines. It is imperative to address this situation to help Reliability and Safety disciplines improve their effectiveness and efficiency. The paper poses an important question to address, "Safety and Reliability - Are they unique or unisonous?" To answer the question, the paper reviewed several most commonly used analyses from each of the disciplines, namely, FMEA, reliability allocation and prediction, reliability design involvement, system safety hazard analysis, Fault Tree Analysis, and Probabilistic Risk Assessment. The paper pointed out uniqueness and unison of Safety and Reliability in their respective roles, requirements, approaches, and tools, and presented some suggestions for enhancing and improving the individual disciplines, as well as promoting the integration of the two. The paper concludes that Safety and Reliability are unique, but compensating each other in many aspects, and need to be integrated. Particularly, the individual roles of Safety and Reliability need to be differentiated, that is, Safety is to ensure and assure the product meets safety requirements, goals, or desires, and Reliability is to ensure and assure maximum achievability of intended design functions. With the integration of Safety and Reliability, personnel can be shared, tools and analyses have to be integrated, and skill sets can be possessed by the same person with the purpose of providing the best value to a product development.

Huang, Zhaofeng↗

Parameter Transient Behavior Analysis on Fault Tolerant Control System

In a fault tolerant control (FTC) system, a parameter varying FTC law is reconfigured based on fault parameters estimated by fault detection and isolation (FDI) modules. FDI modules require some time to detect fault occurrences in aero-vehicle dynamics. This paper illustrates analysis of a FTC system based on estimated fault parameter transient behavior which may include false fault detections during a short time interval. Using Lyapunov function analysis, the upper bound of an induced-L2 norm of the FTC system performance is calculated as a function of a fault detection time and the exponential decay rate of the Lyapunov function.

Belcastro, Christine↗

Eucalyptus – An Analysis Suite for Fault Trees with Uncertainty Quantification

Eucalyptus is a novel code developed at Lawrence Livermore National Laboratory to incorporate uncertainty quantification into Fault Tree Analysis (FTA). This tool addresses the challenge of imperfect knowledge in “grey-box” systems by allowing analysts to incorporate and propagate uncertainty from component-level assessments to system-level effects. Eucalyptus facilitates a consistent evaluation of the impact of subject matter expert judgment and knowledge gaps on overall system response by Monte Carlo generation of possible system fault trees, sampling probabilities of the existence of subsystems and components. Here, the code supports the specification of fault trees through text and allows export to various formats, including auto-generated images, easing analysis and reducing errors. It has undergone extensive verification testing, demonstrating its reliability and readiness for deployment, and leverages on-node parallelism for rapid analysis. Example analyses are shown that include the identification of system failure paths and quantification of the value of further information about system components.

Fault Tree Analysis↗

Analysis of faults and pit chains in Noctis Labyrinthus: Implications for early extension and possible magmatic plumbing

Noctis Labyrinthus has been a region of disputed origin due to its complexity and poor understanding of how various processes and mechanisms may have combined to form it. The surface is an integrated record of intensive tectonic activity expressed by a multiple extended sets of dip-slip faults oriented in different directions, and thought to have acted on this region over its history. These faults are always coalescent to pits and pit chains displaying a complicated geological history in the region. To understand this geological history, we mapped the surface features in Noctis Labyrinthus using the High-Resolution Stereo Camera (HRSC) onboard Mars Express ND2 nadir channel basemaps, and we adapted the Digital Terrain Map (DTM) from the Mission Experiment Gridded Data Record (MEGDR) of Mars Orbiter Laser Altimeter (MOLA) onboard Mars Global Surveyor (MGS) for the topography. We have investigated the spatial distribution and trend of fault systems, the pit chains' morphology, and the correlation between these two types of features. Our results show three fault systems: i) NS and NNE-SSW, ii) EW and ENE-WSW, and iii) NNW-SSE and NW. The analysis of the faults trending, cross-cutting correlation and the superimposition led to identify multiple intersections between these faults that have been alongside with the reactivations of some inherited faults. We interpreted the first system of fault to be related to coeval lateral extension, generated by regional stress tensor, which is probably related to the slight bending of Valles Marineris within two phases of bidirectional extension. The second system of faults has been generated by the radial oblate stress tensor related to the formation of the small shield volcanoes in Syria Planum. However, the third system is likely related to the external driving process, probably in the Tharsis province. We classified pits in four evolutionary stages based on their morphometric attributes. We believe that the formation of the pit chains in Noctis Labyrinthus is related to a surface collapse after a pressure drop related to the magma chamber deflation associated with Syria Planum volcanic province. We propose a deformational model based on early extension and magmatic plumbing as driving processes for the formation of Noctis Labyrinthus.

Mayssa El Yazidi↗

Analysis of faults detected in a large-scale multi-version software development experiment

In a multiversion software experiment, twenty programs were built to the same specification of an inertial navigation problem. The programs were then subjected to a three-phase testing and debugging process: an acceptance test, a certification test, and an operational test. Less than 20 percent of the faults discovered during the certification and operational testing were nonunique, i.e., the same or very similar faults would be found in more than one program. However, some of these common faults spanned as many as half of the versions. Faults discovered during the certification testing were due to specification errors and ambiguities, inadequate programmer background knowledge, insufficient programming experience, incomplete analysis, and insufficient acceptance testing. Faults discovered during the operational testing were of a more subtle nature, and were mostly due to various programmer knowledge defects and incomplete analysis errors. Techniques that might have prevented the observed faults are discussed.

Vouk, Mladen A.↗

Reliability computation using fault tree analysis

A method is presented for calculating event probabilities from an arbitrary fault tree. The method includes an analytical derivation of the system equation and is not a simulation program. The method can handle systems that incorporate standby redundancy and it uses conditional probabilities for computing fault trees where the same basic failure appears in more than one fault path.

Chelson, P. O.↗

Performance Analysis on Fault Tolerant Control System

In a fault tolerant control (FTC) system, a parameter varying FTC law is reconfigured based on fault parameters estimated by fault detection and isolation (FDI) modules. FDI modules require some time to detect fault occurrences in aero-vehicle dynamics. In this paper, an FTC analysis framework is provided to calculate the upper bound of an induced-L(sub 2) norm of an FTC system with existence of false identification and detection time delay. The upper bound is written as a function of a fault detection time and exponential decay rates and has been used to determine which FTC law produces less performance degradation (tracking error) due to false identification. The analysis framework is applied for an FTC system of a HiMAT (Highly Maneuverable Aircraft Technology) vehicle. Index Terms fault tolerant control system, linear parameter varying system, HiMAT vehicle.

Shin, Jong-Yeob↗

A Markov model reduction technique for fault tolerant processor reliability analysis

A fault tolerant processor (FTP) plays a key role in many high performance, safety-critical control system applications. Realistic modeling of an FTP is crucial to gaining a high degree of confidence in the reliability and safety analysis of such a system. While fidelity is clearly a major consideration, a practical model must also be kept to a moderate size to allow its incorporation into the overall system model. This paper presents a systematic reduction technique that starts from a complex, detailed model of a triple, redundant FTP and produces a low order approximation of very high accuracy. The existence of two distinct time scales represents the key to the success of the technique. No eigenvalue solution or coordinate transformation are needed. The reduced model captures all the important features of the detailed model, is amenable to an analytical solution and provides insight into the reconfiguration behavior of an FTP.

Schor, Andrei L.↗

Thermal Fault Tolerance Analysis of Carbon Fiber Rope Barrier Systems for Use in the Reusable Solid Rocket Motor ( RSRM) Nozzle Joints

Carbon Fiber Rope (CFR) thermal barrier systems are being considered for use in several RSRM (Reusable Solid Rocket Motor) nozzle joints as a replacement for the current assembly gap close-out process/design. This study provides for development and test verification of analysis methods used for flow-thermal modeling of a CFR thermal barrier subject to fault conditions such as rope combustion gas blow-by and CFR splice failure. Global model development is based on a 1-D (one dimensional) transient volume filling approach where the flow conditions are calculated as a function of internal 'pipe' and porous media 'Darcy' flow correlations. Combustion gas flow rates are calculated for the CFR on a per-linear inch basis and solved simultaneously with a detailed thermal-gas dynamic model of a local region of gas blow by (or splice fault). Effects of gas compressibility, friction and heat transfer are accounted for the model. Computational Fluid Dynamic (CFD) solutions of the fault regions are used to characterize the local flow field, quantify the amount of free jet spreading and assist in the determination of impingement film coefficients on the nozzle housings. Gas to wall heat transfer is simulated by a large thermal finite element grid of the local structure. The employed numerical technique loosely couples the FE (Finite Element) solution with the gas dynamics solution of the faulted region. All free constants that appear in the governing equations are calibrated by hot fire sub-scale test. The calibrated model is used to make flight predictions using motor aft end environments and timelines. Model results indicate that CFR barrier systems provide a near 'vented joint' style of pressurization. Hypothetical fault conditions considered in this study (blow by, splice defect) are relatively benign in terms of overall heating to nozzle metal housing structures.

Clayton, J. Louie↗

A System for Fault Management and Fault Consequences Analysis for NASA's Deep Space Habitat

NASA's exploration program envisions the utilization of a Deep Space Habitat (DSH) for human exploration of the space environment in the vicinity of Mars and/or asteroids. Communication latencies with ground control of as long as 20+ minutes make it imperative that DSH operations be highly autonomous, as any telemetry-based detection of a systems problem on Earth could well occur too late to assist the crew with the problem. A DSH-based development program has been initiated to develop and test the automation technologies necessary to support highly autonomous DSH operations. One such technology is a fault management tool to support performance monitoring of vehicle systems operations and to assist with real-time decision making in connection with operational anomalies and failures. Toward that end, we are developing Advanced Caution and Warning System (ACAWS), a tool that combines dynamic and interactive graphical representations of spacecraft systems, systems modeling, automated diagnostic analysis and root cause identification, system and mission impact assessment, and mitigation procedure identification to help spacecraft operators (both flight controllers and crew) understand and respond to anomalies more effectively. In this paper, we describe four major architecture elements of ACAWS: Anomaly Detection, Fault Isolation, System Effects Analysis, and Graphic User Interface (GUI), and how these elements work in concert with each other and with other tools to provide fault management support to both the controllers and crew. We then describe recent evaluations and tests of ACAWS on the DSH testbed. The results of these tests support the feasibility and strength of our approach to failure management automation and enhanced operational autonomy

System Effects Analysis↗

EV-EVSE Fault Study: An Analysis of Thermal Events Caused by Electrical Faults during DC Charging

This report outlines multiple avenues of analysis of electrical faults associated with electric vehicles (EVs) during charging, focusing specifically on the interactions between EVs and EV supply equipment (EVSE). Key concerns include the identification and mitigation of overtemperature events that can result in fires, which are commonly initiated by localized heating of connectors, wiring, or high-impedance fault current paths.

42 - ENGINEERING↗

Modular techniques for dynamic fault-tree analysis

It is noted that current approaches used to assess the dependability of complex systems such as Space Station Freedom and the Air Traffic Control System are incapable of handling the size and complexity of these highly integrated designs. A novel technique for modeling such systems which is built upon current techniques in Markov theory and combinatorial analysis is described. It enables the development of a hierarchical representation of system behavior which is more flexible than either technique alone. A solution strategy which is based on an object-oriented approach to model representation and evaluation is discussed. The technique is virtually transparent to the user since the fault tree models can be built graphically and the objects defined automatically. The tree modularization procedure allows the two model types, Markov and combinatoric, to coexist and does not require that the entire fault tree be translated to a Markov chain for evaluation. This effectively reduces the size of the Markov chain required and enables solutions with less truncation, making analysis of longer mission times possible. Using the fault-tolerant parallel processor as an example, a model is built and solved for a specific mission scenario and the solution approach is illustrated in detail.

Patterson-Hine, F. A.↗

The application of Skylab imagery to analysis of fault tectonics and earthquake hazards in the Peninsular Ranges, southern California

The author has identified the following significant results. Frame 114 of the Salton Sea area was studied in all bands to analyze the appearance of important faults. These faults were also studied in the field as well as from aircraft and in aerial photography. The San Andreas/Banning and the Mission Creek faults can be traced across Coachella Valley even though they are buried by alluvium. The faults form ground water barriers and the near surface ground water on the northeast sides of the faults supports patches of vegetation (mesquite and palms) in an otherwise barren desert. These oases are best seen in band 3 (color IR). Otherwise, faults are best seen in band 4 (aerial color). Of the B and W bands, 5 (red) is best for delineating faults. Bands 1 and 2 are excessively grainy and the resolution is considerably inferior to the other bands.

Merifield, P. M.↗

Software Requirements Analysis as Fault Predictor

Waiting until the integration and system test phase to discover errors leads to more costly rework than resolving those same errors earlier in the lifecycle. Costs increase even more significantly once a software system has become operational. WE can assess the quality of system requirements, but do little to correlate this information either to system assurance activities or long-term reliability projections - both of which remain unclear and anecdotal. Extending earlier work on requirements accomplished by the ARM tool, measuring requirements quality information against code complexity and test data for the same system may be used to predict specific software modules containing high impact or deeply embedded faults now escaping in operational systems. Such knowledge would lead to more effective and efficient test programs. It may enable insight into whether a program should be maintained or started over.

Wallace, Dolores↗

NASA Taxonomies for Searching Problem Reports and FMEAs

Many types of hazard and risk analyses are used during the life cycle of complex systems, including Failure Modes and Effects Analysis (FMEA), Hazard Analysis, Fault Tree and Event Tree Analysis, Probabilistic Risk Assessment, Reliability Analysis and analysis of Problem Reporting and Corrective Action (PRACA) databases. The success of these methods depends on the availability of input data and the analysts knowledge. Standard nomenclature can increase the reusability of hazard, risk and problem data. When nomenclature in the source texts is not standard, taxonomies with mapping words (sets of rough synonyms) can be combined with semantic search to identify items and tag them with metadata based on a rich standard nomenclature. Semantic search uses word meanings in the context of parsed phrases to find matches. The NASA taxonomies provide the word meanings. Spacecraft taxonomies and ontologies (generalization hierarchies with attributes and relationships, based on terms meanings) are being developed for types of subsystems, functions, entities, hazards and failures. The ontologies are broad and general, covering hardware, software and human systems. Semantic search of Space Station texts was used to validate and extend the taxonomies. The taxonomies have also been used to extract system connectivity (interaction) models and functions from requirements text. Now the Reconciler semantic search tool and the taxonomies are being applied to improve search in the Space Shuttle PRACA database, to discover recurring patterns of failure. Usual methods of string search and keyword search fall short because the entries are terse and have numerous shortcuts (irregular abbreviations, nonstandard acronyms, cryptic codes) and modifier words cannot be used in sentence context to refine the search. The limited and fixed FMEA categories associated with the entries do not make the fine distinctions needed in the search. The approach assigns PRACA report titles to problem classes in the taxonomy. Each ontology class includes mapping words - near-synonyms naming different manifestations of that problem class. The mapping words for Problems, Entities and Functions are converted to a canonical form plus any of a small set of modifier words (e.g. non-uniformity NOT + UNIFORM.) The report titles are parsed as sentences if possible, or treated as a flat sequence of word tokens if parsing fails. When canonical forms in the title match mapping words, the PRACA entry is associated with the corresponding Problem, Entity or Function in the ontology. The user can search for types of failures associated with types of equipment, clustering by type of problem (e.g., all bearings found with problems of being uneven: rough, irregular, gritty ). The results could also be used for tagging PRACA report entries with rich metadata. This approach could also be applied to searching and tagging failure modes, failure effects and mitigations in FMEAs. In the pilot work, parsing 52K+ truncated titles (the test cases that were available), has resulted in identification of both a type of equipment and type of problem in about 75% of the cases. The results are displayed in a manner analogous to Google search results. The effort has also led to the enrichment of the taxonomy, adding some new categories and many new mapping words. Further work would make enhancements that have been identified for improving the clustering and further reducing the false alarm rate. (In searching for recurring problems, good clustering is more important than reducing false alarms). Searching complete PRACA reports should lead to immediate improvement.

Malin, Jane T.↗

Machine learning of fault characteristics from rocket engine simulation data

Transformation of data into knowledge through conceptual induction has been the focus of our research described in this paper. We have developed a Machine Learning System (MLS) to analyze the rocket engine simulation data. MLS can provide to its users fault analysis, characteristics, and conceptual descriptions of faults, and the relationships of attributes and sensors. All the results are critically important in identifying faults.

Ke, Min↗