Search NASA⌕ Search

SEARCH · Search NASA

Results for “Systems Engineering, Failure Prevention”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Systems Engineering with a Focus on Failure Prevention

A discussion of Systems Engineering principles as related to failure prevention methodologies and technologies. The presentation outlines requirements development, design reviews, verification and validation, and finally risk management prectices.

Systems Engineering, Failure Prevention↗

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering↗

Key Reliability Drivers of Liquid Propulsion Engines and A Reliability Model for Sensitivity Analysis

This paper is to address the in-flight reliability of a liquid propulsion engine system for a launch vehicle. We first establish a comprehensive list of system and sub-system reliability drivers for any liquid propulsion engine system. We then build a reliability model to parametrically analyze the impact of some reliability parameters. We present sensitivity analysis results for a selected subset of the key reliability drivers using the model. Reliability drivers identified include: number of engines for the liquid propulsion stage, single engine total reliability, engine operation duration, engine thrust size, reusability, engine de-rating or up-rating, engine-out design (including engine-out switching reliability, catastrophic fraction, preventable failure fraction, unnecessary shutdown fraction), propellant specific hazards, engine start and cutoff transient hazards, engine combustion cycles, vehicle and engine interface and interaction hazards, engine health management system, engine modification, engine ground start hold down with launch commit criteria, engine altitude start (1 in. start), Multiple altitude restart (less than 1 restart), component, subsystem and system design, manufacturing/ground operation support/pre and post flight check outs and inspection, extensiveness of the development program. We present some sensitivity analysis results for the following subset of the drivers: number of engines for the propulsion stage, single engine total reliability, engine operation duration, engine de-rating or up-rating requirements, engine-out design, catastrophic fraction, preventable failure fraction, unnecessary shutdown fraction, and engine health management system implementation (basic redlines and more advanced health management systems).

Huang, Zhao-Feng↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Text-based information about past engineering failures presents one possible solution by facilitating the retrieval of information that can inform new designs. However, identifying documents containing relevant information and extracting required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing-based framework is proposed to discover relevant knowledge from documents containing failure-related design information. Documents containing usable information are filtered using sentiment analysis based on a custom lexicon specialized for engineering design and by filtering out documents containing only irrelevant topics. Next, from the identified usable documents, information relating to engineering failures, contributing factors that can be controlled at design time (“risk factors”), and recommended preventative actions are extracted. Semantic similarity is then used to group similar pieces of extracted information for improved generalizability. The proposed framework is applied to NASA’s Lessons Learned Information System (LLIS). The framework can be used to identify documents containing usable failure-related design information from other databases, extract relevant information from these documents, and generalize the acquired knowledge such that it can be applied to novel systems.

Sequoia R. Andrade↗

Explicit Finite Element Modeling of Multilayer Composite Fabric for Gas Turbine Engine Containment Systems: Ballistic Impact Testing - Part 2

Under the Federal Aviation Administration's Airworthiness Assurance Center of Excellence and the Aircraft Catastrophic Failure Prevention Program, National Aeronautics and Space Administration Glenn Research Center collaborated with Arizona State University, Honeywell Engines, Systems and Services, and SRI International to develop improved computational models for designing fabric-based engine containment systems. In the study described in this report, ballistic impact tests were conducted on layered dry fabric rings to provide impact response data for calibrating and verifying the improved numerical models. This report provides data on projectile velocity, impact and residual energy, and fabric deformation for a number of different test conditions.

ZYLON↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science↗

Algorithm Helps Monitor Engine Operation

Real-Time Failure Control (RTFC) algorithm part of automated monitoring-and-shutdown system being developed to ensure safety and prevent major damage to equipment during ground tests of main engine of space shuttle. Includes redundant sensors, controller voting logic circuits, automatic safe-limit logic circuits, and conditional-decision logic circuits, all monitored by human technicians. Basic principles of system also applicable to stationary powerplants and other complex machinery systems.

Eckerling, Sherry J.↗

Probabilistic Risk Analysis and Margin Process for a Flexible Thermal Protection System

Atmospheric entry vehicle thermal protection systems are margined due to the uncertainties that exist in entry aeroheating environments and the thermal response of the materials and structures. Entry vehicle thermal protections systems are traditionally over-margined for the heat loads that are experienced along the entry trajectory by designing to survive stacked worst-case scenarios. Additionally, the conventional heat shield design and margin process offers very little insight into the risk of over-temperature during flight and the corresponding reliability of the heat shield performance. A probabilistic margin process can be used to appropriately margin the thermal protection system based on rigorously calculated risk of failure. This probabilistic margin process allows engineers to make informed aeroshell design, entry-trajectory design, and risk trades while preventing excessive margin from being applied. This study presents the methods of the probabilistic margin process and how the uncertainty analysis is used to determine the reliability of the entry vehicle thermal protection system and associated risks of failure.

Tobin, Steven A.↗

J-2X Turbopump Cavitation Diagnostics

The J-2X is the upper stage engine currently being designed by Pratt & Whitney Rocketdyne (PWR) for the Ares I Crew Launch Vehicle (CLV). Propellant supply requirements for the J-2X are defined by the Ares Upper Stage to J-2X Interface Control Document (ICD). Supply conditions outside ICD defined start or run boxes can induce turbopump cavitation leading to interruption of J-2X propellant flow during hot fire operation. In severe cases, cavitation can lead to uncontained engine failure with the potential to cause a vehicle catastrophic event. Turbopump and engine system performance models supported by system design information and test data are required to predict existence, severity, and consequences of a cavitation event. A cavitation model for each of the J-2X fuel and oxidizer turbopumps was developed using data from pump water flow test facilities at Pratt & Whitney Rocketdyne (PWR) and Marshall Space Flight Center (MSFC) together with data from Powerpack 1A testing at Stennis Space Center (SSC) and from heritage systems. These component models were implemented within the PWR J-2X Real Time Model (RTM) to provide a foundation for predicting system level effects following turbopump cavitation. The RTM serves as a general failure simulation platform supporting estimation of J-2X redline system effectiveness. A study to compare cavitation induced conditions with component level structural limit thresholds throughout the engine was performed using the RTM. Results provided insight into system level turbopump cavitation effects and redline system effectiveness in preventing structural limit violations. A need to better understand structural limits and redline system failure mitigation potential in the event of fuel side cavitation was indicated. This paper examines study results, efforts to mature J-2X turbopump cavitation models and structural limits, and issues with engine redline detection of cavitation and the use of vehicle-side abort triggers to augment the engine redline system.

Santi, I. Michael↗

Embedded expert system for space shuttle main engine maintenance

The SPARTA Embedded Expert System (SEES) is an intelligent health monitoring system that directs analysis by placing confidence factors on possible engine status and then recommends a course of action to an engineer or engine controller. The technique can prevent catastropic failures or costly rocket engine down time because of false alarms. Further, the SEES has potential as an on-board flight monitor for reusable rocket engine systems. The SEES methodology synergistically integrates vibration analysis, pattern recognition and communications theory techniques with an artificial intelligence technique - the Embedded Expert System (EES).

Pooley, J.↗

Real-Time Diagnosis of Faults Using a Bank of Kalman Filters

A new robust method of automated real-time diagnosis of faults in an aircraft engine or a similar complex system involves the use of a bank of Kalman filters. In order to be highly reliable, a diagnostic system must be designed to account for the numerous failure conditions that an aircraft engine may encounter in operation. The method achieves this objective though the utilization of multiple Kalman filters, each of which is uniquely designed based on a specific failure hypothesis. A fault-detection-and-isolation (FDI) system, developed based on this method, is able to isolate faults in sensors and actuators while detecting component faults (abrupt degradation in engine component performance). By affording a capability for real-time identification of minor faults before they grow into major ones, the method promises to enhance safety and reduce operating costs. The robustness of this method is further enhanced by incorporating information regarding the aging condition of an engine. In general, real-time fault diagnostic methods use the nominal performance of a "healthy" new engine as a reference condition in the diagnostic process. Such an approach does not account for gradual changes in performance associated with aging of an otherwise healthy engine. By incorporating information on gradual, aging-related changes, the new method makes it possible to retain at least some of the sensitivity and accuracy needed to detect incipient faults while preventing false alarms that could result from erroneous interpretation of symptoms of aging as symptoms of failures. The figure schematically depicts an FDI system according to the new method. The FDI system is integrated with an engine, from which it accepts two sets of input signals: sensor readings and actuator commands. Two main parts of the FDI system are a bank of Kalman filters and a subsystem that implements FDI decision rules. Each Kalman filter is designed to detect a specific sensor or actuator fault. When a sensor or actuator fault occurs, large estimation errors are generated by all filters except the one using the correct hypothesis. By monitoring the residual output of each filter, the specific fault that has occurred can be detected and isolated on the basis of the decision rules. A set of parameters that indicate the performance of the engine components is estimated by the "correct" Kalman filter for use in detecting component faults. To reduce the loss of diagnostic accuracy and sensitivity in the face of aging, the FDI system accepts information from a steady-state-condition-monitoring system. This information is used to update the Kalman filters and a data bank of trim values representative of the current aging condition.

Kobayashi, Takahisa↗

Spacecraft Testing Programs: Adding Value to the Systems Engineering Process

Testing has long been recognized as a critical component of spacecraft development activities - yet many major systems failures may have been prevented with more rigorous testing programs. The question is why is more testing not being conducted? Given unlimited resources, more testing would likely be included in a spacecraft development program. Striking the right balance between too much testing and not enough has been a long-term challenge for many industries. The objective of this paper is to discuss some of the barriers, enablers, and best practices for developing and sustaining a strong test program and testing team. This paper will also explore the testing decision factors used by managers; the varying attitudes toward testing; methods to develop strong test engineers; and the influence of behavior, culture and processes on testing programs. KEY WORDS: Risk, Integration and Test, Validation, Verification, Test Program Development

Britton, Keith J.↗

Pyrotechnic system failures: Causes and prevention

Although pyrotechnics have successfully accomplished many critical mechanical spacecraft functions, such as ignition, severance, jettisoning and valving (excluding propulsion), failures continue to occur. Provided is a listing of 84 failures of pyrotechnic hardware with completed design over a 23-year period, compiled informally by experts from every NASA Center, as well as the Air Force Space Division and the Naval Surface Warfare Center. Analyses are presented as to when and where these failures occurred, their technical source or cause, followed by the reasons why and how these kinds of failures persist. The major contributor is a fundamental lack of understanding of the functional mechanisms of pyrotechnic devices and systems, followed by not recognizing pyrotechnics as an engineering technology, insufficient manpower with hands-on experience, too few test facilities, and inadequate guidelines and specifications for design, development, qualification and acceptance. Recommendations are made on both a managerial and technical basis to prevent failures, increase reliability, improve existing and future designs, and develop the technology to meet future requirements.

Bement, Laurence J.↗

Flight experience with Apollo spacecraft propulsion systems

Apollo 17 ended the most successful application of rocket propulsion systems in man's history. A total of 23 developmental and manned operational flights were made. Seven hundred and sixty-three spacecraft rocket engines were flown in the program. Over 6 h of manned rocket flights were logged by the spacecraft propulsion systems and approximately one million rocket engine firings were made. One engine failure was encountered on an early unmanned flight as a result of a failure in the guidance programmer which caused the engine to operate in a manner known to cause failures. Numerous operational problems and malfunctions were observed; however, system and component redundancy prevented loss of mission objectives and never jeopardized crew safety. Performance of all systems was usually nominal and most problems were merely nuisances. This paper will present some highlights of Apollo propulsion performance and will provide a bibliography of all flight results.

Thibodaux, J. G., Jr.↗

Applicability of a Crack-Detection System for Use in Rotor Disk Spin Test Experiments Being Evaluated

Engine makers and aviation safety government institutions continue to have a strong interest in monitoring the health of rotating components in aircraft engines to improve safety and to lower maintenance costs. To prevent catastrophic failure (burst) of the engine, they use nondestructive evaluation (NDE) and major overhauls for periodic inspections to discover any cracks that might have formed. The lowest cost fluorescent penetrant inspection NDE technique can fail to disclose cracks that are tightly closed during rest or that are below the surface. The NDE eddy current system is more effective at detecting both crack types, but it requires careful setup and operation and only a small portion of the disk can be practically inspected. So that sensor systems can sustain normal function in a severe environment, health-monitoring systems require the sensor system to transmit a signal if a crack detected in the component is above a predetermined length (but below the length that would lead to failure) and lastly to act neutrally upon the overall performance of the engine system and not interfere with engine maintenance operations. Therefore, more reliable diagnostic tools and high-level techniques for detecting damage and monitoring the health of rotating components are very essential in maintaining engine safety and reliability and in assessing life.

Abdul-Aziz, Ali↗

A Prognostic Launch Vehicle Probability of Failure Assessment Methodology for Conceptual Systems Predicated on Human Causal Factors

Create an improved method to calculate reliability of a conceptual launch vehicle system prior to fabrication by using historic data of actual root causes of failures. While failures have unique "proximate causes", there are typically a finite amount of common "root causes". Heretofore launch vehicle reliability evaluation typically hardware-centric statistical analyses, while most root causes of failures are been shown to be human-centric. A method based on human-centric root causes can be used to quantify reliability assessments and focus proposed actions to mitigate problems. Existing methods have been optimistic in their projections of launch vehicle reliability compared to actuals. Hypothesis: reliability of a conceptual launch vehicle can be more accurately evaluated based on a rational, probabilistic approach using past failure assessment teams' findings predicated on human-centric causes."Human Reliability Analysis Methods Selection Guidance for NASA"Chandler F.T., et al., NASA HQ/OSMA study group, July 2006. Outside HRA experts from academia, other federal labs, and the private sector. 50 system reliability methods considered, fourteen selected for further study, four finally selected as best suited for human spaceflight. Probabilistic Risk Analysis (PRA) + Human Reliability Analysis (HRA) enabled incorporating effects and probabilities of human errors. While four down-selected methods deemed appropriate for failure assessment, it did not appear that these methods could be concisely applied to perform major system-wide assessment of probability of failure of a conceptual design without becoming unwieldy."Engineering a Safer World", Detailed, comprehensive study external to NASA Leveson N. G., MIT, 2011.Systems-Theoretic Accident Model and Processes (STAMP). All-encompassing accident model based on systems theory analyzed accidents after they occurred and created approaches to prevent occurrence in developing systems not focused on failure prevention per se, but rather reducing hazards by influencing human behavior through use of constraints, hierarchical control structures, and process models to improve system safetySystem Theoretic Process Analysis (STPA) addresses predictive part of problem (a "hazard analysis"). Includes all causal factors identified in STAMP: "...design errors, software flaws, component interaction accidents, cognitively complex human decision-making errors, and social organizational and management factors contributing to accidents" can guide design process rather than require it to exist before-hand did not appear capable of concise application for system-wide assessment of probability of failure of a conceptual design without becoming unwieldy.

Williams, Craig H.↗

Space Transportation System Availability Requirements and Its Influencing Attributes Relationships

It is essential that management and engineering understand the need for an availability requirement for the customer's space transportation system as it enables the meeting of his needs, goal, and objectives. There are three types of availability, e.g., operational availability, achieved availability, or inherent availability. The basic definition of availability is equal to the mean uptime divided by the sum of the mean uptime plus the mean downtime. The major difference is the inclusiveness of the functions within the mean downtime and the mean uptime. This paper will address tIe inherent availability which only addresses the mean downtime as that mean time to repair or the time to determine the failed article, remove it, install a replacement article and verify the functionality of the repaired system. The definitions of operational availability include the replacement hardware supply or maintenance delays and other non-design factors in the mean downtime. Also with inherent availability the mean uptime will only consider the mean time between failures (other availability definitions consider this as mean time between maintenance - preventive and corrective maintenance) that requires the repair of the system to be functional. It is also essential that management and engineering understand all influencing attributes relationships to each other and to the resultant inherent availability requirement. This visibility will provide the decision makers with the understanding necessary to place constraints on the design definition for the major drivers that will determine the inherent availability, safety, reliability, maintainability, and the life cycle cost of the fielded system provided the customer. This inherent availability requirement may be driven by the need to use a multiple launch approach to placing humans on the moon or the desire to control the number of spare parts required to support long stays in either orbit or on the surface of the moon or mars. It is the intent of this paper to provide the visibility of relationships of these major attribute drivers (variables) to each other and the resultant system inherent availability, but also provide the capability to bound the variables providing engineering the insight required to control the system's engineering solution. An example of this visibility will be the need to provide integration of similar discipline functions to allow control of the total parts count of the space transportation system. Also the relationship visibility of selecting a reliability requirement will place a constraint on parts count to achieve a given inherent availability requirement or accepting a larger parts count with the resulting higher reliability requirement. This paper will provide an understanding for the relationship of mean repair time (mean downtime) to maintainability, e.g., accessibility for repair, and both mean time between failure, e.g., reliability of hardware and the system inherent availability. Having an understanding of these relationships and resulting requirements before starting the architectural design concept definition will avoid considerable time and money required to iterate the design to meet the redesign and assessment process required to achieve the results required of the customer's space transportation system. In fact the impact to the schedule to being able to deliver the system that meets the customer's needs, goals, and objectives may cause the customer to compromise his desired operational goal and objectives resulting in considerable increased life cycle cost of the fielded space transportation system.

Rhodes, Russel E.↗