Search NASASearch

SEARCH · Search NASA

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Requirements Engineering Scorecard and the Next-Generation Space Suit

The objective for a NASA contractor, the performing organization in this case study, is to provide engineering services to develop and deliver the next generation space suit to NASA, the customer in this case study. A case study with qualitative and quantitative analyses regarding a new process and approach to requirements engineering is described, with the intent that if utilized, these tools may have contributed to improvements across the project in terms of meeting cost, scope, budget and quality while appropriately accounting for risk management. The procedure entails a research method in which the current state of the project, current state of the art, and the identified systems engineering challenges are evaluated. Iterative models are tempered through development by continual improvements by engineering evaluation of engineers on the project. The current results have produced a prototype of a requirements engineering scorecard with implementations of FMEA and quantitative analysis to (i) identify root cause of underdeveloped requirements and (ii) project management impacts with regards to project risk. Forward work includes customer, performing organization, acceptance against applicable INCOSE community accepted practices.

INCOSE

JPL/NASA/IEEE Test Effectiveness Workshop

(none given)From OBJECTIVES: Specific objectives of the working group are to support the innovation, development, evaluation and implementation of test methods, metrics and tools based on failure engineering/physics and/or root cause evaluations. Data sources systems and tools shall be developed and implemented that: 1) provide improved preventions, controls, analyses and tests (PACT) & field failure data collection, 2) facilities data analysis, archiving, retrieval, failure physics and/or root cause evaluations and 3) enable new and existing technology suitability evaluations to be performed.

effectiveness concurrent engineering metrics test

Space Shuttle Stiffener Ring Foam Failure, a Non-Conventional Approach

The Space Shuttle makes use of the excellent properties of rigid polyurethane foam for cryogenic tank insulation and as structural protection on the solid rocket boosters. When foam applications debond, classical methods of analysis do not always provide root cause of the failure of the foam. Realizing that foam is the ideal media to document and preserve its own mode of failure, thin sectioning was seen as a logical approach for foam failure analysis. Thin sectioning in two directions, both horizontal and vertical to the application, was chosen to observe the three dimensional morphology of the foam cells. The cell foam morphology provided a much greater understanding of the failure modes than previously achieved.

Howard, Philip M.

Pendulum Motion in Main Parachute Clusters

The coupled dynamics of a cluster of parachutes to a payload are notoriously difficult to predict. Often the payload is designed to be insensitive to the range of attitude and rates that might occur, but spacecraft generally do not have the mass and volume budgeted for this robust of a design. The National Aeronautics and Space Administration (NASA) Orion Capsule Parachute Assembly System (CPAS) implements a cluster of three mains for landing. During testing of the Engineering Development Unit (EDU) design, it was discovered that with a cluster of two mains (a fault tolerance required for human rating) the capsule coupled to the parachute cluster could get into a limit cycle pendulum motion which would exceed the spacecraft landing capability. This pendulum phenomenon could not be predicted with the existing models and simulations. A three phased effort has been undertaken to understand the consequence of the pendulum motion observed, and explore potential design changes that would mitigate this phenomenon. This paper will review the early analysis that was performed of the pendulum motion observed during EDU testing, summarize the analysis ongoing to understand the root cause of the pendulum phenomenon, and discuss the modeling and testing that is being pursued to identify design changes that would mitigate the risk.

Ray, Eric S.

Report of the Odyssey FPGA Independent Assessment Team

An independent assessment team (IAT) was formed and met on April 2, 2001, at Lockheed Martin in Denver, Colorado, to aid in understanding a technical issue for the Mars Odyssey spacecraft scheduled for launch on April 7, 2001. An RP1280A field-programmable gate array (FPGA) from a lot of parts common to the SIRTF, Odyssey, and Genesis missions had failed on a SIRTF printed circuit board. A second FPGA from an earlier Odyssey circuit board was also known to have failed and was also included in the analysis by the IAT. Observations indicated an abnormally high failure rate for flight RP1280A devices (the first flight lot produced using this flow) at Lockheed Martin and the causes of these failures were not determined. Standard failure analysis techniques were applied to these parts, however, additional diagnostic techniques unique for devices of this class were not used, and the parts were prematurely submitted to a destructive physical analysis, making a determination of the root cause of failure difficult. Any of several potential failure scenarios may have caused these failures, including electrostatic discharge, electrical overstress, manufacturing defects, board design errors, board manufacturing errors, FPGA design errors, or programmer errors. Several of these mechanisms would have relatively benign consequences for disposition of the parts currently installed on boards in the Odyssey spacecraft if established as the root cause of failure. However, other potential failure mechanisms could have more dire consequences. As there is no simple way to determine the likely failure mechanisms with reasonable confidence before Odyssey launch, it is not possible for the IAT to recommend a disposition for the other parts on boards in the Odyssey spacecraft based on sound engineering principles.

Mayer, Donald C.

Ground Operations Autonomous Control and Integrated Health Management

An intelligent autonomous control capability has been developed and is currently being validated in ground cryogenic fluid management operations. The capability embodies a physical architecture consistent with typical launch infrastructure and control systems, augmented by a higher level autonomous control (AC) system enabled to make knowledge-based decisions. The AC system is supported by an integrated system health management (ISHM) capability that detects anomalies, diagnoses causes, determines effects, and could predict future anomalies. AC is implemented using the concept of programmed sequences that could be considered to be building blocks of more generic mission plans. A sequence is a series of steps, and each executes actions once conditions for the step are met (e.g. desired temperatures or fluid state are achieved). For autonomous capability, conditions must consider also health management outcomes, as they will determine whether or not an action is executed, or how an action may be executed, or if an alternative action is executed instead. Aside from health, higher level objectives can also drive how a mission is carried out. The capability was developed using the G2 software environment (www.gensym.com) augmented by a NASA Toolkit that significantly shortens time to deployment. G2 is a commercial product to develop intelligent applications. It is fully object oriented. The core of the capability is a Domain Model of the system where all elements of the system are represented as objects (sensors, instruments, components, pipes, etc.). Reasoning and decision making can be done with all elements in the domain model. The toolkit also enables implementation of failure modes and effects analysis (FMEA), which are represented as root cause trees. FMEA's are programmed graphically, they are reusable, as they address generic FMEA referring to classes of subsystems or objects and their functional relationships. User interfaces for integrated awareness by operators have been created.

Figueroa, Fernando

Digital Image Correlation Data Processing and Analysis Techniques to Enhance Test Data Assessment and Improve Structural Simulations

The NASA Shell Buckling Knockdown Factor Project (SBKF) was established in 2007 by the NASA Engineering and Safety Center (NESC) with the primary goal to develop new analysis-based buckling design factors (a.k.a. knockdown factors) and high-fidelity buckling simulations for selected launch-vehicle-like cylindrical shell structures. A series of tests are being conducted on large-scale metallic and composite cylindrical shells in order to provide validation data for these new factors and simulations. However, the validation of these new factors and simulations is quite demanding and requires test data that is commensurate with their fidelity. Traditional instrumentation, such as linear variable displacement transducers (LVDTs) and electrical-resistance strain gages serve a critical role in providing accurate displacement and strain measurements in these tests, but only allow for data to be recorded at a select number of point locations and are not sufficient to provide all the necessary validation data. Advanced measurement technologies can be used effectively to complement traditional instrumentation and gather additional data required to validate these structural simulations. In particular, three-dimensional digital image correlation (DIC) was implemented during SBKF cylinder testing to characterize the full-field displacement and strain behavior. Commercially available VIC-3DTM software and user-written data processing scripts were used to generate valuable data and insight into the complex buckling response of the cylinders that otherwise would be impossible to gather using traditional instrumentation. In addition, the measured data from DIC was used to verify measured test data obtained from other instrumentation, enhance test and analysis correlation, and help identify the root cause of anomalous test results that may have gone unexplained if only traditional instrumentation was used. Selected test results that demonstrate the use of DIC on the SBKF cylinders are presented and a portion of the data processing methods are described.

Gardner, Nathaniel W.

Application of PoF Based Virtual Qualification Methods for Reliability Assessment of Mission Critical PCBs

Reliability is the ability of a product to perform the function for which it was intended for a specified period of time (or cycles) for a given set of life cycle conditions. In today's compressed mission development cycles where designing, building and testing the physical models has to occur in a matter of months not years, Projects don't have the luxury of iteratively building and testing those models. Physics of failure (PoF) is an engineering-based approach to reliability that begins with an understanding of materials, processes, physical interactions, degradation and failure mechanisms, as well as identifying failure models. The PoF approach uses modeling and simulation to qualify a design and manufacturing process, with the ultimate intent of eliminating failures early in the design process by addressing the root cause. The physics-of-failure analysis proactively incorporates reliability into the design process by establishing a scientific basis for evaluating new materials, structures and technologies. Virtual physics-of-failure modeling allows engineers to determine if new technological node can be added to an existing system. This presentation will illustrate an application of a PoF based tool during the initial phases of a printed circuit board assembly development and how the NASA GSFC team was able to dynamically study the effects of electronics parts and printed circuit board material configuration changes under simulated thermal and vibrational stresses

Electronics packaging

Application of PoF Based Virtual Qualification Methods for Reliability Assessment of Mission Critical PCBs

Reliability is the ability of a product to perform the function for which it was intended for a specified period of time (or cycles) for a given set of life cycle conditions. In today's compressed mission development cycles where designing, building and testing the physical models has to occur in a matter of months not years, Projects don't have the luxury of iteratively building and testing those models. Physics of failure (PoF) is an engineering-based approach to reliability that begins with an understanding of materials, processes, physical interactions, degradation and failure mechanisms, as well as identifying failure models. The PoF approach uses modeling and simulation to qualify a design and manufacturing process, with the ultimate intent of eliminating failures early in the design process by addressing the root cause. The physics-of-failure analysis proactively incorporates reliability into the design process by establishing a scientific basis for evaluating new materials, structures and technologies. Virtual physics-of-failure modeling allows engineers to determine if new technological node can be added to an existing system. This presentation will illustrate an application of a PoF based tool during the initial phases of a printed circuit board assembly development and how the NASA GSFC team was able to dynamically study the effects of electronics parts and printed circuit board material configuration changes under simulated thermal and vibrational stresses.

Physics of Failure

A Risk Analysis Tool for Estimating the Risk of Electrical Failures Due to Human Induced Defects

Aerospace electrical systems are required to withstand and adequately operate in extremely harsh environments that include, for example, high radiation exposure, temperature extremes, intense vibrational stress and drastic temperature cycling. The nature of aerospace electronics also demands high reliability since, with very few exceptions, there is no chance for hardware servicing or repairs. Common risk mitigation techniques for this type of situation are to perform a Reliability Analysis of the system throughout the development cycle, and to use electrical components that are regarded as “high reliability” because of additional controls and requirements applied in their design, manufacturing and testing. Unfortunately, studies have shown that even though these techniques are used, many systems fail to meet mission requirements well before the predicted lifetimes. This paper presents the analysis of failures of electrical parts, experienced during various stages of system development, at NASA Goddard Space Flight Center, Greenbelt MD, between the years 2001 and 2013. These components were subjected to qualification, screening and testing in which the goal was to ensure that the components would survive the stresses of the mission. The analysis categorizes failures by part type and failure mechanisms. One of the results of the analysis was the realization that a surprising proportion of failures experienced during system integration and testing were caused by human error (i.e. human induced defect). Further analysis included the determination of root failure mechanisms and any influencing factors contributing to these failures. The major causes of these defects were attributed to electrostatic damage (ESD), electrical overstress (EOS), mechanical overstress (MOS), and thermal overstress (TOS). Finally, the study proposes a risk analysis tool which incorporates these major causes for the failures, termed error-producing conditions (EPCs), and a proportionality factor representing the number of each type of failure that has occurred at the facility under study. These factors are quantified and used to communicate the risk of human induced defects for the assembly, integration and testing of space hardware based on the system’s electrical parts list. The new risk identification can trigger risk-mitigating actions more effectively, based on the presence of component categories or other hazardous conditions that have a history of failure due to human error.

Majewicz, Peter J.

Command Process Modeling & Risk Analysis

Commanding Errors may be caused by a variety of root causes. It's important to understand the relative significance of each of these causes for making institutional investment decisions. One of these causes is the lack of standardized processes and procedures for command and control. We mitigate this problem by building periodic tables and models corresponding to key functions within it. These models include simulation analysis and probabilistic risk assessment models.

functional analysis

Materials Analysis: A Key to Unlocking the Mystery of the Columbia Tragedy

Materials analyses of key forensic evidence helped unlock the mystery of the loss of space shuttle Columbia that disintegrated February 1, 2003 while returning from a 16-day research mission. Following an intensive four-month recovery effort by federal, state, and local emergency management and law officials, Columbia debris was collected, catalogued, and reassembled at the Kennedy Space Center. Engineers and scientists from the Materials and Processes (M&P) team formed by NASA supported Columbia reconstruction efforts, provided factual data through analysis, and conducted experiments to validate the root cause of the accident. Fracture surfaces and thermal effects of selected airframe debris were assessed, and process flows for both nondestructive and destructive sampling and evaluation of debris were developed. The team also assessed left hand (LH) airframe components that were believed to be associated with a structural breach of Columbia. Analytical data collected by the M&P team showed that a significant thermal event occurred at the left wing leading edge in the proximity of LH reinforced carbon carbon (RCC) panels 8 and 9. The analysis also showed exposure to temperatures in excess of 1,649 C, which would severely degrade the support structure, tiles, and RCC panel materials. The integrated failure analysis of wing leading edge debris and deposits strongly supported the hypothesis that a breach occurred at LH RCC panel 8.

Mayeaux, Brian M.

Examining the Feasibility of Detailed Decomposition of Radiative Flux Anomalies to Cloud Changes

This presentation will explore the degree to which radiative flux anomalies in the last two decades seen by CERES can be interpreted as the result of changes in particular cloud types. A study of this kind is potentially enabled by the new CERES FlxByCldTyp (FBCT) product which provides combined Terra and Aqua daytime 1°-regional gridded daily and monthly top-of-the-atmosphere radiative fluxes and associated MODIS-derived cloud properties stratified by Cloud Top Pressure (CTP) and Cloud Optical Thickness (TAU), i.e., arrays of cloud properties and radiative fluxes resolved in CTP-TAU bins. It has been shown (Sun et al. 2022) that the FBCT time series of global flux anomalies tracks closely its SSF1deg and EBAF counterparts, indicating that these three products provide fluxes of consistent stability for interannual variability studies. We are thus in position to examine whether CERES regional or global radiative fluxes anomalies and trends are driven by changes in certain elements of FBCT flux arrays corresponding to particular cloud types. Anomalies in flux array elements will in turn be investigated in terms of cloud property variations within the CTP-TAU bins where the biggest flux anomalies are encountered. We will essentially seek to obtain the relative contributions to the anomalies of binned all-sky fluxes of changes in cloud fraction within cloud types and of changes in their overcast fluxes modulated by other cloud properties. Such an analysis holds great promise of unveiling the root causes of cloud-driven changes in the planetary radiation budget of the last two decades.

cloud

Late-Notice HIE Investigation

Provide a response to MOWG action item 1410-01: Analyze close approaches which have required mission team action on short notice. Determine why the approaches were identified later in the process than most other events. Method: Performed an analysis to determine whether there is any correlation between late notice event identification and space weather, sparse tracking, or high drag objects, which would allow preventive action to be taken Examined specific late notice events identified by missions as problematic to try to identify root cause and attempt to relate them to the correlation analysis.

HIE

Review of Orbiter Flight Boundary Layer Transition Data

In support of the Shuttle Return to Flight program, a tool was developed to predict when boundary layer transition would occur on the lower surface of the orbiter during reentry due to the presence of protuberances and cavities in the thermal protection system. This predictive tool was developed based on extensive wind tunnel tests conducted after the loss of the Space Shuttle Columbia. Recognizing that wind tunnels cannot simulate the exact conditions an orbiter encounters as it re-enters the atmosphere, a preliminary attempt was made to use the documented flight related damage and the orbiter transition times, as deduced from flight instrumentation, to calibrate the predictive tool. After flight STS-114, the Boundary Layer Transition Team decided that a more in-depth analysis of the historical flight data was needed to better determine the root causes of the occasional early transition times of some of the past shuttle flights. In this paper we discuss our methodology for the analysis, the various sources of shuttle damage information, the analysis of the flight thermocouple data, and how the results compare to the Boundary Layer Transition prediction tool designed for Return to Flight.

Mcginley, Catherine B.

ISHM Decision Analysis Tool: Operations Concept

The state-of-the-practice Shuttle caution and warning system warns the crew of conditions that may create a hazard to orbiter operations and/or crew. Depending on the severity of the alarm, the crew is alerted with a combination of sirens, tones, annunciator lights, or fault messages. The combination of anomalies (and hence alarms) indicates the problem. Even with much training, determining what problem a particular combination represents is not trivial. In many situations, an automated diagnosis system can help the crew more easily determine an underlying root cause. Due to limitations of diagnosis systems,however, it is not always possible to explain a set of alarms with a single root cause. Rather, the system generates a set of hypotheses that the crew can select from. The ISHM Decision Analysis Tool (IDAT) assists with this task. It presents the crew relevant information that could help them resolve the ambiguity of multiple root causes and determine a method for mitigating the problem. IDAT follows graphical user interface design guidelines and incorporates a decision analysis system. I describe both of these aspects.

Source record

Dynamic tooth loads and stressing for high contact ratio spur gears

An analysis and computer program were developed for calculating the dynamic gear tooth loading and root stressing for high contact ratio gearing (HCRG) as well as LCRG. The analysis includes the effects of the variable tooth stiffness during the mesh, tooth profile modification, and gear errors. The calculation of the tooth root stressing caused by the dynamic gear tooth loads is based on a modified Heywood gear tooth stress analysis, which appears more universally applicable to both LCRG and HCRG. The computer program is presently being expanded to calculate the tooth contact stressing and PV values. Sample application of the gear program to equivalent LCRG (1.566 contact ratio) and HCRG (2.40 contact ratio) revealed the following: (1) the operating conditions and dynamic characteristics of the gear system an affect the gear tooth loading and root stressing, and therefore, life significantly; (2) the length of the profile modification affect the tooth loading and root stressing significantly, the amount depending on the applied load, speed, and contact ratio; and (3) the effect of variable tooth stiffness is small, shifting and increasing the response peaks slightly from those for constant tooth stiffness.

Cornell, R. W.

Business Intelligence Modeling in Launch Operations

This technology project is to advance an integrated Planning and Management Simulation Model for evaluation of risks, costs, and reliability of launch systems from Earth to Orbit for Space Exploration. The approach builds on research done in the NASA ARC/KSC developed Virtual Test Bed (VTB) to integrate architectural, operations process, and mission simulations for the purpose of evaluating enterprise level strategies to reduce cost, improve systems operability, and reduce mission risks. The objectives are to understand the interdependency of architecture and process on recurring launch cost of operations, provide management a tool for assessing systems safety and dependability versus cost, and leverage lessons learned and empirical models from Shuttle and International Space Station to validate models applied to Exploration. The systems-of-systems concept is built to balance the conflicting objectives of safety, reliability, and process strategy in order to achieve long term sustainability. A planning and analysis test bed is needed for evaluation of enterprise level options and strategies for transit and launch systems as well as surface and orbital systems. This environment can also support agency simulation .based acquisition process objectives. The technology development approach is based on the collaborative effort set forth in the VTB's integrating operations. process models, systems and environment models, and cost models as a comprehensive disciplined enterprise analysis environment. Significant emphasis is being placed on adapting root cause from existing Shuttle operations to exploration. Technical challenges include cost model validation, integration of parametric models with discrete event process and systems simulations. and large-scale simulation integration. The enterprise architecture is required for coherent integration of systems models. It will also require a plan for evolution over the life of the program. The proposed technology will produce long-term benefits in support of the NASA objectives for simulation based acquisition, will improve the ability to assess architectural options verses safety/risk for future exploration systems, and will facilitate incorporation of operability as a systems design consideration, reducing overall life cycle cost for future systems. The future of business intelligence of space exploration will focus on the intelligent system-of-systems real-time enterprise. In present business intelligence, a number of technologies that are most relevant to space exploration are experiencing the greatest change. Emerging patterns of set of processes rather than organizational units leading to end-to-end automation is becoming a major objective of enterprise information technology. The cost element is a leading factor of future exploration systems.

Bardina, Jorge E.