Root Cause Analysis of Data Refinement Process – Medical Conditions Capability Resource Tables
- Objective - Background - Capability Resource Tables (CRT) - Root Cause Analysis(RCA) - Error types - RCA - Fishbone; 5 whys - Lessons Learned
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
- Objective - Background - Capability Resource Tables (CRT) - Root Cause Analysis(RCA) - Error types - RCA - Fishbone; 5 whys - Lessons Learned
Explore the source record for details and available documents.
The medical system for spaceflight thus far has been designed to support missions in low earth orbit (LEO). Crew capabilities are limited and heavily dependent on the team of medical support staff at Mission Control Center (MCC) to guide diagnosis and management. However, missions to the Moon and Mars will suffer from several constraints that will make this ground support focused approach to care ineffective. In order to update and modify medical system design, NASA has relied on Probabilistic Risk Assessment (PRA) modeling to mitigate medical risk through trade space analysis. Specifically, capability resource tables (CRT’s) were developed to create a dataset of resources required to manage a list of accepted medical conditions significant in exploration spaceflight. With 120 conditions, this dataset contained hundreds of capabilities and thousands of resources with tens of thousands of cells of data. Initially these tables were built in excel for high throughput during development, but ultimately had to be transferred, managed, and modified into the Evidence Library database for modeling purposes. The process of collating and reviewing the Evidence Library revealed numerous errors in the dataset that had to be corrected through iterative changes. Several error types emerged during this process and can be broken into specific classifications defined as “input”, “transcription”, “structural”, “branching”, and “information”. In reviewing these error types through the root cause analysis (RCA) approach, we were able to identify the contributors to these errors which included single data review points, changing product end goals, limited software selection, time constraints and several others. By reviewing and evaluating the underlying causes we can provide possible system improvements that can be implemented for current and future data management in PRA model inputs.
The NASA's Evolutionary Xenon Thruster (NEXT) project is developing an advanced ion propulsion system for future NASA missions for solar system exploration. A critical element of the propulsion system is the Power Processing Unit (PPU) which supplies regulated power to the key components of the thruster. The PPU contains six different power supplies including the beam, discharge, discharge heater, neutralizer, neutralizer heater, and accelerator supplies. The beam supply is the largest and processes up to 93+% of the power. The NEXT PPU had been operated for approximately 200+ hours and has experienced a series of three capacitor failures in the beam supply. The capacitors are in the same, nominally non-critical location the input filter capacitor to a full wave switching inverter. The three failures occurred after about 20, 30, and 135 hours of operation. This paper provides background on the NEXT PPU and the capacitor failures. It discusses the failure investigation approach, the beam supply power switching topology and its operating modes, capacitor characteristics and circuit testing. Finally, it identifies root cause of the failures to be the unusual confluence of circuit switching frequency, the physical layout of the power circuits, and the characteristics of the capacitor.
The NASA s Evolutionary Xenon Thruster (NEXT) project is developing an advanced ion propulsion system for future NASA missions for solar system exploration. A critical element of the propulsion system is the Power Processing Unit (PPU) which supplies regulated power to the key components of the thruster. The PPU contains six different power supplies including the beam, discharge, discharge heater, neutralizer, neutralizer heater, and accelerator supplies. The beam supply is the largest and processes up to 93+% of the power. The NEXT PPU had been operated for approximately 200+ hr and has experienced a series of three capacitor failures in the beam supply. The capacitors are in the same, nominally non-critical location-the input filter capacitor to a full wave switching inverter. The three failures occurred after about 20, 30, and 135 hr of operation. This paper provides background on the NEXT PPU and the capacitor failures. It discusses the failure investigation approach, the beam supply power switching topology and its operating modes, capacitor characteristics and circuit testing. Finally, it identifies root cause of the failures to be the unusual confluence of circuit switching frequency, the physical layout of the power circuits, and the characteristics of the capacitor.
A gradual, but persistent, decrease in the optical throughput was detected during the early commissioning phase for the Suomi National Polar-Orbiting Partnership (SNPP) Visible Infrared Imager Radiometer Suite (VIIRS) Near Infrared (NIR) bands. Its initial rate and unknown cause were coincidently coupled with a decrease in sensitivity in the same spectral wavelength of the Solar Diffuser Stability Monitor (SDSM) raising concerns about contamination or the possibility of a system-level satellite problem. An anomaly team was formed to investigate and provide recommendations before commissioning could resume. With few hard facts in hand, there was much speculation about possible causes and consequences of the degradation. Two different causes were determined as will be explained in this paper. This paper will describe the build and test history of VIIRS, why there were no indicators, even with hindsight, of an on-orbit problem, the appearance of the on-orbit anomaly, the initial work attempting to understand and determine the cause, the discovery of the root cause and what Test-As-You-Fly (TAYF) activities, can be done in the future to greatly reduce the likelihood of similar optical anomalies. These TAYF activities are captured in the lessons learned section of this paper.
This paper discusses the performance analysis effort being carried out in the NASA Deep Space Network. The activity involves root cause analysis of failures and assessment of key performance metrics. The root cause analysis helps pinpoint the true cause of observed problems so that proper correction can be effected. The assessment currently focuses on three aspects: (1) data delivery metrics such as Quantity, Quality, Continuity, and Latency; (2) link-performance metrics such as antenna pointing, system noise temperature, Doppler noise, frequency and time synchronization, wide-area-network loading, link-configuration setup time; and (3) reliability, maintainability, availability metrics. The analysis establishes whether the current system is meeting its specifications and if so, how much margin is available. The findings help identify the weak points in the system and direct attention of programmatic investment for performance improvement.
This two day workshop will discuss a range of topics, including root cause analysis, physics-of-failure principles and failure mechanisms in printed circuit boards. Printed circuit boards (PCBs) are the baseline for electronics manufacturing upon which electronic components are mounted and formed into electronic systems. PCBs are used in a variety of electronic circuits from simple one-transistor amplifiers to large super computers. A PCB serves three main functions: 1) it provides the necessary mechanical support for the components in the circuit 2) it provides the necessary electrical interconnections, and 3) it bears some form of legend which identifies the components it carries. The failure modes on the PCBs can be categorized in a hierarchical structure, in which the mechanisms and causes are site or location dependant. Specimen preparation techniques, non-destructive and destructive analysis, and materials characterization will also be discussed. The first day of the workshop will present methodologies for identifying potential failure mechanisms in electronics based on the failure history and, systematic approaches to root cause analysis. The second day will cover failure analysis techniques geared towards various failure mechanisms, along with numerous component and PCB assembly failure analysis case studies that illustrate the techniques and analysis. Failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies.
Mission Assurance independent assessments started during the development cycle and continued through post launch operations. In operations, Health and Safety of the Observatory is of utmost importance. Therefore, Mission Assurance must ensure requirements compliance and focus on process improvements required across the operational systems including new/modified products, tools, and procedures. The deployment of the interactive model involves three objectives: Team member Interaction, Good Root Cause Analysis Practices, and Risk Assessment to avoid reoccurrences. In applying this model, we use a metric based measurement process and was found to have the most significant effect, which points to the importance of focuses on a combination of root cause analysis and risk approaches allowing the engineers the ability to prioritize and quantify their corrective actions based on a well-defined set of root cause definitions (i.e. closure criteria for problem reports), success criteria and risk rating definitions.
ISHM capability enables a system to detect anomalies, determine causes and effects, predict future anomalies, and provides an integrated awareness of the health of the system to users (operators, customers, management, etc.). NASA Stennis Space Center, NASA Ames Research Center, and Pratt & Whitney Rocketdyne have implemented a core ISHM capability that encompasses the A1 Test Stand and the J-2X Engine. The implementation incorporates all aspects of ISHM; from anomaly detection (e.g. leaks) to root-cause-analysis based on failure mode and effects analysis (FMEA), to a user interface for an integrated visualization of the health of the system (Test Stand and Engine). The implementation provides a low functional capability level (FCL) in that it is populated with few algorithms and approaches for anomaly detection, and root-cause trees from a limited FMEA effort. However, it is a demonstration of a credible ISHM capability, and it is inherently designed for continuous and systematic augmentation of the capability. The ISHM capability is grounded on an integrating software environment used to create an ISHM model of the system. The ISHM model follows an object-oriented approach: includes all elements of the system (from schematics) and provides for compartmentalized storage of information associated with each element. For instance, a sensor object contains a transducer electronic data sheet (TEDS) with information that might be used by algorithms and approaches for anomaly detection, diagnostics, etc. Similarly, a component, such as a tank, contains a Component Electronic Data Sheet (CEDS). Each element also includes a Health Electronic Data Sheet (HEDS) that contains health-related information such as anomalies and health state. Some practical aspects of the implementation include: (1) near real-time data flow from the test stand data acquisition system through the ISHM model, for near real-time detection of anomalies and diagnostics, (2) insertion of the J-2X predictive model providing predicted sensor values for comparison with measured values and use in anomaly detection and diagnostics, and (3) insertion of third-party anomaly detection algorithms into the integrated ISHM model.
The Space Launch System (SLS) Core Stage (CS) Thrust Vector Control (TVC) system is comprised of eight mechanical feedback Shuttle heritage Type III TVC actuators and four RS-25 engines, each attached to a Shuttle heritage gimbal block/bearing. Two actuators are used to move each engine in two planes perpendicular to one another (i.e., pitch and yaw). The TVC system design leverages hardware from the Space Shuttle program as well as new hardware designed specifically for the Core Stage. The Green Run Hot Fire (GRHF) of the SLS Core Stage provided a flight-like ground test environment for verification of integrated vehicle TVC performance. A TVC model coupled to a vehicle structural dynamic model has been developed previously and incrementally validated in subsystem tests and simulations. Still, some aspects of TVC performance in GRHF were not anticipated. The ensuing investigation demonstrated the need for well-instrumented test environments, various levels of modeling fidelity, test-representative structural models, and caution in reuse of legacy components. This paper is the sixth installment in a seven-paper series surveying the design, engineering, test validation, and flight performance of the Core Stage Thrust Vector Control system. It introduces the salient structural dynamic phenomena uncovered in ambient and hot fire testing. During the Green Run test campaign, a comparison of ambient and hot fire step responses showed a significant change in apparent damping due to the presence of friction, challenging long standing assumptions that friction could be neglected. Additionally, the characteristic response of the engine and thrust structure during GRHF proved to be more complex than anticipated, as evidenced by the available actuator, thrust structure, and engine measurements. While the string-potentiometer based test instrumentation was intended to allow for reconstruction of the engine angles along the two control axes, the geometric placement, location uncertainty, and responses in overlapping frequency spectra revealed additional phenomena requiring further analysis and post-processing. The observations from both modal and frequency response testing during the Green Run ambient and hot fire configurations led to Engine and Core Stage FEM (finite element model) updates. When evidence of unexpected engine motion was found in engine section accelerometer data, the authors pursued additional structural analysis leading to FEM updates associated with the TVC gimbal and thrust structure. Through collaboration between structures, TVC, and flight control disciplines, the test-informed models and root-cause analysis led to confident flight rationale for the first flight of the SLS launch vehicle.
The Space Launch System (SLS) Core Stage (CS) Thrust Vector Control (TVC) system is comprised of eight mechanical feedback Shuttle heritage Type III TVC actuators and four RS-25 engines, each attached to a Shuttle heritage gimbal block/bearing. Two actuators are used to move each engine in two planes perpendicular to one another (i.e., pitch and yaw). The TVC system design leverages hardware from the Space Shuttle program as well as new hardware designed specifically for the Core Stage. The Green Run Hot Fire (GRHF) of the SLS Core Stage provided a flight-like ground test environment for verification of integrated vehicle TVC performance. A TVC model coupled to a vehicle structural dynamic model has been developed previously and incrementally validated in subsystem tests and simulations. Still, some aspects of TVC performance in GRHF were not anticipated. The ensuing investigation demonstrated the need for well-instrumented test environments, various levels of modeling fidelity, test-representative structural models, and caution in reuse of legacy components. This paper is the sixth installment in a seven-paper series surveying the design, engineering, test validation, and flight performance of the Core Stage Thrust Vector Control system. It introduces the salient structural dynamic phenomena uncovered in ambient and hot fire testing. During the Green Run test campaign, a comparison of ambient and hot fire step responses showed a significant change in apparent damping due to the presence of friction, challenging long standing assumptions that friction could be neglected. Additionally, the characteristic response of the engine and thrust structure during GRHF proved to be more complex than anticipated, as evidenced by the available actuator, thrust structure, and engine measurements. While the string-potentiometer based test instrumentation was intended to allow for reconstruction of the engine angles along the two control axes, the geometric placement, location uncertainty, and responses in overlapping frequency spectra revealed additional phenomena requiring further analysis and post-processing. The observations from both modal and frequency response testing during the Green Run ambient and hot fire configurations led to Engine and Core Stage FEM (finite element model) updates. When evidence of unexpected engine motion was found in engine section accelerometer data, the authors pursued additional structural analysis leading to FEM updates associated with the TVC gimbal and thrust structure. Through collaboration between structures, TVC, and flight control disciplines, the test-informed models and root-cause analysis led to confident flight rationale for the first flight of the SLS launch vehicle.
ISHM capability includes: detection of anomalies, diagnosis of causes of anomalies, prediction of future anomalies, and user interfaces that enable integrated awareness (past, present, and future) by users. This is achieved by focused management of data, information and knowledge (DIaK) that will likely be distributed across networks. Management of DIaK implies storage, sharing (timely availability), maintaining, evolving, and processing. Processing of DIaK encapsulates strategies, methodologies, algorithms, etc. focused on achieving high ISHM Functional Capability Level (FCL). High FCL means a high degree of success in detecting anomalies, diagnosing causes, predicting future anomalies, and enabling health integrated awareness by the user. A model that enables ISHM capability, and hence, DIaK management, is denominated the ISHM Model of the System (IMS). We describe aspects of the IMS that focus on processing of DIaK. Strategies, methodologies, and algorithms require proper context. We describe an approach to define and use contexts, implementation in an object-oriented software environment (G2), and validation using actual test data from a methane thruster test program at NASA SSC. Context is linked to existence of relationships among elements of a system. For example, the context to use a strategy to detect leak is to identify closed subsystems (e.g. bounded by closed valves and by tanks) that include pressure sensors, and check if the pressure is changing. We call these subsystems Pressurizable Subsystems. If pressure changes are detected, then all members of the closed subsystem become suspect of leakage. In this case, the context is defined by identifying a subsystem that is suitable for applying a strategy. Contexts are defined in many ways. Often, a context is defined by relationships of function (e.g. liquid flow, maintaining pressure, etc.), form (e.g. part of the same component, connected to other components, etc.), or space (e.g. physically close, touching the same common element, etc.). The context might be defined dynamically (if conditions for the context appear and disappear dynamically) or statically. Although this approach is akin to case-based reasoning, we are implementing it using a software environment that embodies tools to define and manage relationships (of any nature) among objects in a very intuitive manner. Context for higher level inferences (that use detected anomalies or events), primarily for diagnosis and prognosis, are related to causal relationships. This is useful to develop root-cause analysis trees showing an event linked to its possible causes and effects. The innovation pertaining to RCA trees encompasses use of previously defined subsystems as well as individual elements in the tree. This approach allows more powerful implementations of RCA capability in object-oriented environments. For example, if a pressurizable subsystem is leaking, its root-cause representation within an RCA tree will show that the cause is that all elements of that subsystem are suspect of leak. Such a tree would apply to all instances of leak-events detected and all elements in all pressurizable subsystems in the system. Example subsystems in our environment to build IMS include: Pressurizable Subsystem, Fluid-Fill Subsystem, Flow-Thru-Valve Subsystem, and Fluid Supply Subsystem. The software environment for IMS is designed to potentially allow definition of any relationship suitable to create a context to achieve ISHM capability.
Root Source Analysis (RoSA) is a systems engineering methodology that has been developed at NASA over the past five years. It is designed to reduce costs, schedule, and technical risks by systematically examining critical assumptions and the state of the knowledge needed to bring to fruition the products that satisfy mission-driven requirements, as defined for each element of the Work (or Product) Breakdown Structure (WBS or PBS). This methodology is sometimes referred to as the ValuStream method, as inherent in the process is the linking and prioritizing of uncertainties arising from knowledge shortfalls directly to the customer's mission driven requirements. RoSA and ValuStream are synonymous terms. RoSA is not simply an alternate or improved method for identifying risks. It represents a paradigm shift. The emphasis is placed on identifying very specific knowledge shortfalls and assumptions that are the root sources of the risk (the why), rather than on assessing the WBS product(s) themselves (the what). In so doing RoSA looks forward to anticipate, identify, and prioritize knowledge shortfalls and assumptions that are likely to create significant uncertainties/ risks (as compared to Root Cause Analysis, which is most often used to look back to discover what was not known, or was assumed, that caused the failure). Experience indicates that RoSA, with its primary focus on assumptions and the state of the underlying knowledge needed to define, design, build, verify, and operate the products, can identify critical risks that historically have been missed by the usual approaches (i.e., design review process and classical risk identification methods). Further, the methodology answers four critical questions for decision makers and risk managers: 1. What s been included? 2. What's been left out? 3. How has it been validated? 4. Has the real source of the uncertainty/ risk been identified, i.e., is the perceived problem the real problem? Users of the RoSA methodology have characterized it as a true bottoms up risk assessment.
The purpose of the Sentiment of Search Study for NASA Johnson Space Center (JSC) is to gain insight into the intranet search environment. With an initial usability survey, the authors were able to determine a usability score based on the Systems Usability Scale (SUS). Created in 1986, the freely available, well cited, SUS is commonly used to determine user perceptions of a system (in this case the intranet search environment). As with any improvement initiative, one must first examine and document the current reality of the situation. In this scenario, a method was needed to determine the usability of a search interface in addition to the user's perception on how well the search system was providing results. The use of the SUS provided a mechanism to quickly ascertain information in both areas, by adding one additional open-ended question at the end. The first ten questions allowed us to examine the usability of the system, while the last questions informed us on how the users rated the performance of the search results. The final analysis provides us with a better understanding of the current situation and areas to focus on for improvement. The power of search applications to enhance knowledge transfer is indisputable. The performance impact for any user unable to find needed information undermines project lifecycle, resource and scheduling requirements. Ever-increasing complexity of content and the user interface make usability considerations for the intranet, especially for search, a necessity instead of a 'nice-to-have'. Despite these arguments, intranet usability is largely disregarded due to lack of attention beyond the functionality of the infrastructure (White, 2013). The data collected from users of the JSC search system revealed their overall sentiment by means of the widely-known System Usability Scale. Results of the scores suggest 75%, +/-0.04, of the population rank the search system below average. In terms of a grading scaled, this equated to D or lower. It is obvious JSC users are not satisfied with the current situation, however they are eager to provide information and assistance in improving the search system. A majority of the respondents provided feedback on the issues most troubling them. This information will be used to enrich the next phase, root cause analysis and solution creation.
NASA's need to trace mistakes to their source to try and eliminate them in the future has resulted in software known as Root Cause Analysis (RoCA). Fair, Isaac & Co., Inc. has applied RoCA software, originally developed under an SBIR contract with Kennedy, to its predictive software technology. RoCA can generate graphic reports to make analysis of problems easier and more efficient.
All of the International Space Station (ISS) systems which require computer control depend upon the hardware and software of the Command and Data Handling System (C&DH) system, currently a network of over 30 386-class computers called Multiplexor/Dimultiplexors (MDMs)[18]. The Caution and Warning System (C&W)[7], a set of software tasks that runs on the MDMs, is responsible for detecting, classifying, and reporting errors in all ISS subsystems including the C&DH. Fault Detection, Isolation and Recovery (FDIR) of these errors is typically handled with a combination of automatic and human effort. We are developing an Advanced Diagnostic System (ADS) to augment the C&W system with decision support tools to aid in root cause analysis as well as resolve differing human and machine C&DH state estimates. These tools which draw from sources in model-based reasoning[ 16,291, will improve the speed and accuracy of flight controllers by reducing the uncertainty in C&DH state estimation, allowing for a more complete assessment of risk. We have run tests with ISS telemetry and focus on those C&W events which relate to the C&DH system itself. This paper describes our initial results and subsequent plans.
There are a number of architecture models for implementing Integrated Systems Health Management (ISHM) capabilities. For example, approaches based on the OSA-CBM and OSA-EAI models, or specific architectures developed in response to local needs. NASA s John C. Stennis Space Center (SSC) has developed one such version of an extensible architecture in support of rocket engine testing that integrates a palette of functions in order to achieve an ISHM capability. Among the functional capabilities that are supported by the framework are: prognostic models, anomaly detection, a data base of supporting health information, root cause analysis, intelligent elements, and integrated awareness. This paper focuses on the role that intelligent elements can play in ISHM architectures. We define an intelligent element as a smart element with sufficient computing capacity to support anomaly detection or other algorithms in support of ISHM functions. A smart element has the capabilities of supporting networked implementations of IEEE 1451.x smart sensor and actuator protocols. The ISHM group at SSC has been actively developing intelligent elements in conjunction with several partners at other Centers, universities, and companies as part of our ISHM approach for better supporting rocket engine testing. We have developed several implementations. Among the key features for these intelligent sensors is support for IEEE 1451.1 and incorporation of a suite of algorithms for determination of sensor health. Regardless of the potential advantages that can be achieved using intelligent sensors, existing large-scale systems are still based on conventional sensors and data acquisition systems. In order to bring the benefits of intelligent sensors to these environments, we have also developed virtual implementations of intelligent sensors.