Search NASA⌕ Search

SEARCH · Search NASA

Results for “Fault detection and characterization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

95 records · Page 6

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASA's Space Launch System

The engineering development of the new Space Launch System (SLS) launch vehicle requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The characteristics of these spacecraft systems must be matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex system engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in specialized Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model based algorithms and their development lifecycle from inception through Flight Software certification are an important focus of this development effort to further insure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. NASA formed a dedicated M&FM team for addressing fault management early in the development lifecycle for the SLS initiative. As part of the development of the M&FM capabilities, this team has developed a dedicated testbed that integrates specific M&FM algorithms, specialized nominal and off-nominal test cases, and vendor-supplied physics-based launch vehicle subsystem models. Additionally, the team has developed processes for implementing and validating these algorithms for concept validation and risk reduction for the SLS program. The flexibility of the Vehicle Management End-to-end Testbed (VMET) enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS. The intent of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software development infrastructure and its related testing entities. In any software development process there is inherent risk in the interpretation and implementation of concepts into software through requirements and test cases into flight software compounded with potential human errors throughout the development lifecycle. Risk reduction is addressed by the M&FM analysis group working with other organizations such as S&MA, Structures and Environments, GNC, Orion, the Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission and Loss of Crew probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses that can be tested in VMET to ensure that failures can be detected, and confirm that responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - ARINC 653 partitioned OS, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by Flight Software. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure the effectiveness of M&FM algorithms performance in the FSW development and test processes.

Trevino, Luis↗

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASAs Space Launch System

The engineering development of the National Aeronautics and Space Administration's (NASA) new Space Launch System (SLS) requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The nominal and off-nominal characteristics of SLS's elements and subsystems must be understood and matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex systems engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model-based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model-based algorithms and their development lifecycle from inception through FSW certification are an important focus of SLS's development effort to further ensure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. To test and validate these M&FM algorithms a dedicated test-bed was developed for full Vehicle Management End-to-End Testing (VMET). For addressing fault management (FM) early in the development lifecycle for the SLS program, NASA formed the M&FM team as part of the Integrated Systems Health Management and Automation Branch under the Spacecraft Vehicle Systems Department at the Marshall Space Flight Center (MSFC). To support the development of the FM algorithms, the VMET developed by the M&FM team provides the ability to integrate the algorithms, perform test cases, and integrate vendor-supplied physics-based launch vehicle (LV) subsystem models. Additionally, the team has developed processes for implementing and validating the M&FM algorithms for concept validation and risk reduction. The flexibility of the VMET capabilities enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS, GNC, and others. One of the principal functions of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software test and validation processes. In any software development process there is inherent risk in the interpretation and implementation of concepts from requirements and test cases into flight software compounded with potential human errors throughout the development and regression testing lifecycle. Risk reduction is addressed by the M&FM group but in particular by the Analysis Team working with other organizations such as S&MA, Structures and Environments, GNC, Orion, Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission (LOM) and Loss of Crew (LOC) probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses to be tested in VMET to ensure reliable failure detection, and confirm responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - the ARINC 6535-partitioned Operating System, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by FSW. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure their effectiveness and performance in the exterior FSW development and test processes. This paper is outlined in a systematic fashion analogous to a lifecycle process flow for engineering development of algorithms into software and testing. Section I describes the NASA SLS M&FM context, presenting the current infrastructure, leading principles, methods, and participants. Section II defines the testing philosophy of the M&FM algorithms as related to VMET followed by section III, which presents the modeling methods of the algorithms to be tested and validated in VMET. Its details are then further presented in section IV followed by Section V presenting integration, test status, and state analysis. Finally, section VI addresses the summary and forward directions followed by the appendices presenting relevant information on terminology and documentation.

Trevino, Luis↗

Understanding software faults and their role in software reliability modeling

This study is a direct result of an on-going project to model the reliability of a large real-time control avionics system. In previous modeling efforts with this system, hardware reliability models were applied in modeling the reliability behavior of this system. In an attempt to enhance the performance of the adapted reliability models, certain software attributes were introduced in these models to control for differences between programs and also sequential executions of the same program. As the basic nature of the software attributes that affect software reliability become better understood in the modeling process, this information begins to have important implications on the software development process. A significant problem arises when raw attribute measures are to be used in statistical models as predictors, for example, of measures of software quality. This is because many of the metrics are highly correlated. Consider the two attributes: lines of code, LOC, and number of program statements, Stmts. In this case, it is quite obvious that a program with a high value of LOC probably will also have a relatively high value of Stmts. In the case of low level languages, such as assembly language programs, there might be a one-to-one relationship between the statement count and the lines of code. When there is a complete absence of linear relationship among the metrics, they are said to be orthogonal or uncorrelated. Usually the lack of orthogonality is not serious enough to affect a statistical analysis. However, for the purposes of some statistical analysis such as multiple regression, the software metrics are so strongly interrelated that the regression results may be ambiguous and possibly even misleading. Typically, it is difficult to estimate the unique effects of individual software metrics in the regression equation. The estimated values of the coefficients are very sensitive to slight changes in the data and to the addition or deletion of variables in the regression equation. Since most of the existing metrics have common elements and are linear combinations of these common elements, it seems reasonable to investigate the structure of the underlying common factors or components that make up the raw metrics. The technique we have chosen to use to explore this structure is a procedure called principal components analysis. Principal components analysis is a decomposition technique that may be used to detect and analyze collinearity in software metrics. When confronted with a large number of metrics measuring a single construct, it may be desirable to represent the set by some smaller number of variables that convey all, or most, of the information in the original set. Principal components are linear transformations of a set of random variables that summarize the information contained in the variables. The transformations are chosen so that the first component accounts for the maximal amount of variation of the measures of any possible linear transform; the second component accounts for the maximal amount of residual variation; and so on. The principal components are constructed so that they represent transformed scores on dimensions that are orthogonal. Through the use of principal components analysis, it is possible to have a set of highly related software attributes mapped into a small number of uncorrelated attribute domains. This definitively solves the problem of multi-collinearity in subsequent regression analysis. There are many software metrics in the literature, but principal component analysis reveals that there are few distinct sources of variation, i.e. dimensions, in this set of metrics. It would appear perfectly reasonable to characterize the measurable attributes of a program with a simple function of a small number of orthogonal metrics each of which represents a distinct software attribute domain.

Munson, John C.↗

Supporting Crew Autonomy in Deep Space Exploration: Preliminary Onboard Capability Requirements and Proposed Research Questions. Technical Report of the Autonomous Crew Operations Technical Interchange Meeting

Communication delays are a critical challenge posed by long duration deep space exploration. Space missions historically have relied on an ever-present Mission Control Center (MCC) to direct operations in near real-time. As unanticipated anomalies that defeat fault detection and resolution systems do arise, the lack of real-time communication will significantly weaken what the MCC support represents: a reliable safety net for the flight crew through its deep and diverse areas of expertise and investigative resources. As a consequence, future space vehicles and habitats need to be equipped with capabilities to support the flight crew to operate with little or no ground support. Considerations must be given to vehicle and mission designs that will fortify the traditionally ground-centered safety net and forge new support systems, when communication delays exist. In August 2018, NASA’s Human Research Program, through its Human Factors and Behavioral Performance Element, convened a Technical Interchange Meeting (TIM) on Autonomous Crew Operations at NASA Ames Research Center. The goal of the meeting was to gather input from NASA centers, industry, academia, and branches of the Department of Defense (DoD) to address how intelligent technologies can be applied to augment onboard capabilities to support crew anomaly response. The TIM featured 24 presentations by 29 speakers and hosted a total of 59 attendees, including 43 from 5 NASA centers (Ames, Johnson, Langley, Marshall, and Jet Propulsion Lab) and 4 from the DoD (3 from Army Research Lab and 1 from Naval Postgraduate School), with remaining attendees from academia (e.g., UC Davis, CMU) and industry (e.g., IBM, Siemens). Discussions were centered around three themes: standards and guidelines, lessons learned in analog environments, and technologies. To help provide a framework for discussion, a concept matrix describing anomaly response processes was created prior to the TIM (Figure 1, page 6). The matrix captures the steps involved (monitoring and detection, diagnosis, solution development and evaluation, solution implementation and verification, resolution documentation) as well as the resources and capabilities required to support these steps (data, knowledge, analysis, synthesis, resource management). A wallpaper size printout of the matrix was utilized at the TIM to solicit attendee inputs along the three themes; the activity garnered 108 submissions of ideas. Overall, what emerged from TIM discussions was a picture of mismatch between crew anomaly response needs and support that can be provided by existing intelligent technologies. The needs are broad, spanning multiple steps and processes/resources, with many of which lacking support from existing technologies, such as knowledge management throughout the steps of problem solving (especially in resolution documentation) and manpower management. The solutions provided by existing intelligent technologies are specific to the steps/processes that they are designed to support and constrained to solving only problems similar to those that have occurred before. What is lacking from technologies is typically made up by humans, specifically their complex critical thinking, creative problem solving, and domain expertise. In the end, the TIM highlighted the pressing need to support responses to onboard anomalies during autonomous crew operations, particularly those that have eluded the system tests, inspection, and other assurance processes. Such anomalies can potentially threaten crew and vehicle safety, as well as significantly impact overall operations with additional workload. These fairly rare events are difficult to anticipate and prepare for, given the state-of-the-art in intelligent technologies. This is true even for anomalies that stem from “unknown knowns”—cases in which there is sufficient external information to characterize the problem but the overall pattern fails to be recognized by the problem solver, or in which the internal knowledge needed to solve a problem is held tacitly and potentially accessible by the problem solver but not articulated. It follows that the ability to tackle anomalies lies not only with the availability of relevant information and knowledge but also their accessibility in times of need. To that end, we propose research questions along the following three broad themes: • How intelligent technologies can help make relevant knowledge and information available? • How intelligent technologies can help make relevant knowledge and information accessible? • How intelligent technologies can help support the crew operating as a team in anomaly response processes?

autonomous crew operations↗

Validation of Atmosphere/Ionosphere Signals Associated with Major Earthquakes by Multi-Instrument Space-Borne and Ground Observations

The latest catastrophic earthquake in Japan (March 2011) has renewed interest in the important question of the existence of pre-earthquake anomalous signals related to strong earthquakes. Recent studies have shown that there were precursory atmospheric/ionospheric signals observed in space associated with major earthquakes. The critical question, still widely debated in the scientific community, is whether such ionospheric/atmospheric signals systematically precede large earthquakes. To address this problem we have started to investigate anomalous ionospheric / atmospheric signals occurring prior to large earthquakes. We are studying the Earth's atmospheric electromagnetic environment by developing a multisensor model for monitoring the signals related to active tectonic faulting and earthquake processes. The integrated satellite and terrestrial framework (ISTF) is our method for validation and is based on a joint analysis of several physical and environmental parameters (thermal infrared radiation, electron concentration in the ionosphere, lineament analysis, radon/ion activities, air temperature and seismicity) that were found to be associated with earthquakes. A physical link between these parameters and earthquake processes has been provided by the recent version of Lithosphere-Atmosphere-Ionosphere Coupling (LAIC) model. Our experimental measurements have supported the new theoretical estimates of LAIC hypothesis for an increase in the surface latent heat flux, integrated variability of outgoing long wave radiation (OLR) and anomalous variations of the total electron content (TEC) registered over the epicenters. Some of the major earthquakes are accompanied by an intensification of gas migration to the surface, thermodynamic and hydrodynamic processes of transformation of latent heat into thermal energy and with vertical transport of charged aerosols in the lower atmosphere. These processes lead to the generation of external electric currents in specific regions of the atmosphere and the modifications, by dc electric fields, in the ionosphere-atmosphere electric circuit. We retrospectively analyzed temporal and spatial variations of four different physical parameters (gas/radon counting rate, lineaments change, long-wave radiation transitions and ionospheric electron density/plasma variations) characterizing the state of the lithosphere/atmosphere coupling several days before the onset of the earthquakes. Validation processes consist in two phases: A. Case studies for seven recent major earthquakes: Japan (M9.0, 2011), China (M7.9, 2008), Italy (M6.3, 2009), Samoa (M7, 2009), Haiti (M7.0, 2010) and, Chile (M8.8, 2010) and B. A continuous retrospective analysis was preformed over two different regions with high seismicity- Taiwan and Japan for 2003-2009. Satellite, ground surface, and troposphere data were obtained from Terra/ASTER, Aqua/AIRS, POES and ionospheric variations from DEMETER and COSMIC-I data. Radon and GPS/TEC were obtaining from monitoring sites in Taiwan, Japan and Italy and from global ionosphere maps (GIM) respectively. Our analysis of ground and satellite data during the occurrence of 7 global earthquakes has shown the presence of anomalies in the atmosphere. Our results for Tohoku M9.0 earthquake show that on March 7th, 2011 (4 days before the main shock and 1 day before the M7.2 foreshock of March 8, 2011) a rapid increase of emitted infrared radiation was observed by the satellite data and an anomaly was developed near the epicenter. The GPS/TEC data indicate an increase and variation in electron density reaching a maximum value on March 8. From March 3 to 11 a large increase in electron concentration was recorded at all four Japanese ground-based ionosondes, which returned to normal after the main earthquake. Similar approach for analyzing atmospheric and ionospheric parameters has been applied for China (M7.9, 2008), Italy (M6.3, 2009), Samoa (M7, 2009), Haiti (M7.0, 2010) and Chile (M8.8, 2010) eahquakes. Results have revealed the presence of related variations of these parameters implying their connection with the earthquake process. The second phase (B) of this validation included 102 major earthquakes (M>5.9) in Taiwan and Japan. We have found anomalous behavior before all of these events with no false negatives. False alarm ratio for false positives is less then 10% and has been calculated for the same month of the earthquake occurrence for the entire period of analysis (2003-2009). The commonalities for detecting atmospheric/ionospheric anomalies are: i.) Regularly appearance over regions of maximum stress (i.e., along plate boundaries); ii.) Anomaly existence over land and sea; and iii) association with M>5.9 earthquakes not deeper than 100km. Due to their long duration over the same region these anomalies are not consistent with a meteorological origin. Our initial results from the ISTF validation of multi-instrument space-borne and ground observations show a systematic appearance of atmospheric anomalies near the epicentral area, one to seven (average) days prior to the largest earthquakes, and suggest that it could be explained by a coupling process between the observed physical parameters and the pre-earthquake preparation processes.

Ouzounov, Dimitar↗