Search NASASearch

SEARCH · Search NASA

Results for “fault mitigation methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

49 records · Page 3

Passive Islanding Detection of Inverter-Based Resources in a Noisy Environment

Islanding occurs when a load is energized solely by local generators and can result in frequency and voltage instability, changes in current, and poor power quality. Poor power quality can interrupt industrial operations, damage sensitive electrical equipment, and induce outages upon the resynchronization of the island with the grid. This study proposes an islanding detection method employing a Duffing oscillator to analyze voltage fluctuations at the point of common coupling (PCC) under a high-noise environment. Unlike existing methods, which overlook the noise effect, this paper mitigates noise impact on islanding detection. Power system noise in PCC measurements arises from switching transients, harmonics, grounding issues, voltage sags and swells, electromagnetic interference, and power quality issues that affect islanding detection. Transient events like lightning-induced traveling waves to the PCC can also introduce noise levels exceeding the voltage amplitude by more than seven times, thus disturbing conventional detection techniques. The noise interferes with measurements and increases the nondetection zone (NDZ), causing failed or delayed islanding detection. The Duffing oscillator nonlinear dynamics enable detection capabilities at a high noise level. The proposed method is designed to detect the PCC voltage fluctuations based on the IEEE standard 1547 through the Duffing oscillator. For the voltages beyond the threshold, the Duffing oscillator phase trajectory changes from periodic to chaotic mode and sends an islanded operation command to the inverter. The proposed islanding detection method distinguishes switching transients and faults from an islanded operation. Experimental validation of the method is conducted using a 3.6 kW PV setup.

Energy & Fuels

Impact of Geomagnetic Induced Current Neutral Blocking Devices on Distance Relays in Sub-transmission Networks with IBR

Geomagnetically induced currents (GICs) can flow through transmission lines during geomagnetic disturbances, such as solar flares or coronal mass ejections. These currents can cause problems like transformer saturation and equipment damage. The most common method of mitigating GICs involves installing GIC neutral blocking devices (NBDs) in transformer neutrals. However, the wide application of capacitive GIC blocking devices may have unintended adverse effects on other devices, such as distance protection relays. As the number of inverter-based resources being connected to the transmission and sub-transmission systems increases, the likelihood of a sub-transmission line protected by a distance relay connected to a transformer with NBDs is increasing. Therefore, distance relays fed by IBRs and transmission lines with GIC-NBDs must be studied. This paper studies the effect of GIC-NBDs on a 69kV sub-transmission line of various lengths fed by a 25 MVA IBR and synchronous source. This work focused on the behavior of the GIC-NBDs using the measured apparent phase-to-ground and phase-to-phase impedance calculated by the relay and the source impedance ratio (SIR) during various electrical faults.

Patel, Trupal R [Sandia National Laboratories (SNL

A Sensitivity-driven Wide Area Protection (SWAP) Coordination Tool for High Penetration of Inverter-based Resources (IBR)

Traditionally, power system generation sources have been composed of synchronous generators, of which the fault current behavior is understood with minimal differences between generation size and types due to the physics of their construction. Present protection schemes and modeling methods are based upon these understood characteristics. Most renewable generation is composed of inverter-based resources (IBR), in which fault current is determined by switching control software and hardware limitations, each of which can vary between manufacturers and even between models of the same manufacturer. The resulting fault current is low in magnitude, low in negative-sequence current, unpredictable phase angles, and is a challenge to model. These characteristics also result in a challenge to traditional protection schemes and fault simulation software. To address several of these concerns, the project has the following goals: 1. Improve IBR models: Improve IBR models used in short circuit (SC) programs to accurately capture the response of IBRs at the bulk power system (BPS) level for fault and protection studies. 2. Develop automation tool: Develop an automation tool that allows engineers to identify protection coordination and sensitivity issues by performing SC and protection coordination studies in a high IBR-penetrated grid by applying variations to the IBR models, faults, contingencies, etc. 3. Develop schemes: Develop new protection mitigation solution schemes that complement the existing protection systems to ensure safe operation of the BPS with higher IBR penetration levels. The project team did not achieve this final goal, as the Department of Energy (DOE) stopped the project early due to changes in DOE funding priorities. The termination notice came at the beginning of the final project phase, while the team was identifying and beginning to investigate protection issues. It should be noted that the team discussed a 100% penetration scenario. However, this scenario would require the use of grid-forming IBR models that are not presently available. Since developing these models requires additional effort, the 100% penetration scenario was not pursued during this project. In the future, developing the methodology and models for the 100% scenario could benefit the industry.

14 SOLAR ENERGY

FPGA-Based, Self-Checking, Fault-Tolerant Computers

A proposed computer architecture would exploit the capabilities of commercially available field-programmable gate arrays (FPGAs) to enable computers to detect and recover from bit errors. The main purpose of the proposed architecture is to enable fault-tolerant computing in the presence of single-event upsets (SEUs). [An SEU is a spurious bit flip (also called a soft error) caused by a single impact of ionizing radiation.] The architecture would also enable recovery from some soft errors caused by electrical transients and, to some extent, from intermittent and permanent (hard) errors caused by aging of electronic components. A typical FPGA of the current generation contains one or more complete processor cores, memories, and highspeed serial input/output (I/O) channels, making it possible to shrink a board-level processor node to a single integrated-circuit chip. Custom, highly efficient microcontrollers, general-purpose computers, custom I/O processors, and signal processors can be rapidly and efficiently implemented by use of FPGAs. Unfortunately, FPGAs are susceptible to SEUs. Prior efforts to mitigate the effects of SEUs have yielded solutions that degrade performance of the system and require support from external hardware and software. In comparison with other fault-tolerant- computing architectures (e.g., triple modular redundancy), the proposed architecture could be implemented with less circuitry and lower power demand. Moreover, the fault-tolerant computing functions would require only minimal support from circuitry outside the central processing units (CPUs) of computers, would not require any software support, and would be largely transparent to software and to other computer hardware. There would be two types of modules: a self-checking processor module and a memory system (see figure). The self-checking processor module would be implemented on a single FPGA and would be capable of detecting its own internal errors. It would contain two CPUs executing identical programs in lock step, with comparison of their outputs to detect errors. It would also contain various cache local memory circuits, communication circuits, and configurable special-purpose processors that would use self-checking checkers. (The basic principle of the self-checking checker method is to utilize logic circuitry that generates error signals whenever there is an error in either the checker or the circuit being checked.) The memory system would comprise a main memory and a hardware-controlled check-pointing system (CPS) based on a buffer memory denoted the recovery cache. The main memory would contain random-access memory (RAM) chips and FPGAs that would, in addition to everything else, implement double-error-detecting and single-error-correcting memory functions to enable recovery from single-bit errors.

Some, Raphael

Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network

Ionizing radiation from cosmic rays and gammas can induce discontinuous jumps in the environmental charge of superconducting qubits (charge jumps), causing correlated errors that challenge fault-tolerant quantum computing while simultaneously providing a detection signature for quantum sensing applications. Current detection methods operate offline, introducing latency incompatible with in-the-loop qubit control. In this paper, an online detector of charge jumps for superconducting qubits, based on a dilated causal convolutional neural network (DCCNN) designed for in-the-loop deployment on the Quantum Instrumentation Control Kit (QICK) platform, is presented. The network is trained on synthetic Ramsey tomography scans generated from qubit templates measured at the Northwestern Experimental Underground Site (NEXUS) at Fermilab, and translated to FPGA firmware via hls4ml with ap_fixed$\langle 16,6 \rangle$ quantization, reaching a per-inference latency of $6.19 μ$s on the Zynq UltraScale+ RFSoC ZCU216. At this operating point the DCCNN matches the detection efficiency of the established offline $χ^2$ algorithm ($0.843 \pm 0.022$ vs. $0.866 \pm 0.020$ on $|Δq| \in [0.1, 0.5] e$ at matched false-positive rate), while requiring no per-qubit hyperparameter tuning. This shifts charge-jump detection from a post-hoc diagnostic to a control-loop primitive, enabling adaptive protocols that respond to radiation-induced events in situ, with applications to quantum-computing error mitigation and to the use of superconducting qubits as particle detectors.

Gaytan-Villarreal, Daniel [Carnegie Mellon U.]

The Use of Probabilistic Methods to Evaluate the Systems Impact of Component Design Improvements on Large Turbofan Engines

Probabilistic Structural Analysis (PSA) is now commonly used for predicting the distribution of time/cycles to failure of turbine blades and other engine components. These distributions are typically based on fatigue/fracture and creep failure modes of these components. Additionally, reliability analysis is used for taking test data related to particular failure modes and calculating failure rate distributions of electronic and electromechanical components. How can these individual failure time distributions of structural, electronic and electromechanical component failure modes be effectively combined into a top level model for overall system evaluation of component upgrades, changes in maintenance intervals, or line replaceable unit (LRU) redesign? This paper shows an example of how various probabilistic failure predictions for turbine engine components can be evaluated and combined to show their effect on overall engine performance. A generic model of a turbofan engine was modeled using various Probabilistic Risk Assessment (PRA) tools (Quantitative Risk Assessment Software (QRAS) etc.). Hypothetical PSA results for a number of structural components along with mitigation factors that would restrict the failure mode from propagating to a Loss of Mission (LOM) failure were used in the models. The output of this program includes an overall failure distribution for LOM of the system. The rank and contribution to the overall Mission Success (MS) is also given for each failure mode and each subsystem. This application methodology demonstrates the effectiveness of PRA for assessing the performance of large turbine engines. Additionally, the effects of system changes and upgrades, the application of different maintenance intervals, inclusion of new sensor detection of faults and other upgrades were evaluated in determining overall turbine engine reliability.

Packard, Michael H.

Wellbore Stability and Mud Loss Management in Geothermal Drilling: Optimizing Mud Weight to Mitigate Tensile Wellbore Fracturing at The Geysers, California

As part of a U.S. Department of Energy (DOE) Geothermal Technologies Office-funded initiative, Geysers Power Company, LLC, a subsidiary of Calpine Corporation, has been working to enhance drilling performance at the world’s largest geothermal field, The Geysers, in northern California. In a recent drilling operation of the GDC-36 well, excessive mud losses were encountered, initially addressed through repeated but largely ineffective cement plugging. Ultimately, the most effective strategy was to drill blind through the loss zones, made feasible by the high rate of penetration (ROP) achieved with PDC bits, allowing significant progress before the mud tanks were depleted and water-sensitive argillic formation layers could collapse. In response to these challenges, the project team explored alternative methods to minimize downtime and risks associated with cement plugging and continuous mud loss and to contemplate the driving mechanisms for the losses. Wellbore imaging using Formation MicroImager (FMI) and Ultrasonic Borehole Imager (UBI) tools revealed longitudinal tensile fractures, which were attributed to mud weights exceeding the minimum circumferential stress resulting from the native stress field and formation pressure. This study examines the mud losses encountered and leverages wellbore imaging data to understand the mechanisms behind mud induced tensile fracturing in specific rock facies. Understanding fracture behavior across different lithologies is crucial, as fractures within the reservoir can enhance steam migration throughout the system. The reservoir at The Geysers lies within the Mesozoic Franciscan Assemblage, a tectonic mélange formed by subduction. It consists of metamorphosed turbidite sandstone (greywacke) and mudstone (argillite), oceanic upper crust (including greenstone and chert), and serpentinized ultramafic rocks - each exhibiting distinct geomechanical fracturing properties. The structural fabric of the Franciscan Assemblage was shaped by low-angle Mesozoic thrust faulting and later overprinted by sub-vertical strike-slip structures related to the Pacific-North American plate boundary. A wellbore stability model was developed using core measurements and logs to simulate fracturing scenarios during drilling under varying stress conditions. These simulations guided the development of an optimized mud weight management strategy that should enable adaptive adjustments during drilling, reducing the likelihood of tensile fracturing and mud losses, ultimately improving operational efficiency.

15 GEOTHERMAL ENERGY

Open Circuit Resonant (SansEC) Sensor Technology for Lightning Mitigation and Damage Detection and Diagnosis for Composite Aircraft Applications

Traditional methods to protect composite aircraft from lightning strike damage rely on a conductive layer embedded on or within the surface of the aircraft composite skin. This method is effective at preventing major direct effect damage and minimizes indirect effects to aircraft systems from lightning strike attachment, but provides no additional benefit for the added parasitic weight from the conductive layer. When a known lightning strike occurs, the points of attachment and detachment on the aircraft surface are visually inspected and checked for damage by maintenance personnel to ensure continued safe flight operations. A new multi-functional lightning strike protection (LSP) method has been developed to provide aircraft lightning strike protection, damage detection and diagnosis for composite aircraft surfaces. The method incorporates a SansEC sensor array on the aircraft exterior surfaces forming a "Smart skin" surface for aircraft lightning zones certified to withstand strikes up to 100 kiloamperes peak current. SansEC sensors are open-circuit devices comprised of conductive trace spiral patterns sans (without) electrical connections. The SansEC sensor is an electromagnetic resonator having specific resonant parameters (frequency, amplitude, bandwidth & phase) which when electromagnetically coupled with a composite substrate will indicate the electrical impedance of the composite through a change in its resonant response. Any measureable shift in the resonant characteristics can be an indication of damage to the composite caused by a lightning strike or from other means. The SansEC sensor method is intended to diagnose damage for both in-situ health monitoring or ground inspections. In this paper, the theoretical mathematical framework is established for the use of open circuit sensors to perform damage detection and diagnosis on carbon fiber composites. Both computational and experimental analyses were conducted to validate this new method and system for aircraft composite damage detection and diagnosis. Experimental test results on seeded fault damage coupons and computational modeling simulation results are presented. This paper also presents the shielding effectiveness along with the lightning direct effect test results from several different SansEC LSP and baseline protected and unprotected carbon fiber reinforced polymer (CFRP) test panels struck at 40 and 100 kiloamperes following a universal common practice test procedure to enable damage comparisons between SansEC LSP configurations and common practice copper mesh LSP approaches. The SansEC test panels were mounted in a LSP test bed during the lightning test. Electrical, mechanical and thermal parameters were measured during lightning attachment and are presented with post test nondestructive inspection comparisons. The paper provides correlational results between the SansEC sensors computed electric field distribution and the location of the lightning attachment on the sensor trace and visual observations showing the SansEC sensor's affinity for dispersing the lightning attachment.

Szatkowski, George N.

Practical Application of PRA as an Integrated Design Tool for Space Systems

This paper presents the application of the first comprehensive Probabilistic Risk Assessment (PRA) during the design phase of a joint NASA/NOAA weather satellite program, Geostationary Operational Environmental Satellite Series R (GOES-R). GOES-R is the next generation weather satellite primarily to help understand the weather and help save human lives. PRA has been used at NASA for Human Space Flight for many years. PRA was initially adopted and implemented in the operational phase of manned space flight programs and more recently for the next generation human space systems. Since its first use at NASA, PRA has become recognized throughout the Agency as a method of assessing complex mission risks as part of an overall approach to assuring safety and mission success throughout project lifecycles. PRA is now included as a requirement during the design phase of both NASA next generation manned space vehicles as well as for high priority robotic missions. The influence of PRA on GOES-R design and operation concepts are discussed in detail. The GOES-R PRA is unique at NASA for its early implementation. It also represents a pioneering effort to integrate risks from both Spacecraft (SC) and Ground Segment (GS) to fully assess the probability of achieving mission objectives. PRA analysts were actively involved in system engineering and design engineering to ensure that a comprehensive set of technical risks were correctly identified and properly understood from a design and operations perspective. The analysis included an assessment of SC hardware and software, SC fault management system, GS hardware and software, common cause failures, human error, natural hazards, solar weather and infrastructure (such as network and telecommunications failures, fire). PRA findings directly resulted in design changes to reduce SC risk from micro-meteoroids. PRA results also led to design changes in several SC subsystems, e.g. propulsion, guidance, navigation and control (GNC), communications, mechanisms, and command and data handling (C&DH). The fault tree approach assisted in the development of the fault management system design. Human error analysis, which examined human response to failure, indicated areas where automation could reduce the overall probability of gaps in operation by half. In addition, the PRA brought to light many potential root causes of system disruptions, including earthquakes, inclement weather, solar storms, blackouts and other extreme conditions not considered in the typical reliability and availability analyses. Ultimately the PRA served to identify potential failures that, when mitigated, resulted in a more robust design, as well as to influence the program's concept of operations. The early and active integration of PRA with system and design engineering provided a well-managed approach for risk assessment that increased reliability and availability, optimized lifecyc1e costs, and unified the SC and GS developments.

Kalia, Prince

Detecting Short Circuits: Post Accident Electric Vehicle Battery Safety Check

Fast and accurate detection of soft short circuits (SCs) in the battery packs of damaged electric vehicles is needed by first responders and mechanics to mitigate the potential risk from battery fires that may occur hours, days, or weeks after an accident. Here, this paper presents an SC-detection algorithm for potentially damaged lithium-ion batteries that works quickly and without a priori knowledge of the battery-pack chemistry, capacity, state of charge, or state of health. The proposed universal SC-detection algorithm is designed to be implemented on an inexpensive handheld device that can connect to and monitor the voltages of all cells in a pack. Transient filtering and linear-quadratic state observation provide estimates of normalized SC current for every cell in the pack. Cells with SC-current estimates outside a sigma-based threshold are detected. Simulations, experiments, and electric vehicle (EV) crash data are used to verify the speed, sensitivity, and accuracy of the method, demonstrating 96% accurate detection of 0.0027 C SCs in under 1 h for 5S cell groups in the lab and no false positives for crashed Volkswagen, Chevrolet, and Tesla vehicles without SCs.

25 - ENERGY STORAGE

NASA Taxonomies for Searching Problem Reports and FMEAs

Many types of hazard and risk analyses are used during the life cycle of complex systems, including Failure Modes and Effects Analysis (FMEA), Hazard Analysis, Fault Tree and Event Tree Analysis, Probabilistic Risk Assessment, Reliability Analysis and analysis of Problem Reporting and Corrective Action (PRACA) databases. The success of these methods depends on the availability of input data and the analysts knowledge. Standard nomenclature can increase the reusability of hazard, risk and problem data. When nomenclature in the source texts is not standard, taxonomies with mapping words (sets of rough synonyms) can be combined with semantic search to identify items and tag them with metadata based on a rich standard nomenclature. Semantic search uses word meanings in the context of parsed phrases to find matches. The NASA taxonomies provide the word meanings. Spacecraft taxonomies and ontologies (generalization hierarchies with attributes and relationships, based on terms meanings) are being developed for types of subsystems, functions, entities, hazards and failures. The ontologies are broad and general, covering hardware, software and human systems. Semantic search of Space Station texts was used to validate and extend the taxonomies. The taxonomies have also been used to extract system connectivity (interaction) models and functions from requirements text. Now the Reconciler semantic search tool and the taxonomies are being applied to improve search in the Space Shuttle PRACA database, to discover recurring patterns of failure. Usual methods of string search and keyword search fall short because the entries are terse and have numerous shortcuts (irregular abbreviations, nonstandard acronyms, cryptic codes) and modifier words cannot be used in sentence context to refine the search. The limited and fixed FMEA categories associated with the entries do not make the fine distinctions needed in the search. The approach assigns PRACA report titles to problem classes in the taxonomy. Each ontology class includes mapping words - near-synonyms naming different manifestations of that problem class. The mapping words for Problems, Entities and Functions are converted to a canonical form plus any of a small set of modifier words (e.g. non-uniformity NOT + UNIFORM.) The report titles are parsed as sentences if possible, or treated as a flat sequence of word tokens if parsing fails. When canonical forms in the title match mapping words, the PRACA entry is associated with the corresponding Problem, Entity or Function in the ontology. The user can search for types of failures associated with types of equipment, clustering by type of problem (e.g., all bearings found with problems of being uneven: rough, irregular, gritty ). The results could also be used for tagging PRACA report entries with rich metadata. This approach could also be applied to searching and tagging failure modes, failure effects and mitigations in FMEAs. In the pilot work, parsing 52K+ truncated titles (the test cases that were available), has resulted in identification of both a type of equipment and type of problem in about 75% of the cases. The results are displayed in a manner analogous to Google search results. The effort has also led to the enrichment of the taxonomy, adding some new categories and many new mapping words. Further work would make enhancements that have been identified for improving the clustering and further reducing the false alarm rate. (In searching for recurring problems, good clustering is more important than reducing false alarms). Searching complete PRACA reports should lead to immediate improvement.

Malin, Jane T.

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASA's Space Launch System

The development of the Space Launch System (SLS) launch vehicle requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The characteristics of these systems must be matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large complex systems engineering challenge being addressed in part by focusing on the specific subsystems handling of off-nominal mission and fault tolerance. Using traditional model based system and software engineering design principles from the Unified Modeling Language (UML), the Mission and Fault Management (M&FM) algorithms are crafted and vetted in specialized Integrated Development Teams composed of multiple development disciplines. NASA also has formed an M&FM team for addressing fault management early in the development lifecycle. This team has developed a dedicated Vehicle Management End-to-End Testbed (VMET) that integrates specific M&FM algorithms, specialized nominal and off-nominal test cases, and vendor-supplied physics-based launch vehicle subsystem models. The flexibility of VMET enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the algorithms utilizing actual subsystem models. The intent is to validate the algorithms and substantiate them with performance baselines for each of the vehicle subsystems in an independent platform exterior to flight software test processes. In any software development process there is inherent risk in the interpretation and implementation of concepts into software through requirements and test processes. Risk reduction is addressed by working with other organizations such as S&MA, Structures and Environments, GNC, Orion, the Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission and Loss of Crew probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and detection and responses that can be tested in VMET and confirm that responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - ARINC 653 partitioned OS, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM. The plan for VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by Flight Software. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure the effectiveness of M&FM algorithms performance in the FSW development and test processes. This paper is outlined in a systematic fashion analogous to a lifecycle process flow for engineering development of algorithms into software and testing. Section I describes the NASA SLS M&FM context, presenting the current infrastructure, leading principles, methods, and participants. Section II defines the testing philosophy of the M&FM algorithms as related to VMET followed by section III, which presents the modeling methods of the algorithms to be tested and validated in VMET. Its details are then further presented in section IV followed by Section V presenting integration, test status, and state analysis. Finally, section VI addresses the summary and forward directions followed by the appendices presenting relevant information on terminology and documentation.

Trevino, Luis

A Vehicle Management End-to-End Testing and Analysis Platform for Validation of Mission and Fault Management Algorithms to Reduce Risk for NASAs Space Launch System

The engineering development of the National Aeronautics and Space Administration's (NASA) new Space Launch System (SLS) requires cross discipline teams with extensive knowledge of launch vehicle subsystems, information theory, and autonomous algorithms dealing with all operations from pre-launch through on orbit operations. The nominal and off-nominal characteristics of SLS's elements and subsystems must be understood and matched with the autonomous algorithm monitoring and mitigation capabilities for accurate control and response to abnormal conditions throughout all vehicle mission flight phases, including precipitating safing actions and crew aborts. This presents a large and complex systems engineering challenge, which is being addressed in part by focusing on the specific subsystems involved in the handling of off-nominal mission and fault tolerance with response management. Using traditional model-based system and software engineering design principles from the Unified Modeling Language (UML) and Systems Modeling Language (SysML), the Mission and Fault Management (M&FM) algorithms for the vehicle are crafted and vetted in Integrated Development Teams (IDTs) composed of multiple development disciplines such as Systems Engineering (SE), Flight Software (FSW), Safety and Mission Assurance (S&MA) and the major subsystems and vehicle elements such as Main Propulsion Systems (MPS), boosters, avionics, Guidance, Navigation, and Control (GNC), Thrust Vector Control (TVC), and liquid engines. These model-based algorithms and their development lifecycle from inception through FSW certification are an important focus of SLS's development effort to further ensure reliable detection and response to off-nominal vehicle states during all phases of vehicle operation from pre-launch through end of flight. To test and validate these M&FM algorithms a dedicated test-bed was developed for full Vehicle Management End-to-End Testing (VMET). For addressing fault management (FM) early in the development lifecycle for the SLS program, NASA formed the M&FM team as part of the Integrated Systems Health Management and Automation Branch under the Spacecraft Vehicle Systems Department at the Marshall Space Flight Center (MSFC). To support the development of the FM algorithms, the VMET developed by the M&FM team provides the ability to integrate the algorithms, perform test cases, and integrate vendor-supplied physics-based launch vehicle (LV) subsystem models. Additionally, the team has developed processes for implementing and validating the M&FM algorithms for concept validation and risk reduction. The flexibility of the VMET capabilities enables thorough testing of the M&FM algorithms by providing configurable suites of both nominal and off-nominal test cases to validate the developed algorithms utilizing actual subsystem models such as MPS, GNC, and others. One of the principal functions of VMET is to validate the M&FM algorithms and substantiate them with performance baselines for each of the target vehicle subsystems in an independent platform exterior to the flight software test and validation processes. In any software development process there is inherent risk in the interpretation and implementation of concepts from requirements and test cases into flight software compounded with potential human errors throughout the development and regression testing lifecycle. Risk reduction is addressed by the M&FM group but in particular by the Analysis Team working with other organizations such as S&MA, Structures and Environments, GNC, Orion, Crew Office, Flight Operations, and Ground Operations by assessing performance of the M&FM algorithms in terms of their ability to reduce Loss of Mission (LOM) and Loss of Crew (LOC) probabilities. In addition, through state machine and diagnostic modeling, analysis efforts investigate a broader suite of failure effects and associated detection and responses to be tested in VMET to ensure reliable failure detection, and confirm responses do not create additional risks or cause undesired states through interactive dynamic effects with other algorithms and systems. VMET further contributes to risk reduction by prototyping and exercising the M&FM algorithms early in their implementation and without any inherent hindrances such as meeting FSW processor scheduling constraints due to their target platform - the ARINC 6535-partitioned Operating System, resource limitations, and other factors related to integration with other subsystems not directly involved with M&FM such as telemetry packing and processing. The baseline plan for use of VMET encompasses testing the original M&FM algorithms coded in the same C++ language and state machine architectural concepts as that used by FSW. This enables the development of performance standards and test cases to characterize the M&FM algorithms and sets a benchmark from which to measure their effectiveness and performance in the exterior FSW development and test processes. This paper is outlined in a systematic fashion analogous to a lifecycle process flow for engineering development of algorithms into software and testing. Section I describes the NASA SLS M&FM context, presenting the current infrastructure, leading principles, methods, and participants. Section II defines the testing philosophy of the M&FM algorithms as related to VMET followed by section III, which presents the modeling methods of the algorithms to be tested and validated in VMET. Its details are then further presented in section IV followed by Section V presenting integration, test status, and state analysis. Finally, section VI addresses the summary and forward directions followed by the appendices presenting relevant information on terminology and documentation.

Trevino, Luis