Search NASASearch

Engineering topics

Daniel Hulse

Publications and source records attributed to Daniel Hulse.

At least 19 records

Workshop on the Role of Design Assurance in System-Wide Safety’s Safety Demonstrator Series

This report summarizes a three-hour hybrid in-person/online workshop on June 7, 2024 on the topic of design assurance and the Safety Demonstrator Series (SDS), which was held at NASA Ames Research Center. The SDS provides an operational demonstration of, and recommendations for, requirements and standards necessary to monitor, assess, and mitigate risks to assure safety in disaster-oriented operations. Over sixty NASA personnel participated in the workshop. Four main topics were discussed: (1) assurance needs for the Safety Demonstrator Series, (2) assurance and the In-Time Aviation Safety Management System (IASMS), (3) in-time assurance: existing efforts and future opportunities, and (4) demonstrating assurance tools in the Safety Demonstrators. Key takeaways are as follows: 1. Design assurance tools can be used to assure an In-Time Aviation Safety Management System, the systems that comprise it, and other systems or missions. Assurance must consider both systems and components and include the interactions between elements in both a systems/aircraft context and a systems-of-systems/airspace context. 2. Design-time assurance activities can support the identification of monitors needed for operational assurance activities (i.e., the “monitor” function in the monitor-assess-mitigate paradigm at the heart of the IASMS concept). 3. A major opportunity for design-time assurance tools to contribute to the IASMS concept is to support rapid re-validation of systems. This will be particularly important for (1) supporting novel operations in the IASMS, where operational data may disprove design-time assumptions (motivating re-analysis of system safety) and (2) adapting technologies (e.g., AI/ML for autonomous operations) to new operational domains, where there may be new or different safety considerations not included in the initial scope of operations. 4. There are several design assurance tools under development in the System-Wide Safety project that can support assurance of Services, Functions, and Capabilities (SFCs) in the Safety Demonstrators. Transitioning these tools from one-off research projects into a functioning part of IASMS assurance will require closer integration between these tools. 5. It is not clear whether the role of design assurance tools is primarily as a part of IASMS architecture or as an external check on IASMS. The workshop consisted of four discussion topics initiated via four lightning talks by System-Wide Safety researchers. Discussions utilized Mural to engage both in-person and online participants in the hybrid format. Polls and surveys were also utilized to gather participant input. The workshop closed with a reflection activity for participants, as well as new ideas for collaboration and coordination of ongoing System-Wide Safety research.

design assurance

Recommendations on Evidence and Process for Certification of Learning-enabled Components in Aerospace Systems

This report primarily identifies a collection of relevant and necessary evidence for assurance of machine learnt components (MLCs)—also known as learning-enabled components—integrated into aircraft systems, and gives preliminary suggestions on the elements of a certification process that invoke the identified evidence. The main focus is on feedforward neural networks that are static and trained offline through supervised learning. A brief background on the generic elements of the lifecycle of an MLC is given to contextualize the assurance considerations and, consequently, the evidence that is relevant and necessary to support certification. At the level of an MLC, those considerations relate to: (i) the consistency and correctness of MLC contributions to system functions in the context of a validated functional intent; and (ii) the absence of MLC contributions to aircraft-level failure conditions. At an ML model level, confidence in model and data properties contribute to assurance of the containing MLC, in particular: (a) generalizability and robustness of models, in the presence of inputs not previously seen during training, disturbances to inputs, and unexpected inputs; and (b) valid data, i.e., data that are at least representative, relevant, complete, and accurate. Evidence for the above span the elements of the ML lifecycle, and includes, at a minimum, lifecycle artifacts that pertain to: (1) properties of requirements capturing functional intent, safety constraints, and aspects of the intended use and operating environment; (2) model performance, model complexity and design, and algorithm choice; (3) achievement of required performance at the levels of a trained model during model development, a trained model after model development is complete, and a trained model that is transformed into an executable equivalent; (4) model implementation aspects necessary for transforming a trained model into the executable equivalent; (5) integration of the executable trained model into the containing MLC, and eventually the larger system; and, (6) lastly, the verification and validation (V&V) of each of the above. Such V&V lifecycle artifacts themselves include: aspects of coverage, e.g., of various levels of requirements by the input space of the model and the data; traceability (where applicable); application of formal methods for property specification, analysis, and checking. Examples of evidence generation methods and tools further ground the discussion on what constitutes evidence, and the contribution to assurance during certification. The identified assurance considerations and supporting evidence is not a comprehensive set. Additionally, neither what should be considered as sufficient evidence relative to the assigned criticality of an MLC, nor how criticality ought to be determined and adjusted, have been considered in this report. However, suggestions are made for potential activities of the ML lifecycle that are aimed at providing confidence that an MLC can be relied upon when integrated into its containing (aircraft) system. Those activities are proposed as candidate elements of a certification process for MLCs. The main purpose of this report to inform regulatory guidance and consensus standards that may be used to meet the safety intent of the applicable regulations.

Aviation safety

Supporting Hazard Analysis for Wildfire Response Using fmdtools and MIKA

The System Wide Safety (SWS) Safety Demonstrator (SD) Series drives development of an increasingly capable In-Time Aviation Safety Management System (IASMS) focusing on humanitarian applications, starting with wildfire response (SD-1). The goals of this report are to (1) provide an early hazard analysis and mitigation evaluation of wildfire response to support these efforts and (2) provide a demonstration of capabilities of the Fault Model Design Tools (fmdtools) and Manager for Intelligent Knowledge Access (MIKA) tools. fmdtools provides a modeling, simulation, and resiliency analysis framework in which a wildfire response model, the System Modeling and Analysis of Resiliency in Scalable Traffic Management for Emergency Response Operations (SMARt-STEReO), is built. MIKA is an intelligent knowledge manager with several capabilities, including assisting in hazard analysis by extracting and analyzing hazards from historical incident reports. The following topics are covered in the report: Understanding Wildfire Hazard Dynamics. We provide a description and simulated examples of how hazards occur in the SMARt-STEReO model of wildfire response and their effect on its outcome. This provides a common mental model and focuses the analysis presented in the remainder of the report. Wildfire Hazard Identification. MIKA identifies wildfire hazards from three relevant datasets: the ICS-209-PLUS, SAFECOM, and SAFENET. Hazards are manually organized into a taxonomy and MIKA analyzes each hazard’s effects, likelihood, severity, and risk. Evaluating Mitigation Strategies. The SMARt-STEReO wildfire response model built in fmdtools evaluates a subset of identified hazards. Specifically, we simulate the effect of communications faults and equipment faults on operator safety, the effect of changing winds and flammability, and a scenario with multiple ignition points and heavy smoke. Tool Limitations and Usage Considerations. We provide a discussion of appropriate tool use cases as well as limitations and considerations for usage. The tool findings are used to synthesize recommendations for wildfire response operations, which can be captured as part of an IASMS. Key recommendations are as follows: Hazards are identified from a broad spectrum of sources including aircraft subsystems, operational sources, and ground crew operations. Highest risk operational environment hazards identified are Evacuations. The highest risk manned aerial operations hazard categorized is Jumper Operations Mishap. Ground crew hazards that are highest risk are Burns, Cargo Operations Overhead, Dehydration, Entrapment, Falling Objects, Heart Attacks, Heat Exhaustion, Inadequate Training or Certification, Vehicle Breakdown, and Vehicle Collision. Modelled containment failures arise from a mismatch between the difficulty of the firefighting scenario and the capacity (e.g., speed, effectiveness, awareness) of the response. In firefighting scenarios where containment is possible (e.g., because the fire does not spread too quickly), these mismatches can occur because of a change in environmental conditions (e.g., wind, flammability, etc) or because of planning, equipment, or communications faults. Improvements to communications increase the capacity of the firefighting response by reducing the time needed to respond to the fire. While surveillance does not increase this capacity by itself, it increases operator safety by increasing state awareness, enabling firefighters to evade approaching fires. Increasing both has a synergistic effect. In general, these performance and resilience increases generalize over fault scenarios as well as unforeseen changes to circumstances (i.e., wind, aridity, etc.). However, these improvements need to be designed so as not to make the system prone to persistent large-scale communications outages, which can reduce performance.

Hazard analysis

The System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO)

Wildfire emergency response has remained rooted in relatively low-tech solutions for coordination between ground and aerial assets. These low-tech solutions are robust for the remote environments in which wildfires are usually fought, but limit strategic cross-organizational support and the ability to deploy and effectively utilize aerial assets. As aircraft become more advanced and new technology, including drones, become available to firefighters, a new, more modern method of asset coordination is needed. NASA is working on a project called ‘Scalable Traffic Management for Emergency Response Operations’ (STEReO) to integrate unmanned aerial systems (UAS)and UAS traffic management (UTM)into wildfire response. STEReO’s goals include simplifying the coordination of aerial assets, improving the existing UAS framework, and increasing the role of additional autonomous systems to reduce human risk and to increase system resilience. This paper describes the development of the ‘System Modeling and Analysis of Resiliency in STEReO’ (SMARt-STEReO) project, which aims to model wildfire response and to quantify the additional system resilience that STEReO technology provides firefighters. This paper verifies SMARt-STEReO and defines its scope; it includes experimental and statistical analysis of the impact that the addition of UAS has on both performance metrics and also on performance resiliency response to a given fault. SMARt-STEReO is a grid-based model of fire propagation that incorporates varying crew responses. Through the use of a Python package called ‘fmdtools’, the model easily allows for the addition of faults to the system. These faults allow analysts to investigate various response parameters. Factors including terrain, fuel type and wind speed can be modified to affect the fire propagation; additionally, the number of ground crews, engines, fixed wing aircraft, helicopters, and UAS can be changed to affect the crew response. The communication lines between actors mimic those used in real life situations. This paper explains the development of SMARt-STEReO including background research, verification and validation, and preliminary experimental analysis of system resilience to both a minor and major fault in systems with and without UAS.

Resiliency

Understanding Resilience Optimization Architectures With an Optimization Problem Repository

Optimizing a system’s resilience can be challenging, especially when it involves considering both the inherent resilience of a robust design and the active resilience of a health management system to a set of computationally-expensive hazard simulations. While prior work has developed specialized architectures to effectively and efficiently solve combined design and resilience optimization problems, the comparison of these architectures has been limited to a single case study. To further study resilience optimization formulations, this work develops a problem repository which includes previously-developed resilience optimization problems and additional problems presented in this work: a notional system resilience model, a pandemic response model, and a cooling tank hazard prevention model. This work then uses models in the repository at large to understand the characteristics of resilience optimization problems and study the applicability of optimization architectures and decomposition strategies. Based on the comparisons in the repository, applying an optimization architecture effectively requires understanding the alignment and coupling relationships between the design and resilience models, as well as the efficiency characteristics of the algorithms. While alignment determines the necessity of a surrogate of resilience cost in the upper-level design problem, coupling determines the overall applicability of a sequential, alternating, or bilevel structure. Additionally, the application of decomposition strategies is dependent on there being limited interactions between variable sets, which often does not hold when a resilience policy is parameterized in terms of actions to take in hazardous model states rather than specific given scenarios.

Resilience

The System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO)

NASA's Scalable Traffic Management for Emergency Response Operations (STEReO) project aims to leverage Unmanned Aerial Systems (UAS) and UAS Traffic Management (UTM) to improve asset coordination and overall emergency response. One application of STEReO is wildfire response, which is the focus of this research. In order to implement the operations described in the STEReO project, these additions must have tangible benefits and proven safety. To this end, the System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO) project constructs a simulation model, developed through the Python modeling and resiliency analysis package fmdtools. The model describes wildfire response operations, including current operational concepts and emerging concepts utilizing UAS as described in STEReO. While previous simulation models focus primarily on fire propagation with some models including emergency response intervention, SMART-STEReO evaluates the system performance and resilience benefits gained by the addition of UAS and UTM. Due to the novelty and complexity of the model, initial model verification and validation efforts are conducted and a detailed description of the model is provided. Preliminary results from experimental analysis on the SMARt-STEReO model indicate that when compared to current operations, the addition of UAS in wildfire operations results in improved response efforts, in terms of fewer acres burned, as well as improved system resilience in response to a given fault.

Sequoia Andrade

Synthetic Failure Mode Generation for Resilience Analysis and Failure Mechanism Discovery

Traditional risk-based design processes seek to mitigate operational hazards by manually identifying possible faults and corresponding mitigation strategies—a tedious process which critically relies on the designer’s limited knowledge. Resilience-based design, on the other hand, seeks to embody generic hazard-mitigating properties in the system to mitigate unknown hazards, often by modelling the system's response to potential hazardous events. This work adapts this approach to the traditional risk-based design process to synthetically generate hazardous modes, by representing them as a unique combination of internal component health-states which can then be injected and simulated in a model of the system failure dynamics. The design process may then reduce the risk of unknown internal hazards by iteratively mitigating the effects of these modes. The performance of this approach is evaluated in a model of an autonomous rover, where cluster analysis shows that elaborating the space of synthetic faults in the drive system using this approach uncovers a wider range of possible hazardous trajectories and failure consequences within each trajectory. However, this increase in hazard information comes at a high computational expense, highlighting the need for advanced, efficient methods to search and sample the hazard space.

Simulation

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models–specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

resilience

Synthetic Failure Mode Generation for Resilience Analysis and Failure Mechanism Discovery

Traditional risk-based design processes seek to mitigate operational hazards by manually identifying possible faults and corresponding mitigation strategies—a tedious process which critically relies on the designer’s limited knowledge. Resilience-based design, on the other hand, seeks to embody generic hazard-mitigating properties in the system to mitigate unknown hazards, often by modelling the system's response to potential hazardous events. This work adapts this approach to the traditional risk-based design process to synthetically generate hazardous modes, by representing them as a unique combination of internal component health-states which can then be injected and simulated in a model of the system failure dynamics. The design process may then reduce the risk of unknown internal hazards by iteratively mitigating the effects of these modes. The performance of this approach is evaluated in a model of an autonomous rover, where cluster analysis shows that elaborating the space of synthetic faults in the drive system using this approach uncovers a wider range of possible hazardous trajectories and failure consequences within each trajectory. However, this increase in hazard information comes at a high computational expense, highlighting the need for advanced, efficient methods to search and sample the hazard space.

Simulation

Modeling Distributed Situation Awareness in Resilience-Based Design of Complex Engineered Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad

Uncovering Hazards Using a Multi-Objective Optimization to Explore the Faulty State-Space

Considering resilience when designing complex engineered systems is crucial to ensure the system is safe under unexpected hazardous scenarios. Traditional risk-based approaches, such as Failure Modes and Effects Analysis (FMEA) are useful for designing the system to mitigate hazardous scenarios that can be identified by the designer, but often require experience or prior knowledge of system failures to generate. More recently, researchers have developed simulation tools that enable the designer to model large sets of hazardous scenarios (driven by both internal faults and external factors) through simulation. While these tools enable a wider scope of fault modes to be evaluated (e.g., by injecting combined set of fault modes or injecting modes at different times), the resulting assessments (like FMEA) still require knowledge of the specific modes to be evaluated. However, failure to analyze a wide variety of fault scenarios can lead to an incomplete picture of the system resilience, especially to "surprise events'' which may be difficult for the designer to identify and predict beforehand. To overcome this challenge, previous work developed a fault sampling approach for resilience simulations which would procedurally-generate a wide variety of potential faults by systematically perturbing the health states of the system. While the resulting fault modes generated covered a much larger space hazards than would be otherwise considered (and identified many unique failure trajectories which would not have otherwise been identified), it also significantly increased the computational cost of the analysis and resulted in the simulation and analysis of a large set of essentially duplicate scenarios. Additionally, as the number of dimensions in the faulty state-space increases, the full elaboration of possible modes becomes computationally infeasible, justifying the use of a more targeted search. To resolve this limitation, this work proposes the use of a multiobjective optimization algorithm to search the health state space for potential fault modes that are both (1) hazardous and (2) unique. To solve this type of problem, this work proposes the use of a cooperative co-evolutionary algorithm. To demonstrate this approach, it will be applied to a model of an autonomous rover which uses line markings to navigate, focusing on potential hazards in the drive system which could cause the rover to crash. To determine the merit of the approach, it will further be compared with the previously-presented range elaboration approach and a random mode generation approach on the basis of computational efficiency and found modes.

Resilience

fmdtools Tutorial: Intro to Resilience Modelling, Simulation, and Visualization in Python With fmdtools

This workshop will cover the basics of using the fmdtools package for the simulation of hazardous scenarios for resilience simulation. The fmdtools simulation package is an open-source python toolkit for simulating the dynamic response of a system to internal and externally-driven hazardous scenarios, including faults and environmental conditions, that can be used to analyze the risks related to these hazards. Prior to the development of fmdtools, researchers had to either adapt an (often limited) propriety toolkit or develop their own design/simulation/analysis codes to develop their models of hazardous events, a significant technical burden to both (1) leveraging resilience modeling methodologies and (2) extending these methodologies with their own contributions. The fmdtools package provides a number of model constructs and simulation and analysis methods to enable the designer to focus to solely on their modeling case-study while still enabling a significant degree of model expressiveness and adaptability via Python-based model definition. This tutorial will present the setup of fmdtools and a high-level overview of its use, as well as some simple examples for understanding how to leverage its modelling, simulation, and analysis capabilities. Familiarity with jupyter notebook and basic python will be assumed.

Daniel Hulse

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models--specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

Daniel Hulse

Can Resilience Assessments Inform Early Design Human Factors Decision-making?

There is a growing call among researchers for tighter coupling between human factors and human reliability assessments. In this research, we explore if early design stage resilience assessments can help bridge some of the gaps between human factors and human reliability assessments. Resilience in systems is their ability to recover reasonably and operate within acceptable bounds during failures and unexpected events. The fmdtools toolkit allows designers to assess the resilience of a system by modeling the human error and machine-related failure propagation in both nominal and faulty scenarios during the early design stages. As a result, the fmdtools toolkit has a low-fidelity dynamic human reliability assessment component built into it. In this paper, we study if the results from the fmdtools simulations can help inform and prioritize human factors design decision-making, resulting in tighter coupling between human factors and human reliability assessments during the design process. Specifically, we explore the results from a rover design example to understand the types of information that can help guide human factor-related decision-making.

Resilience-based Design