Search NASASearch

Engineering topics

Lukman Irshad

Publications and source records attributed to Lukman Irshad.

At least 19 records

Workshop on the Role of Design Assurance in System-Wide Safety’s Safety Demonstrator Series

This report summarizes a three-hour hybrid in-person/online workshop on June 7, 2024 on the topic of design assurance and the Safety Demonstrator Series (SDS), which was held at NASA Ames Research Center. The SDS provides an operational demonstration of, and recommendations for, requirements and standards necessary to monitor, assess, and mitigate risks to assure safety in disaster-oriented operations. Over sixty NASA personnel participated in the workshop. Four main topics were discussed: (1) assurance needs for the Safety Demonstrator Series, (2) assurance and the In-Time Aviation Safety Management System (IASMS), (3) in-time assurance: existing efforts and future opportunities, and (4) demonstrating assurance tools in the Safety Demonstrators. Key takeaways are as follows: 1. Design assurance tools can be used to assure an In-Time Aviation Safety Management System, the systems that comprise it, and other systems or missions. Assurance must consider both systems and components and include the interactions between elements in both a systems/aircraft context and a systems-of-systems/airspace context. 2. Design-time assurance activities can support the identification of monitors needed for operational assurance activities (i.e., the “monitor” function in the monitor-assess-mitigate paradigm at the heart of the IASMS concept). 3. A major opportunity for design-time assurance tools to contribute to the IASMS concept is to support rapid re-validation of systems. This will be particularly important for (1) supporting novel operations in the IASMS, where operational data may disprove design-time assumptions (motivating re-analysis of system safety) and (2) adapting technologies (e.g., AI/ML for autonomous operations) to new operational domains, where there may be new or different safety considerations not included in the initial scope of operations. 4. There are several design assurance tools under development in the System-Wide Safety project that can support assurance of Services, Functions, and Capabilities (SFCs) in the Safety Demonstrators. Transitioning these tools from one-off research projects into a functioning part of IASMS assurance will require closer integration between these tools. 5. It is not clear whether the role of design assurance tools is primarily as a part of IASMS architecture or as an external check on IASMS. The workshop consisted of four discussion topics initiated via four lightning talks by System-Wide Safety researchers. Discussions utilized Mural to engage both in-person and online participants in the hybrid format. Polls and surveys were also utilized to gather participant input. The workshop closed with a reflection activity for participants, as well as new ideas for collaboration and coordination of ongoing System-Wide Safety research.

design assurance

Recommendations on Evidence and Process for Certification of Learning-enabled Components in Aerospace Systems

This report primarily identifies a collection of relevant and necessary evidence for assurance of machine learnt components (MLCs)—also known as learning-enabled components—integrated into aircraft systems, and gives preliminary suggestions on the elements of a certification process that invoke the identified evidence. The main focus is on feedforward neural networks that are static and trained offline through supervised learning. A brief background on the generic elements of the lifecycle of an MLC is given to contextualize the assurance considerations and, consequently, the evidence that is relevant and necessary to support certification. At the level of an MLC, those considerations relate to: (i) the consistency and correctness of MLC contributions to system functions in the context of a validated functional intent; and (ii) the absence of MLC contributions to aircraft-level failure conditions. At an ML model level, confidence in model and data properties contribute to assurance of the containing MLC, in particular: (a) generalizability and robustness of models, in the presence of inputs not previously seen during training, disturbances to inputs, and unexpected inputs; and (b) valid data, i.e., data that are at least representative, relevant, complete, and accurate. Evidence for the above span the elements of the ML lifecycle, and includes, at a minimum, lifecycle artifacts that pertain to: (1) properties of requirements capturing functional intent, safety constraints, and aspects of the intended use and operating environment; (2) model performance, model complexity and design, and algorithm choice; (3) achievement of required performance at the levels of a trained model during model development, a trained model after model development is complete, and a trained model that is transformed into an executable equivalent; (4) model implementation aspects necessary for transforming a trained model into the executable equivalent; (5) integration of the executable trained model into the containing MLC, and eventually the larger system; and, (6) lastly, the verification and validation (V&V) of each of the above. Such V&V lifecycle artifacts themselves include: aspects of coverage, e.g., of various levels of requirements by the input space of the model and the data; traceability (where applicable); application of formal methods for property specification, analysis, and checking. Examples of evidence generation methods and tools further ground the discussion on what constitutes evidence, and the contribution to assurance during certification. The identified assurance considerations and supporting evidence is not a comprehensive set. Additionally, neither what should be considered as sufficient evidence relative to the assigned criticality of an MLC, nor how criticality ought to be determined and adjusted, have been considered in this report. However, suggestions are made for potential activities of the ML lifecycle that are aimed at providing confidence that an MLC can be relied upon when integrated into its containing (aircraft) system. Those activities are proposed as candidate elements of a certification process for MLCs. The main purpose of this report to inform regulatory guidance and consensus standards that may be used to meet the safety intent of the applicable regulations.

Aviation safety

Supporting Hazard Analysis for Wildfire Response Using fmdtools and MIKA

The System Wide Safety (SWS) Safety Demonstrator (SD) Series drives development of an increasingly capable In-Time Aviation Safety Management System (IASMS) focusing on humanitarian applications, starting with wildfire response (SD-1). The goals of this report are to (1) provide an early hazard analysis and mitigation evaluation of wildfire response to support these efforts and (2) provide a demonstration of capabilities of the Fault Model Design Tools (fmdtools) and Manager for Intelligent Knowledge Access (MIKA) tools. fmdtools provides a modeling, simulation, and resiliency analysis framework in which a wildfire response model, the System Modeling and Analysis of Resiliency in Scalable Traffic Management for Emergency Response Operations (SMARt-STEReO), is built. MIKA is an intelligent knowledge manager with several capabilities, including assisting in hazard analysis by extracting and analyzing hazards from historical incident reports. The following topics are covered in the report: Understanding Wildfire Hazard Dynamics. We provide a description and simulated examples of how hazards occur in the SMARt-STEReO model of wildfire response and their effect on its outcome. This provides a common mental model and focuses the analysis presented in the remainder of the report. Wildfire Hazard Identification. MIKA identifies wildfire hazards from three relevant datasets: the ICS-209-PLUS, SAFECOM, and SAFENET. Hazards are manually organized into a taxonomy and MIKA analyzes each hazard’s effects, likelihood, severity, and risk. Evaluating Mitigation Strategies. The SMARt-STEReO wildfire response model built in fmdtools evaluates a subset of identified hazards. Specifically, we simulate the effect of communications faults and equipment faults on operator safety, the effect of changing winds and flammability, and a scenario with multiple ignition points and heavy smoke. Tool Limitations and Usage Considerations. We provide a discussion of appropriate tool use cases as well as limitations and considerations for usage. The tool findings are used to synthesize recommendations for wildfire response operations, which can be captured as part of an IASMS. Key recommendations are as follows: Hazards are identified from a broad spectrum of sources including aircraft subsystems, operational sources, and ground crew operations. Highest risk operational environment hazards identified are Evacuations. The highest risk manned aerial operations hazard categorized is Jumper Operations Mishap. Ground crew hazards that are highest risk are Burns, Cargo Operations Overhead, Dehydration, Entrapment, Falling Objects, Heart Attacks, Heat Exhaustion, Inadequate Training or Certification, Vehicle Breakdown, and Vehicle Collision. Modelled containment failures arise from a mismatch between the difficulty of the firefighting scenario and the capacity (e.g., speed, effectiveness, awareness) of the response. In firefighting scenarios where containment is possible (e.g., because the fire does not spread too quickly), these mismatches can occur because of a change in environmental conditions (e.g., wind, flammability, etc) or because of planning, equipment, or communications faults. Improvements to communications increase the capacity of the firefighting response by reducing the time needed to respond to the fire. While surveillance does not increase this capacity by itself, it increases operator safety by increasing state awareness, enabling firefighters to evade approaching fires. Increasing both has a synergistic effect. In general, these performance and resilience increases generalize over fault scenarios as well as unforeseen changes to circumstances (i.e., wind, aridity, etc.). However, these improvements need to be designed so as not to make the system prone to persistent large-scale communications outages, which can reduce performance.

Hazard analysis

Synthetic Failure Mode Generation for Resilience Analysis and Failure Mechanism Discovery

Traditional risk-based design processes seek to mitigate operational hazards by manually identifying possible faults and corresponding mitigation strategies—a tedious process which critically relies on the designer’s limited knowledge. Resilience-based design, on the other hand, seeks to embody generic hazard-mitigating properties in the system to mitigate unknown hazards, often by modelling the system's response to potential hazardous events. This work adapts this approach to the traditional risk-based design process to synthetically generate hazardous modes, by representing them as a unique combination of internal component health-states which can then be injected and simulated in a model of the system failure dynamics. The design process may then reduce the risk of unknown internal hazards by iteratively mitigating the effects of these modes. The performance of this approach is evaluated in a model of an autonomous rover, where cluster analysis shows that elaborating the space of synthetic faults in the drive system using this approach uncovers a wider range of possible hazardous trajectories and failure consequences within each trajectory. However, this increase in hazard information comes at a high computational expense, highlighting the need for advanced, efficient methods to search and sample the hazard space.

Simulation

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models–specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

resilience

Synthetic Failure Mode Generation for Resilience Analysis and Failure Mechanism Discovery

Traditional risk-based design processes seek to mitigate operational hazards by manually identifying possible faults and corresponding mitigation strategies—a tedious process which critically relies on the designer’s limited knowledge. Resilience-based design, on the other hand, seeks to embody generic hazard-mitigating properties in the system to mitigate unknown hazards, often by modelling the system's response to potential hazardous events. This work adapts this approach to the traditional risk-based design process to synthetically generate hazardous modes, by representing them as a unique combination of internal component health-states which can then be injected and simulated in a model of the system failure dynamics. The design process may then reduce the risk of unknown internal hazards by iteratively mitigating the effects of these modes. The performance of this approach is evaluated in a model of an autonomous rover, where cluster analysis shows that elaborating the space of synthetic faults in the drive system using this approach uncovers a wider range of possible hazardous trajectories and failure consequences within each trajectory. However, this increase in hazard information comes at a high computational expense, highlighting the need for advanced, efficient methods to search and sample the hazard space.

Simulation

Modeling Distributed Situation Awareness in Resilience-Based Design of Complex Engineered Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad

Uncovering Hazards Using a Multi-Objective Optimization to Explore the Faulty State-Space

Considering resilience when designing complex engineered systems is crucial to ensure the system is safe under unexpected hazardous scenarios. Traditional risk-based approaches, such as Failure Modes and Effects Analysis (FMEA) are useful for designing the system to mitigate hazardous scenarios that can be identified by the designer, but often require experience or prior knowledge of system failures to generate. More recently, researchers have developed simulation tools that enable the designer to model large sets of hazardous scenarios (driven by both internal faults and external factors) through simulation. While these tools enable a wider scope of fault modes to be evaluated (e.g., by injecting combined set of fault modes or injecting modes at different times), the resulting assessments (like FMEA) still require knowledge of the specific modes to be evaluated. However, failure to analyze a wide variety of fault scenarios can lead to an incomplete picture of the system resilience, especially to "surprise events'' which may be difficult for the designer to identify and predict beforehand. To overcome this challenge, previous work developed a fault sampling approach for resilience simulations which would procedurally-generate a wide variety of potential faults by systematically perturbing the health states of the system. While the resulting fault modes generated covered a much larger space hazards than would be otherwise considered (and identified many unique failure trajectories which would not have otherwise been identified), it also significantly increased the computational cost of the analysis and resulted in the simulation and analysis of a large set of essentially duplicate scenarios. Additionally, as the number of dimensions in the faulty state-space increases, the full elaboration of possible modes becomes computationally infeasible, justifying the use of a more targeted search. To resolve this limitation, this work proposes the use of a multiobjective optimization algorithm to search the health state space for potential fault modes that are both (1) hazardous and (2) unique. To solve this type of problem, this work proposes the use of a cooperative co-evolutionary algorithm. To demonstrate this approach, it will be applied to a model of an autonomous rover which uses line markings to navigate, focusing on potential hazards in the drive system which could cause the rover to crash. To determine the merit of the approach, it will further be compared with the previously-presented range elaboration approach and a random mode generation approach on the basis of computational efficiency and found modes.

Resilience

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models--specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

Daniel Hulse

Can Resilience Assessments Inform Early Design Human Factors Decision-making?

There is a growing call among researchers for tighter coupling between human factors and human reliability assessments. In this research, we explore if early design stage resilience assessments can help bridge some of the gaps between human factors and human reliability assessments. Resilience in systems is their ability to recover reasonably and operate within acceptable bounds during failures and unexpected events. The fmdtools toolkit allows designers to assess the resilience of a system by modeling the human error and machine-related failure propagation in both nominal and faulty scenarios during the early design stages. As a result, the fmdtools toolkit has a low-fidelity dynamic human reliability assessment component built into it. In this paper, we study if the results from the fmdtools simulations can help inform and prioritize human factors design decision-making, resulting in tighter coupling between human factors and human reliability assessments during the design process. Specifically, we explore the results from a rover design example to understand the types of information that can help guide human factor-related decision-making.

Resilience-based Design

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad

Defining A Modelling Language to Support Functional Hazard Assessment

Functional Hazard Assessment (FHA) is a key early-stage engineering process that supports the incorporation of safety in design by identifying the high-level functional hazards the system may encounter. While many FHA-like methodologies have been proposed in the design engineering literature, many of these methodologies have had difficulty becoming accepted industry practice. Industry standards, on the other hand, either provide too little recommendation on how to represent the function of the system to perform FHA, or rely on existing design artefacts which insufficiently support the goals of the process. This paper presents some of the problems with current modeling languages (both proposed and used) for FHA which limit the scope, expressiveness, flexibility, and precision of the analysis. It then outlines desirable principles an FHA-supporting analysis language should embody, and introduces the Functional Reasoning Design Language (FRDL), a formal modeling language for describing the functional elements of a system and their interactions, which aims to satisfy these principles. To demonstrate the use of this language, the modeling and hazard analysis of a disaster response drone is presented. While this case study is limited in scope, it highlights how FRDL can represent system function while reducing the ambiguity present in typical FHA-supporting functional modeling languages

Hazard Assessment

Identifying Human Errors and Error Mechanisms From Accident Reports Using Large Language Models

Emerging operational concepts for aviation hinge on novel paradigms for human machine interaction. Critical to their safe operation is early consideration of human error into the design process. Existing methods for consideration of human error require significant expert input, which is challenging both in early design and in novel systems for which there is little existing safety expertise. In this research, we propose a methodology for identifying human error, error producing factors, and mechanisms in early design from historical incident reports. Additionally, we hypothesize that cross-domain sharing of lessons learned can aid with early design human considerations in circumstances where data is not relevant or incomplete. This is addressed by identifying causes of human error in aviation and railway domains through applying state-of-the art natural language processing techniques to historical incident reports. Using this method, it is possible to extract extensive reports on human error from past incidents. Using the proposed approach, we identify nine human errors from railway reports and fourteen from aviation reports, with three errors common to both domains. There is at least one error producing conditions for each human error while a majority of the errors have more than one error mechanism. We also found that a majority of the human errors, error producing factors, and error mechanisms (even if they are not common between the domains) can be used to inform safe operations across domains as long as the errors are not domain specific and are interpreted and contextualized using engineering judgement.

Human Errors

Defining A Modelling Language to Support Functional Hazard Assessment

Functional Hazard Assessment (FHA) is a key early-stage engineering process that supports the incorporation of safety in design by identifying the high-level functional hazards the system may encounter. While many FHA-like methodologies have been proposed in the design engineering literature, many of these methodologies have had difficulty becoming accepted industry practice. Industry standards, on the other hand, either provide little recommendation on how to represent the function of the system to perform FHA, or rely on readily-available models with little justification in design theory. This paper presents some of the problems with current modelling languages used for FHA which limit the scope, expressiveness, flexibility, and precision of the analysis, as well as desirable principles an FHA-supporting analysis language should embody. It further introduces the Functional Reasoning Design Language (FRDL), a formal modelling language for describing the functional behaviors of a system and their interactions which satisfies these principles. To demonstrate the use of this language, the modelling and hazard analysis of a disaster response drone is presented.

safety analysis

Towards Computational Functional Hazard Assessment (CFHA): A Gap Analysis and Concept for Emerging Aviation Systems

Given the current evolution of the National Airspace and future trajectory towards novel and evolving operations with varying levels of autonomy, complexity, and acceptable risk, there is an opportunity to support safety assurance by extending existing methodologies, such as Functional Hazard Assessment (FHA). In response to challenges in performing FHA for novel aviation concepts, we propose a concept for Computational Functional Hazard Assessment (CFHA), which provides processes, methods, and tools for incorporating external data to facilitate further exploration of the hazard space iterativelty throughout the design process. The core components of CFHA involve knowledge capture from historical and operational data, functional architecture specification via a formal modeling language, and simulation for hazardous scenario analysis. Through this concept, we aim to adapt conventional safety assessment to address the increasingly complex hazard space generated from emerging operations.

Seydou Mbaye