Search NASASearch

SEARCH · Search NASA

Results for “Safety Case Patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

31 records · Page 2

Analysis of severe atmospheric disturbances from airline flight records

Advanced methods were developed to determine time varying winds and turbulence from digital flight data recorders carried aboard modern airliners. Analysis of several cases involving severe clear air turbulence encounters at cruise altitudes has shown that the aircraft encountered vortex arrays generated by destabilized wind shear layers above mountains or thunderstorms. A model was developed to identify the strength, size, and spacing of vortex arrays. This model is used to study the effects of severe wind hazards on operational safety for different types of aircraft. The study demonstrates that small remotely piloted vehicles and executive aircraft exhibit more violent behavior than do large airliners during encounters with high-altitude vortices. Analysis of digital flight data from the accident at Dallas/Ft. Worth in 1985 indicates that the aircraft encountered a microburst with rapidly changing winds embedded in a strong outflow near the ground. A multiple-vortex-ring model was developed to represent the microburst wind pattern. This model can be used in flight simulators to better understand the control problems in severe microburst encounters.

Wingrove, R. C.

Analysis of severe atmospheric disturbances from airline flight records

Advanced methods were developed to determine time varying winds and turbulence from digital flight data recorders carried aboard modern airliners. Analysis of several cases involving severe clear air turbulence encounters at cruise altitudes has shown that the aircraft encountered vortex arrays generated by destabilized wind shear layers above mountains or thunderstorms. A model was developed to identify the strength, size, and spacing of vortex arrays. This model is used to study the effects of severe wind hazards on operational safety for different types of aircraft. It is demonstrated that small remotely piloted vehicles and executive aircraft exhibit more violent behavior than do large airliners during encounters with high-altitude vortices. Analysis of digital flight data from the accident at Dallas/Ft. Worth in 1985 indicates that the aircraft encountered a microburst with rapidly changing winds embedded in a strong outflow near the ground. A multiple-vortex-ring model was developed to represent the microburst wind pattern. This model can be used in flight simulators to better understand the control problems in severe microburst encounters.

Wingrove, R. C.

TruePAL – An AI Assistant for First Responder Safety

This paper presents the development of an AI assistant, Trusted and Explainable Artificial Intelligence for Saving Lives (TruePAL), to provide real-time warning of risks of potential crashes to the first responders. The TruePAL system employs an AI and deep learning technology for saving first responders and roadside crews lives in and around active traffic. A deep neural network (DNN) and a Non-Axiomatic Reasoning System (NARS) are implemented as an AI system. A mobile app with AI interface is developed to perform verbal communication with the first responders. The TruePAL team has developed an explainable AI approach by opening up the DNN blackbox to extract the activation filters of various features and parts of the targeted objects. The combination of DNN and NARS makes the TruePAL system explainable to the users. TruePAL ingests on-board cameras, radar, and other sensor signals, analyzes the environment and traffic patterns to generate timely warning to drivers and roadside crews to avoid crashes. The TruePAL team, in collaboration with the Miami/Dade Police Dept., has designed five use cases and multiple sub-scenarios in a CARLA driving simulator to test the capability of TruePAL in timely warning to the first responder drivers in potential crash scenarios. We have successfully demonstrated its capability of timely warning in over a dozen scenarios based on the use cases. The preliminary test simulation results show that TruePAL could provide the drivers and crew members advanced warning before a crash occurs.

Chow, Edward

GridCoPilot for Thermal Events: An LLM-Based Platform for Power Grid Reliability Analysis

Large Language Models show promise for translating natural language into database queries, but deploying such systems in safety-critical domains requires high reliability. We present an application of GridCoPilot to thermal event analysis (heatwaves and coldwaves) that affect power grid reliability. Our approach uses a LangChain SQL Agent to translate natural language queries into auditable SQL statements, with deterministic visualization routines that parse the structured query results. We introduce structural framing as a design principle, we integrate a NERC-region-level event library with county-level meteorology and decompose the combined data into three relational tables (event metadata, county-level event details, and a county-to-NERC subregion mapping), using prompt-guided joins to direct the model toward correct multi-table queries. For two core analytical patterns (identifying worst events by region and by region-year), the system achieved 100% SQL accuracy across all 16 NERC subregions and both event types (64 queries total). These results validate the approach for target use cases, though performance on diverse natural language formulations requires further investigation. We discuss design trade-offs, failure modes including JSON output truncation, and pathways for extending this approach to other hazard domains.

24 POWER TRANSMISSION AND DISTRIBUTION

Effects of Topography on Tropical Forest Structure Depend on Climate Context

Topography affects abiotic conditions which can influence the structure, function, and dynamics of ecological communities. An increasing number of studies have demonstrated biological consequences of fine-scale topographic heterogeneity but we have a limited understanding of how We merged high-resolution (1 sq. meter) data on topography and canopy height derived from airborne lidar with ground-based data from 15 forest plots in Puerto Rico distributed along a precipitation gradient spanning ca. 800 to 3,500 mm yr(exp -1). Ground-based data included species composition, estimated above-ground biomass (AGB), and two key functional traits (wood density and leaf mass per area, LMA) that reflect resource-use strategies and a trade-off between hydraulic safety and hydraulic efficiency. We used hierarchical Bayesian models to evaluate how the interaction between topography climate is related to metrics of forest structure (i.e., canopy height and AGB), as well as taxonomic and functional alpha- and beta-diversity. Fine-scale topography (characterized with the topographic wetness index, TWI) significantly affected forest structure and the strength (and in some cases direction) of these effects varied across the precipitation gradient. In all plots, canopy height increased with topographic wetness but the effect was much stronger in dry compared to wet forest plots. In dry forest plots, topographically wetter microsites also had higher levels of AGB but in wet forest plots, topographically drier microsites had higher AGB. Fine-scale topography influenced functional composition but had only weak or non-significant effects on taxonomic and functional alpha- and beta-diversity. For instance, community-weighted wood density followed a similar pattern to AGB across plots. We also found a marginally significant association between variation of wood density and topographic heterogeneity that depended on climate context. Synthesis: The effects of fine-scale topographic heterogeneity on tropical forest structure and composition depend on the climate context. Our study demonstrates how a stronger integration of topographic heterogeneity across precipitation gradients could improve estimates of forest structure and biomass, and may provide insight to the ways that topography might mediate species responses to drought and climate change.

Tropical dry forest

Hydrogen Leak Modeling for Development of Smart Distributed Monitoring Under Unintended Releases

Hydrogen is a versatile and clean energy carrier that can be produced from various renewable sources such as wind, solar, and hydropower. Hydrogen has the potential to play a crucial role in decarbonizing industrial processes that are currently reliant on fossil fuels and provide long-duration and/or seasonal energy storage to enable electricity decarbonization. Hydrogen can also be used as a fuel for fuel cell vehicles, providing a zero-emission alternative to traditional internal combustion engines. DOE launched the Hydrogen Energy Earthshot (Hydrogen Shot) in June 2021 to reduce the cost of clean hydrogen by 80% to $1 per 1 kilogram in 1 decade ("1 1 1"). While promising, Hydrogen is highly-flammable, and in the presence of oxygen, it can form explosive mixtures. . Therefore, understanding leak scenarios is essential to evaluate and mitigate the safety risks associated with potential hydrogen leaks. An increased understanding of leak behavior, and having tools to model leaks, can help assess how hydrogen would disperse in different environments, influencing emergency response plans and safety measures, and identify potential issues with materials and design systems that can withstand the challenges posed by hydrogen. Recently, researchers have attempted to study hydrogen leaks for development of risk management strategies. However, the focus has been on closed or semi-closed spaces like storage rooms, vehicles, garages, and fueling stations - all promising locations for future hydrogen infrastructure. In this presentation, the modeling environment extends the span of research further by modeling hydrogen leak in an outdoor, open space. We will present the key challenges with modeling hydrogen leaks in an uncontrollable environment, how they were handled, and how modeling results informed sensor selection and placement. A Hydrogen research facility at the National Renewable Energy Laboratory (NREL) was used as a case study to model hydrogen leaks. In the future, Hydrogen wide area detection methodologies will be developed and tested at this site to monitor for unintended and operational hydrogen releases. The data generated from modeling will be used to develop a predictive model to detect hydrogen leak location based on concentration measured by sensors in this open space. Furthermore, the facility was also chosen because controlled hydrogen releases can be performed. A computational fluid dynamics (CFD) based modeling approach was taken to model hydrogen leak. The full-scale hydrogen facility was modeled with a large ambient domain. The electrolyzer at the facility can produce a controlled release rate of 27 kg-H2/hr. Site-specific atmospheric and weather condition data such as wind direction, wind speed at various altitudes, and temperature were used as inputs to the model. To capture the variability of weather conditions, a subset of the weather conditions experienced during daytime hours without precipitation over the course of three months was generated; using established data clustering techniques, a total of 100 condition sets were chosen. The results show statistical distributions and ranges of hydrogen concentrations at locations throughout the domain. These distributions are compared to experimental data from a constant mass flow, controlled hydrogen release at the facility. The stochastic wind conditions of the release make direct validation difficult, therefore, statistical comparison approaches were used. Wind conditions are found to significantly impact the release behavior, including direction and concentration. Sensor selection and placement is proposed for the facility and is now based on release behavior predicted for the facility given its weather patterns; this is much more informed than without the modeling results. The methodology and analysis procedure can be translated to other facilities using modified geometries and site-specific weather conditions. Hydrogen holds great promise as a renewable energy fuel, but ensuring safety in its production, storage, and use is paramount. Studying potential leak scenarios in an open space will help develop sensors to detect hydrogen on a large spectrum of concentration and eventually build a smart distributed monitoring system.

CFD

Application of Cyber-Informed Engineering for Protecting BESS

This white paper synthesizes an array of crucial grid services provided by BESS technology, assesses its architecture and communications, and presents a case study for analysis against the principles introduced by Cyber-Informed Engineering (CIE). Furthermore, in walking through the analysis, this paper presents a framework to evaluate risks and solutions when considering BESS components. Asset owners and buyers could perform this analysis to assess their BESS product implementations, alternative inverter-based resources (IBR), and energy management systems (EMS). Battery systems fulfill various roles contingent on the unique market demands and the specific challenges presented by regional grid infrastructures. These roles also vary due to the differing utility models for ownership and operation, which are adapted to meet regional and local capabilities and requirements. Concerns have been raised regarding the potential for adversaries to exploit knowledge of battery operational patterns to orchestrate decisive attacks. However, the security of operational data for these systems may not be the primary vulnerability, as much of this information is already well-understood within the community. Applying a modest degree of subject matter expertise can often yield valuable predictions regarding how a battery will respond under certain conditions, such as grid emergencies, high or low-temperature days, Public Safety Power Shutoff (PSPS) events, and outages. The operational characteristics of batteries are well-documented, and their capabilities, including the risks associated with misoperation and the resulting consequences, are published and understood within the industry. CIE practices represent the next step in gaining functional assurance and providing an acceptable level of risk, regardless of whether a battery vendor can support a trusted and validated supply chain. While this issue has exacerbated supply chain challenges, it is not an isolated condition. This foreign supply route is the primary source of BESS for the U.S. market. Significant efforts are underway through the Bipartisan Infrastructure Law (BIL) to change that. Still, strategic short-term operational mitigations are needed to ensure the security of our operational technology (OT) systems, which are enhanced by instilling trust and are separate from vendors implementing CIE principles.

25 ENERGY STORAGE

The Evaluation of the Regional Atmospheric Modeling System in the Eastern Range Dispersion Assessment System

The Applied Meteorology Unit (AMU) evaluated the Regional Atmospheric Modeling System (RAMS) contained within the Eastern Range Dispersion Assessment System (ERDAS). ERDAS provides emergency response guidance for Cape Canaveral Air Force Station and Kennedy Space Center operations in the event of an accidental hazardous material release or aborted vehicle launch. The RAMS prognostic data are available to ERDAS for display and are used to initialize the 45th Space Wing/Range Safety dispersion model. Thus, the accuracy of the dispersion predictions is dependent upon the accuracy of RAMS forecasts. The RAMS evaluation consisted of an objective and subjective component for the 1999 and 2000 Florida warm seasons, and the 1999-2000 cool season. In the objective evaluation, the AMU generated model error statistics at surface and upper-level observational sites, compared RAMS errors to a coarser RAMS grid configuration, and benchmarked RAMS against the nationally-used Eta model. In the subjective evaluation, the AMU compared forecast cold fronts, low-level temperature inversions, and precipitation to observations during the 1999-2000 cool season, verified the development of the RAMS forecast east coast sea breeze during both warm seasons, and examined the RAMS daily thunderstorm initiation and precipitation patterns during the 2000 warm season. This report summarizes the objective and subjective verification for all three seasons.

Case, Jonathan

Crew Activity Analyzer

The crew activity analyzer (CAA) is a system of electronic hardware and software for automatically identifying patterns of group activity among crew members working together in an office, cockpit, workshop, laboratory, or other enclosed space. The CAA synchronously records multiple streams of data from digital video cameras, wireless microphones, and position sensors, then plays back and processes the data to identify activity patterns specified by human analysts. The processing greatly reduces the amount of time that the analysts must spend in examining large amounts of data, enabling the analysts to concentrate on subsets of data that represent activities of interest. The CAA has potential for use in a variety of governmental and commercial applications, including planning for crews for future long space flights, designing facilities wherein humans must work in proximity for long times, improving crew training and measuring crew performance in military settings, human-factors and safety assessment, development of team procedures, and behavioral and ethnographic research. The data-acquisition hardware of the CAA (see figure) includes two video cameras: an overhead one aimed upward at a paraboloidal mirror on the ceiling and one mounted on a wall aimed in a downward slant toward the crew area. As many as four wireless microphones can be worn by crew members. The audio signals received from the microphones are digitized, then compressed in preparation for storage. Approximate locations of as many as four crew members are measured by use of a Cricket indoor location system. [The Cricket indoor location system includes ultrasonic/radio beacon and listener units. A Cricket beacon (in this case, worn by a crew member) simultaneously transmits a pulse of ultrasound and a radio signal that contains identifying information. Each Cricket listener unit measures the difference between the times of reception of the ultrasound and radio signals from an identified beacon. Assuming essentially instantaneous propagation of the radio signal, the distance between that beacon and the listener unit is estimated from this time difference and the speed of sound in air.] In this system, six Cricket listener units are mounted in various positions on the ceiling, and as many as four Cricket beacons are attached to crew members. The three-dimensional position of each Cricket beacon can be estimated from the time-difference readings of that beacon from at least three Cricket listener units

Murray, James

Prediction of Cognitive States During Flight Simulation Using Multimodal Psychophysiological Sensing

The Commercial Aviation Safety Team found the majority of recent international commercial aviation accidents attributable to loss of control inflight involved flight crew loss of airplane state awareness (ASA), and distraction was involved in all of them. Research on attention-related human performance limiting states (AHPLS) such as channelized attention, diverted attention, startle/surprise, and confirmation bias, has been recommended in a Safety Enhancement (SE) entitled "Training for Attention Management." To accomplish the detection of such cognitive and psychophysiological states, a broad suite of sensors was implemented to simultaneously measure their physiological markers during a high fidelity flight simulation human subject study. Twenty-four pilot participants were asked to wear the sensors while they performed benchmark tasks and motion-based flight scenarios designed to induce AHPLS. Pattern classification was employed to predict the occurrence of AHPLS during flight simulation also designed to induce those states. Classifier training data were collected during performance of the benchmark tasks. Multimodal classification was performed, using pre-processed electroencephalography, galvanic skin response, electrocardiogram, and respiration signals as input features. A combination of one, some or all modalities were used. Extreme gradient boosting, random forest and two support vector machine classifiers were implemented. The best accuracy for each modality-classifier combination is reported. Results using a select set of features and using the full set of available features are presented. Further, results are presented for training one classifier with the combined features and for training multiple classifiers with features from each modality separately. Using the select set of features and combined training, multistate prediction accuracy averaged 0.64 +/- 0.14 across thirteen participants and was significantly higher than that for the separate training case. These results support the goal of demonstrating simultaneous real-time classification of multiple states using multiple sensing modalities in high fidelity flight simulators. This detection is intended to support and inform training methods under development to mitigate the loss of ASA and thus reduce accidents and incidents.

Harrivel, Angela R.

Use of Raman Spectroscopy and Delta Volume Growth from Void Collapse to Assess Overwrap Stress Gradients Compromising the Reliability of Large Kevlar/Epoxy COPVs

Composite Overwrapped Pressure Vessels (COPVs) are frequently used for storing pressurized gases aboard spacecraft and aircraft when weight saving is desirable compared to all-metal versions. Failure mechanisms in fibrous COPVs and variability in lifetime can be very different from their metallic counterparts; in the former, catastrophic stress-rupture can occur with virtually no warning, whereas in latter, a leak before burst design philosophy can be implemented. Qualification and certification typically requires only one burst test on a production sample (possibly after several pressure cycles) and the vessel need only meet a design burst strength (the maximum operating pressure divided by a knockdown factor). Typically there is no requirement to assess variability in burst strength or lifetime, much less determine production and materials processing parameters important to control of such variability. Characterizing such variability and its source is crucial to models for calculating required reliability over a given lifetime (e.g. R = 0.9999 for 15 years). In this paper we present a case study of how lack of control of certain process parameters in COPV manufacturing can result in variations among vessels and between production runs that can greatly increase uncertainty and reduce reliability. The vessels considered are 40-inch ( NASA Glenn Research center, Cleveland, OH, 44135 29,500 in3 ) spherical COPVs with a 0.74 in. thick Kevlar49/epoxy overwrap and with a titanium liner of which 34 were originally produced. Two burst tests were eventually performed that unexpectedly differed by almost 5%, and were 10% lower than anticipated from burst tests on 26-inch sister vessels similar in every detail. A major observation from measurements made during proof testing (autofrettage) of the 40-inch vessels was that permanent volume growth from liner yielding varied by a factor of more than two (150 in3 to 360 in3 ), which suggests large differences in the residual stress gradient through their overwraps. This resulted in large uncertainty in true fiber stress ratio (fiber stress at operating pressure divided by fiber stress at burst) which governs lifetime. The vessels were originally designed with tight safety margins, so it became crucial to develop a non-destructive evaluation (NDE) technique to directly measure the overwrap residual stress state of each vessel, and to identify those vessels at highest risk of having poor reliability. This paper describes a Raman Spectroscopy technique for measuring certain patterns of fluctuation in fiber elastic strains over the outside vessel surface (where all but one wrap is exposed at certain locations) that are shown to directly correlate to increased fiber stress ratios and reduced reliability.

Kezirian, Michael T.

Forward Skirt Structural Testing on the Space Launch System (SLS) Program

Structural testing was performed to evaluate heritage forward skirts from the Space Shuttle program for use on the Space Launch System (SLS) program. One forward skirt is located in each solid rocket booster. Heritage forward skirts are aluminum 2219 welded structures. Loads are applied at the forward skirt thrust post and ball assembly. Testing was needed because SLS ascent loads are roughly 40% higher than Space Shuttle loads. Testing objectives were to determine margins of safety, demonstrate reliability, and validate analytical models. Two forward skirts were structurally tested using the test configuration. The test stand applied loads to the thrust post. Four hydraulic actuators were used to apply axial load and two hydraulic actuators were used to apply radial and tangential loads. The first test was referred to as FSTA-1 (Forward Skirt Structural Test Article) and was performed in April/May 2014. The purpose of FSTA-1 was to verify the ultimate capability of the forward skirt subjected to ascent ultimate loads. Testing consisted of two liftoff load cases taken to 100% limit load followed by an ascent load case taken to 110% limit load. The forward skirt was unloaded to no load after each test case. Lastly, the forward skirt was tested to 140% limit and then to failure using the ascent loads. The second test was referred to as FSTA-2 and performed in July/August of 2014. The purpose of FSTA-2 was to verify the ultimate capability of the forward skirt subjected to liftoff ultimate loads. Testing consisted of six liftoff load cases taken to 100% limit load followed by the six liftoff cases taken to 140% limit load. Two ascent load cases were then tested to 100% limit load. The forward skirt was unloaded to no load after each test case. Lastly, the forward skirt was tested to 140% limit and then to failure using the ascent loads. The forward skirts on FSTA-1 and FSTA-2 successfully carried all applied liftoff and ascent load cases. Both FSTA-1 and FSTA-2 were tested to failure by increasing the ascent loads. Failure occurred in the forward skirt thrust post radius. The forward skirts on FSTA-1 and FSTA-2 had nearly identical failure modes. FSTA-1 failed at 1.72 times limit load and FSTA-2 failed at 1.62 times limit load. This difference is primarily attributed to variation in material properties in the thrust post region. Test data were obtained from strain gages, deflection gages, ARAMIS digital strain measurement, acoustic emissions, and high-speed video. Strain gage data and ARAMIS strain were compared to finite element (FE) analysis predictions. Both the forward skirt and tooling were modeled. This allows the analysis to simulate the loading as close as possible to actual test configuration. FSTA-1 and FSTA-2 were instrumented with over 200 strain gages to ensure all possible failure modes could be captured. However, it turned out that three gages provided critical strain data. One was located in the post bore and two on the post radius. More gages were not specified due to space limitations and the desire to not interfere with the use of the ARAMIS system on the post radius. Measured strains were compared to analysis results for the load cycle to failure. Note that FSTA-1 gages were lost before failure was reached. FSTA-2 gages made it to the failure load but one of the radius gages was lost before testing began. This gage was not replaced because of the time and cost associated with disassembly of the test structure. Correlation to analysis was excellent for FSTA-1. FSTA-2 was not quite as good because there was more residual strain from previous load cycles. FSTA-2 was loaded and unloaded with 12 liftoff cases and two ascent cases before taking the skirt to failure. FSTA-1 only had two liftoff cases and one ascent case before taking the skirt to failure. The ARAMIS system was used to determine strain at the post radius by processing digital images of a speckled paint pattern. Digital cameras recorded images of the speckled paint pattern. ARAMIS strain results for FSTA-2 just prior to failure. Note a high strain location develops near the left side. This high strain compares well to analysis prediction for both FSTA-1 and FSTA-2. The strain at this location was also plotted versus limit load. Both FSTA-1 and FSTA-2 had excellent correlation between ARAMIS and analysis strains. Acoustic emission (AE) sensors were used to monitor for damage formation that may occur during testing (e.g., crack formation and growth or propagation). AE was very important because after disassembly of FSTA-1, a crack was observed in the ball fitting radius. The ball fitting did not crack on FSTA-2. AE data was used to reconstruct when the crack occurred. The AE energy versus time plot for FSTA. The energy increased considerably at 850 seconds (152% limit load), indicating a crack could have formed at this point. The only visual evidence found that could have corresponded to this was the crack that initiated in the ball fitting. The cracks in the forward skirt aluminum structures would likely have been lower energy due to a lower modulus and all that were found after failure correlated to occurring after the initial crack in the post radius. This was verified by high-speed cameras used to record the failure.

Lohrer, J. D.

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani