Search NASA⌕ Search

SEARCH · Search NASA

Results for “Factorization machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Large-Scale Visualization of 3D Unstructured Groundwater Model Using Cave Automated Virtual Environment

The immersive three-dimensional (3D) virtual reality (VR) visualization of groundwater models allows us to deepen our understanding of aquifer systems and provide better solutions to present groundwater-related problems, such as groundwater recharge, water quality, and sustainability. Visualization assists in accurately developing groundwater models and revealing important subsurface features, including faulting, folding, and unconformity. However, assessing model accuracy poses challenges due to the complexity of geology and groundwater systems. This research demonstrates a workflow to visualize and analyze raw 3D unstructured groundwater model data using an immersive Cave Automated Virtual Environment (CAVE). To visualize the unstructured groundwater model data, the raw dataset is converted into interactive CAVE-compatible formats utilizing a set of tools: ParaView, Blender, and Unity. This enables researchers to immerse themselves in the data, identifying influential patterns and relationships. e resulting insights can inform the development of sophisticated machine-learning models for groundwater level prediction. The CAVE’s immersive capabilities allow intuitive exploration from various perspectives, providing a more holistic understanding of the factors affecting groundwater levels. These insights are crucial to improve predictive models. The CAVE results also facilitate collaborative analysis and have potential applications in training and education. is research demonstrates the value of immersive VR tools such as the CAVE for unraveling intricacies within high-dimensional scientific data to drive real-world forecasting and modeling applications.

54 ENVIRONMENTAL SCIENCES↗

On the Physical Nature of Lyα Transmission Spikes in High-redshift Quasar Spectra

We investigate Lyman-alpha (Lyα) transmission spikes at 5.2 < z < 6.8 using synthetic quasar spectra from the “Cosmic Reionization on Computers” simulations. We focus on understanding the relationship between these spikes and the properties of the intergalactic medium (IGM). Disentangling the complex interplay between IGM physics and the influence of galaxies on the generation of these spikes presents a significant challenge. To address this, we employ Explainable Boosting machines, an interpretable machine learning algorithm, to quantify the relative impact of various IGM properties on the Lyα flux. Our findings reveal that gas density is the primary factor influencing absorption strength, followed by the intensity of background radiation and the temperature of the IGM. Ionizing radiation from local sources (i.e., galaxies) appears to have a minimal effect on Lyα flux. The simulations show that transmission spikes predominantly occur in regions of low gas density. Our results challenge recent observational studies suggesting the origin of these spikes in regions with enhanced radiation. We demonstrate that Lyα transmission spikes are largely a product of the large-scale structure, of which galaxies are biased tracers.

79 ASTRONOMY AND ASTROPHYSICS↗

A 200-kW wind turbine generator conceptual design study

A conceptual design study was conducted to define a 200 kW wind turbine power system configuration for remote applications. The goal was to attain an energy cost of 1 to 2 cents per kilowatt-hour at a 14-mph site (mean average wind velocity at an altitude of 30 ft.) The costs of the Clayton, New Mexico, Mod-OA (200-kW) were used to identify the components, subsystems, and other factors that were high in cost and thus candidates for cost reduction. Efforts devoted to developing component and subsystem concepts and ideas resulted in a machine concept that is considerably simpler, lighter in weight, and lower in cost than the present Mod-OA wind turbines. In this report are described the various innovations that contributed to the lower cost and lighter weight design as well as the method used to calculate the cost of energy.

Source record↗

Auralization of Tonal Rotor Noise Components of a Quadcopter Flyover

The capabilities offered by small unmanned vertical lift aerial vehicles, for example, quadcopters, continue to captivate entrepreneurs across the private, public, and civil sectors. As this industry rapidly expands, the public will be exposed to these devices (and to the noise these devices generate) with increasing frequency and proximity. Accordingly, an assessment of the human response to these machines will be needed shortly by decision makers in many facets of this burgeoning industry, from hardware manufacturers all the way to government regulators. One factor of this response is that of the annoyance to the noise that is generated by these devices. This paper presents work currently being pursued by NASA toward this goal. First, physics-based (CFD) predictions are performed on a single isolated rotor typical of these devices. The result of these predictions are time records of the discrete tonal components of the rotor noise. These time records are calculated for a number of points that appear on a lattice of locations spread over the lower hemisphere of the rotor. The source noise is then generated by interpolating between these time records. The sound from four rotors are combined and simulated-propagation techniques are used to produce complete flyover auralizations.

Christian, Andrew W.↗

Machine Learning Algorithms for Aerosol and Cloud Detection Using CATS on the ISS

Clouds and aerosols are one of the largest uncertainties in understanding and forecasting the Earth’s changing climate system. The type and height of aerosols are important factors in determining the top-of-atmosphere (TOA) radiation budget, either direct reflection of solar radiation back to space and/or absorption of solar radiation. In addition to their impact on the Earth’s climate system, aerosols near the surface from wildfires, man-made pollution events, and dust storms are hazardous to human health. The phase and height of clouds also play a critical role in determining the role of clouds in the Earth’s climate system. Cirrus clouds in the upper troposphere can induce a significant daytime TOA warming effect, while liquid water clouds near the surface cause a large corresponding cooling effect. Lidar measurements provide accurate vertically resolved information about clouds and aerosols, including complex multi-layer scenes where passive sensors are challenged and at night, when passive sensors are unable to measure cloud and aerosol properties. The Cloud-Aerosol Transport System (CATS) is a lidar instrument that operated for 33 months on the International Space Station (ISS) at the 1064 nm wavelength to measure attenuated total backscatter and depolarization ratio. These fundamental measurements are used to derive “vertical feature mask” cloud and aerosol products, including layer top/base heights, layer geometrical thickness, aerosol type, and cloud phase. While space-based lidar systems like CATS provide cloud and aerosol vertical distributions that improve our understanding of the climate system, averaging of the daytime data from these sensors is required, at the expense of spatial resolution, to improve the daytime signal-to noise (SNR) and thus atmospheric layer detection. This presentation shows results from machine learning (ML) techniques that, when applied to CATS data: 1. improve the 1064 nm SNR 2. enable detection of atmospheric features during daytime with a horizontal resolution of 350 m or 5 km (compared to the 60 km required for standard CATS data products) 3. increase the number of atmospheric layers detected in the CATS data. A Convolutional Neural Network (CNN) trained using CATS standard data products also demonstrated the potential for improved cloud-aerosol discrimination, cloud phase, and aerosol typing compared to the operational CATS algorithms for cloud edges and complex near-surface scenes during daytime. The ML tools described in this paper can facilitate the development of smaller, low-cost lidar systems in the future and enable real-time accessibility of lidar data products from future lidar systems for monitoring and forecasting of hazardous events.

John Yorks↗

Improvement of Radiative Fluxes for the CERES FluxByCldTyp Data Product Based on Machine Learning Technique

The NASA Clouds and the Earth's Radiant Energy System (CERES) product provides over 20 years of accurately observed top-of-the-atmosphere and surface flux data for climate studies. The interaction between clouds and radiation interaction is a key factor that dominate climate feedbacks but is not well understood. To further advance our understanding of the cloud-radiation interaction, a new CERES FluxByCldTyp (FBCT) product has been developed that contains radiative fluxes by cloud-type, which can provide more stringent constraints when validating models. The FBCT product utilizes Moderate Resolution Imaging Spectroradiometer (MODIS) narrow-band (NB) imager channel radiances partitioned by cloud-type within a CERES footprint to estimate their broadband fluxes. The MODIS multi-channel derived broadband fluxes were compared with the CERES observed footprint fluxes and were found to be within 1% and 2.5% for LW and SW, respectively, as well as being mostly free of cloud property dependencies. The FBCT all-sky and clear-sky monthly averaged fluxes were found to be consistent with the CERES SSF1deg product. This study takes advantage of recent progress in machine learning (ML) field by applying deep neural network algorithm to improve fluxes based on MODIS NB radiances. The preliminary study shows ML produce are an improvement over the current FBCT Edition 4 NB2BB algorithm. Furthermore, unlike Ed4 NB2BB, the new ML method convert NB radiances directly to broadband fluxes. For future Ed5, new NB radiances are proposed and used by ML to improve fluxes calculation. Preliminary results show significant LW improvement.

Sun, Moguo↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

Presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Handbook of estimating data, factors, and procedures

Elements to be considered in estimating production costs are discussed in this manual. Guidelines, objectives, and methods for analyzing requirements and work structure are given. Time standards for specific specfic operations are listed for machining, sheet metal working, electroplating and metal treating; painting; silk screening, etching and encapsulating; coil winding; wire preparation and wiring; soldering; and the fabrication of etched circuits and terminal boards. The relation of the various elements of cost to the total cost as proposed for various programs by various contractors is compared with government estimates.

Freeman, L. M.↗

Climatic and Socioeconomic Drivers of Water Use and Their Spatio‐Temporal Patterns for Small and Mid‐Sized Cities in the Contiguous United States

This study explores the drivers of urban water use and their spatial-temporal patterns in 142 small and mid-sized cities across the Contiguous United States (CONUS) by analyzing the data directly collected from these cities and using advanced machine learning techniques. We identify five distinguished clusters across CONUS, each showing unique trends of the impact of drivers on water use. We find that socioeconomic factors significantly influence water use in eastern and southwestern cities, while climatic variables such as precipitation and temperature range dominate in central and northwestern regions. Temporal analysis reveals the impacts of major socioeconomic and climatic disruptions on urban water use in the period 2011–2021, including the COVID lockdown, the rapid growth of data centers, and the drought of 2012. In addition, our analysis suggests that economic growth in small and mid-sized US cities continues to be accompanied by rising water use, contrasting with the opposite trend observed in large cities in prior studies. This implies that as smaller cities develop, their water use may increase above current levels until incomes reach a higher threshold, highlighting the need to improve water use efficiency. This study also presents useful insights for developing effective water demand management strategies in response to climatic variability and socioeconomic growth in small and mid-sized cities.

54 ENVIRONMENTAL SCIENCES↗

Alternating bending-steady torque fatigue reliability

Results generated by three unique fatigue reliability research machines which can apply alternating-bending loads combined with steady torque are presented. Six-inch long, AISI steel, grooved specimens with a stress concentration factor of 1.42 and Rockwell C 35/40 hardness were subjected to various combinations of these loads and cycled to failure. The generated cycles-to-failure and staircase-testing data are statistically analyzed to develop distributional S-N and Goodman diagrams. Various failure theories are investigated to determine which one best represents the data. The effect of the groove and of the various combined bending-torsion loads on the finite and endurance life strength of such components, as well as on the Goodman diagram, are determined. Design applications are presented.

Kececioglu, D.↗

Discussion and theoretical summarization of the experimental data

A summary of research on psychological factors that cause substantial changes in the reliability indicators of an operators work is followed by a conclusion that strong moral-volitional qualities are the basic factors that make the human behavior under conditions of stress effective; emotional subcortical subdominants affect a person's conscious organization and self control in a man machine environment.

Mileryan, Y. A.↗

Combined bending-torsion fatigue reliability of AISI 4340 steel shafting with K sub t = 2.34

Results generated by three, unique fatigue reliability research machines which can apply reversed bending loads combined with steady torque are presented. Six-inch long, AISI 4340 steel, grooved specimens with a stress concentration factor of 2.34 and R sub C 35/40 hardness were subjected to various combinations of these loads and cycled to failure. The generated cycles-to-failure and stress-to-failure data are statistically analyzed to develop distributional S-N and Goodman diagrams. Various failure theories are investigated to determine which one represents the data best. The effect of the groove and of the various combined bending-torsion loads on the S-N and Goodman diagrams are determined. Three design applications are presented. The third one illustrates the weight savings that may be achieved by designing for reliability.

Kececioglu, D.↗

First concept for a tropical area monitoring project

The first concept of a tropical area monitoring project is presented. The project would develop an operational system capable of monitoring land areas by machine processing of satellite data. LANDSAT images would be processed within a controlled isolable unit to detect changes in forest cover, rangeland, soil integrity, and other factors important to conservation of tropical ecology. An introductory developmental effort is described to demonstrate the use of LANDSAT data in this application. The independent unit, which functions as a user organization within the development project, assures that the technology will be transferable to a user organization through well defined, easily monitored interfaces with the rest of the world.

Source record↗

Producibility aspects of advanced composites for an L-1011 Aileron

The design of advanced composite aileron suitable for long-term service on transport aircraft includes Kevlar 49 fabric skins on honeycomb sandwich covers, hybrid graphite/Kevlar 49 ribs and spars, and graphite/epoxy fittings. Weight and cost savings of 28 and 20 percent, respectively, are predicted by comparison with the production metallic aileron. The structural integrity of the design has been substantiated by analysis and static tests of subcomponents. The producibility considerations played a key role in the selection of design concepts with potential for low-cost production. Simplicity in fabrication is a major factor in achieving low cost using advanced tooling and manufacturing methods such as net molding to size, draping, forming broadgoods, and cocuring components. A broadgoods dispensing machine capable of handling unidirectional and bidirectional prepreg materials in widths ranging from 12 to 42 inches is used for rapid layup of component kits and covers. Existing large autoclaves, platen presses, and shop facilities are fully exploited.

Van Hamersveld, J.↗

The human role in space (THURIS)

An overview of the human role in space station activities is presented. Associated factors such as performance cost, and risk are discussed. Benefits gained from previous successful manned space missions are highlighted. Human qualifications and capabilities associated with man machine systems are explored. Candidate procedures to be carried out by extravehicular activity spacecrews are described.

Wolbers, H. L.↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a 3D unstructured-grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data-structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which ease the implementation of computational problems on parallel architecture machines by relieving the user of the low-level machine specific issues. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗