Search NASASearch

SEARCH · Search NASA

Results for “transformer based forecasting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han

Geothermal well testing pressure prediction by using a hybrid transformer model system: FORGE well use case

Geothermal has huge potential to become an indispensable component in achieving the goal of sustainable energy economy, given its capability to provide consistent baseload power to the electric grid. Injection tests are crucial in geothermal energy system as they naturally help to evaluate reservoir properties, understand fluid flow and even enhance reservoir performance. In this research, we developed a hybrid model system that integrates machine learning (ML) regression, a physics-based mathematical model, and transformer deep learning. Trained and validated using FORGE injection test dataset, this system can forecast the pressure variations both upward and downward over time. The pressure prediction achieved prediction accuracy within 3-6% variance of true pressure values. The system can significantly save time and reduce costs by testing only a few cycles and then using model predictions for further analysis, instead of conducting additional real injection cycle tests. The developed model system also holds promise for designing injection test processes and maintaining well production in geothermal energy. Presented at the IMAGE ‘25 Conference led by Shell.

FORGE

Radiometric Modeling and Calibration of the Geostationary Imaging Fourier Transform Spectrometer (GIFTS)Ground Based Measurement Experiment

The ultimate remote sensing benefits of the high resolution Infrared radiance spectrometers will be realized with their geostationary satellite implementation in the form of imaging spectrometers. This will enable dynamic features of the atmosphere s thermodynamic fields and pollutant and greenhouse gas constituents to be observed for revolutionary improvements in weather forecasts and more accurate air quality and climate predictions. As an important step toward realizing this application objective, the Geostationary Imaging Fourier Transform Spectrometer (GIFTS) Engineering Demonstration Unit (EDU) was successfully developed under the NASA New Millennium Program, 2000-2006. The GIFTS-EDU instrument employs three focal plane arrays (FPAs), which gather measurements across the long-wave IR (LWIR), short/mid-wave IR (SMWIR), and visible spectral bands. The GIFTS calibration is achieved using internal blackbody calibration references at ambient (260 K) and hot (286 K) temperatures. In this paper, we introduce a refined calibration technique that utilizes Principle Component (PC) analysis to compensate for instrument distortions and artifacts, therefore, enhancing the absolute calibration accuracy. This method is applied to data collected during the GIFTS Ground Based Measurement (GBM) experiment, together with simultaneous observations by the accurately calibrated AERI (Atmospheric Emitted Radiance Interferometer), both simultaneously zenith viewing the sky through the same external scene mirror at ten-minute intervals throughout a cloudless day at Logan Utah on September 13, 2006. The accurately calibrated GIFTS radiances are produced using the first four PC scores in the GIFTS-AERI regression model. Temperature and moisture profiles retrieved from the PC-calibrated GIFTS radiances are verified against radiosonde measurements collected throughout the GIFTS sky measurement period. Using the GIFTS GBM calibration model, we compute the calibrated radiances from data collected during the moon tracking and viewing experiment events. From which, we derive the lunar surface temperature and emissivity associated with the moon viewing measurements.

Tian, Jialin

Sub-Seasonal Forecasting of the Stratospheric Wave Events, Sudden Stratospheric Warmings, and Their Influence on the Troposphere

Stratospheric wave events and major sudden stratospheric warming (SSW) events are well captured in high-resolution global forecasts out to 10 days. Tropospheric influences of SSW include statistically significant shifts in the storm tracks and associated surface temperature and precipitation pattern changes. Since these events can in turn influence the troposphere on time scales of 30-60 days the ability to predict these events and the subsequent long-term response on time scales beyond 10 days is of interest. Here we examine the prediction of stratospheric wave events and their evolution using the NASA GMAO (Global Modeling and Assimilation Office) Sub-Seasonal to Seasonal (S2S) system. This is a recently released subseasonal to seasonal forecast system, GEOS-S2S version 2.1. Compared to GMAO's previous system, the new version runs at higher atmospheric resolution (approximately 1/2 degree globally), contains a substantially improved model of the cryosphere, includes additional interactive aerosol model components, and the ocean data assimilation system has been replaced with a Local Ensemble Transform Kalman Filter. Results are based on a comprehensive series of hindcasts starting from the year 2000. They show that while the S2S system is not as accurate at 10 days in forecasting major SSW events as the NASA GMAO Forward Processing system, it can usefully predict stratospheric anomalies out to 20 days and the subsequent stratospheric/tropospheric evolution beyond 30 days.

Coy, L.

SPECTER: an instrument concept for CMB spectral distortion measurements with enhanced sensitivity

Deviations of the cosmic microwave background (CMB) energy spectrum from a perfect blackbody uniquely probe a wide range of physics, ranging from fundamental physics in the primordial Universe (μ-distortion) to late-time baryonic feedback processes (y-distortion). While the y-distortion can be detected with a moderate increase in sensitivity over that of COBE/FIRAS, the ΛCDM-predicted μ-distortion is roughly two orders of magnitude smaller and requires substantial improvements, with foregrounds presenting a serious obstacle. Within the standard model, the dominant contribution to μ arises from energy injected via Silk damping, yielding sensitivity to the primordial power spectrum at wavenumbers k ≈ 1-10 4 Mpc -1 . Here, we present a new instrument concept, SPECTER, with the goal of robustly detecting μ. The instrument technology is similar to that of LiteBIRD, but with an absolute temperature calibration system. Using a Fisher approach, we optimize the instrument's configuration to target μ while marginalizing over foreground contaminants. Unlike Fourier-transform-spectrometer-based designs, the specific bands and their individual sensitivities can be independently set in this instrument, allowing significant flexibility. We forecast SPECTER to observe the ΛCDM-predicted μ-distortion at ≈ 5σ (10σ) assuming an observation time of 1 (4) year(s) (corresponding to mission duration of 2 (8) years), after foreground marginalization. Our optimized configuration includes 16 bands spanning 1–2000 GHz with ∼degree-scale angular resolution at ∼ 150 GHz and 1100 total detectors. SPECTER will additionally measure the y-distortion at sub-percent precision and its relativistic correction at percent-level precision, yielding tight constraints on the total thermal energy and mean temperature of ionized gas.

CMBR experiments

A Neural Optimizer With Decision-Focused Learning for Optimal Energy Storage Operation

Here, this article introduces a neural optimizer-based framework for optimizing battery energy storage system (BESS) control for grid services, including demand charge and energy cost reduction. By leveraging decision-focused learning (DFL), the proposed framework ensures seamless integration and adaptation, significantly enhancing control performance. A patch time-series transformer is employed for peak load forecasting, incorporating aleatoric uncertainty quantification to account for forecasting uncertainties within the decision-making process. The framework utilizes a solver-in-the-loop approach to generate optimal BESS actions, which are then used to train the neural optimizer-based agent. By co-optimizing both BESS operational modes and output power within the NN, the system achieves improved performance and robustness. After initial training, the forecasting and control models are jointly fine-tuned to account for forecasting errors, further improving decision precision and efficiency through DFL. Case studies are performed to validate the performance of the framework using multiple real-world datasets, demonstrating superior performance in monthly peak load forecasting compared to state-of-the-art models. In addition, the results are compared against existing decision-making approaches. The results demonstrate a reduction in monthly peak forecasting error by approximately 15% across various performance measures and achieve an optimization gap for BESS operation that is about three times smaller compared to existing methods.

Kim, Hyeonjin [Pacific Northwest National Laborato

Expanding the application of soil moisture monitoring systems through regression-based transformation

Relative to other geophysical variables, soil moisture (SM) estimates derived from land surface models (LSMs) and land data assimilation systems (LDAS) are difficult to transfer between platforms and applications. This difficulty stems from the highly model-dependent nature of LSM SM estimates and differences in the vertical support of discretized SM values. As a result, operational SM estimates generated by one LSM (or LDAS) cannot generally be directly applied to a hydrologic monitoring or forecast system designed around a second LSM. This lack of transferability is particularly problematic for LDAS applications, where the time, expertise, and computational resources required to generate an operational LDAS analysis cannot be practically duplicated for every LSM-specific application. Here, we develop a set of simple regression tools for translating SM estimates between LSMs and multiple LDAS analyses. Results demonstrate that simple multivariate linear regression - utilizing independent variables based on multi-layer and temporally lagged SM estimates - can significantly improve upon baseline transformation approaches using direct percentile matching. The proposed regression approaches are effective for both the LSM-to-LSM and LDAS-to-LDAS transformation of multi-layer SM percentiles. Application of this approach will expand the utility of existing, high-quality (but LSM-specific) operational sources of SM information like the NASA Soil Moisture Active Passive Level-4 Soil Moisture product.

Soil Moisture

Short-Term Probabilistic Solar Forecasting via Reinforcement Learning over ECMWF

In this paper, we present an innovative reinforcement learning approach for short-term solar forecasting, leveraging data from the European Centre for Medium-Range Weather Forecasts (ECMWF). The methodology begins with the application of the System Advisor Model (SAM) to transform various ECMWF numerical weather prediction members into predictive photovoltaic power generation. To enhance the precision of deterministic forecasting, we introduce a dynamic model selection algorithm based on Q-learning. This algorithm dynamically identifies and utilizes the most accurate ensemble member for forecasting purposes. Furthermore, we employ a support vector regression surrogate model with a Gaussian distribution to generate probabilistic forecasts, providing a holistic view of solar energy generation uncertainty. To expedite the training process and make it more practical for real-world applications, we integrate a rolling update workflow. This innovative workflow reduces the training period from months to a mere 19 days, making our method highly efficient. Numerical results of the case study show that in comparison to benchmark models, the proposed method improves the deterministic and probabilistic solar forecasting accuracy by up to 40.84% and 48.42%, respectively.

ensemble forecasting

Autoregressive long-horizon prediction of plasma edge dynamics *

Accurate modeling of scrape-off layer (SOL) and divertor-edge dynamics is vital for designing plasma-facing components in fusion devices. High-fidelity edge fluid/neutral codes such as SOLPS-ITER capture SOL physics with high accuracy, but their computational cost limits broad parameter scans and long transient studies. We present transformer-based, autoregressive surrogates for efficient prediction of 2D, time-dependent plasma edge state fields. Trained on SOLPS-ITER spatiotemporal data for the KSTAR tokamak, the surrogates forecast electron temperature, electron density, and radiated power over extended horizons. We evaluate model variants trained with increasing autoregressive horizons (1–100 steps) on short- and long-horizon prediction tasks. Longer-horizon training systematically improves rollout stability and mitigates error accumulation, enabling stable predictions over hundreds to thousands of steps and reproducing key dynamical features such as the motion of high-radiation regions. Measured end-to-end wall-clock times show the surrogate is orders of magnitude faster than SOLPS-ITER, enabling rapid parameter exploration. Prediction accuracy degrades when the surrogate enters physical regimes not represented in the training dataset, motivating future work on data enrichment and physics-informed constraints. Overall, this approach provides a fast, accurate surrogate for computationally intensive plasma edge simulations, supporting rapid scenario exploration, control-oriented studies, and progress toward real-time applications in fusion devices.

autoregressive deep learning

AIRS Retrieval Validation During the EAQUATE

Atmospheric and surface thermodynamic parameters retrieved with advanced hyperspectral remote sensors of Earth observing satellites are critical for weather prediction and scientific research. The retrieval algorithms and retrieved parameters from satellite sounders must be validated to demonstrate the capability and accuracy of both observation and data processing systems. The European AQUA Thermodynamic Experiment (EAQUATE) was conducted mainly for validation of the Atmospheric InfraRed Sounder (AIRS) on the AQUA satellite, but also for assessment of validation systems of both ground-based and aircraft-based instruments which will be used for other satellite systems such as the Infrared Atmospheric Sounding Interferometer (IASI) on the European MetOp satellite, the Cross-track Infrared Sounder (CrIS) from the NPOESS Preparatory Project and the following NPOESS series of satellites. Detailed inter-comparisons were conducted and presented using different retrieval methodologies: measurements from airborne ultraspectral Fourier transform spectrometers, aircraft in-situ instruments, dedicated dropsondes and radiosondes, and ground based Raman Lidar, as well as from the European Center for Medium range Weather Forecasting (ECMWF) modeled thermal structures. The results of this study not only illustrate the quality of the measurements and retrieval products but also demonstrate the capability of these validation systems which are put in place to validate current and future hyperspectral sounding instruments and their scientific products.

Zhou, Daniel K.

Using a Simple Water Balance Framework to Quantify the Impact of Soil Moisture Initialization on Subseasonal Evapotranspiration and Air Temperature Forecasts

Past studies have shown that accurate soil moisture initialization can contribute significant skill to near-surface air temperature (T2M) forecasts at subseasonal leads. The mechanisms by which soil moisture contributes such skill are examined here with a simple water balance-based model that captures the essence of soil moisture behavior in a state-of-the-art subseasonal-to-seasonal (S2S) forecasting system. The simple model successfully transforms initial soil moisture contents into average “forecasted” ET values at 16-30 day lead that agree well, during summer, with the values forecasted by the full NASA GEOS S2S system, indicating that soil moisture initialization dominates over forecasted meteorology in determining ET fluxes at subseasonal leads. When the simple model’s ET anomalies are interpreted in terms of T2M anomalies, a similar conclusion is reached for T2M: soil moisture initialization explains much (about 50% in the eastern half of the continental US) of the T2M anomalies produced by the full GEOS S2S system at 16-30 day lead, and the T2M forecasts produced by the simple model capture about half of the skill attained by the full system. The simple model’s framework is particularly conducive to an analysis of uncertainty in forecasts. Drier soils are generally found to induce larger uncertainty in ET (and thus T2M) forecasts, a result linked to the functional form relating ET to soil moisture in the simple model and verified by an analysis of the ensemble spreads within the forecasts produced by the full GEOS S2S system

Randal D Koster

Sensitivity-based voltage constraints for optimal power flow in low-voltage distribution feeders

The optimal power flow (OPF) problem for distribution systems can include network details down to the low-voltage (LV) points of interconnection of individual customers. This paper addresses the implementation of voltage magnitude constraints, and sets forth a practicable approach for capturing the effects on voltage from the switching behavior of loads (e.g., heat pumps, air conditioners, water heaters, or pool pumps) and from the variability of renewable generation (e.g., rooftop solar). The proposed method adjusts the OPF voltage constraints based on forecasts of load and generation upper and lower bounds, in conjunction with sensitivity factors derived from the power flow equations. An illustrative OPF formulation is also provided, which incorporates transformer models that include core loss. We demonstrate that accurate modeling of these LV network components is critical to avoid voltage violations at customer points of interconnection. Furthermore, the ideas are validated through numerical case studies on a realistic distribution feeder.

24 POWER TRANSMISSION AND DISTRIBUTION

The Relation Between Atmospheric Humidity and Temperature Trends for Stratospheric Water

We analyze the relation between atmospheric temperature and water vapor-a fundamental component of the global climate system-for stratospheric water vapor (SWV). We compare measurements of SWV (and methane where available) over the period 1980-2011 from NOAA balloon-borne frostpoint hygrometer (NOAA-FPH), SAGE II, Halogen Occultation Experiment (HALOE), Microwave Limb Sounder (MLS)/Aura, and Atmospheric Chemistry Experiment Fourier Transform Spectrometer (ACE-FTS) to model predictions based on troposphere-to-stratosphere transport from ERA-Interim, and temperatures from ERA-Interim, Modern Era Retrospective-Analysis (MERRA), Climate Forecast System Reanalysis (CFSR), Radiosonde Atmospheric Temperature Products for Assessing Climate (RATPAC), HadAT2, and RICHv1.5. All model predictions are dry biased. The interannual anomalies of the model predictions show periods of fairly regular oscillations, alternating with more quiescent periods and a few large-amplitude oscillations. They all agree well (correlation coefficients 0.9 and larger) with observations for higherfrequency variations (periods up to 2-3 years). Differences between SWV observations, and temperature data, respectively, render analysis of the model minus observation residual difficult. However, we find fairly well-defined periods of drifts in the residuals. For the 1980s, model predictions differ most, and only the calculation with ERA-Interim temperatures is roughly within observational uncertainties. All model predictions show a drying relative to HALOE in the 1990s, followed by a moistening in the early 2000s. Drifts to NOAA-FPH are similar (but stronger), whereas no drift is present against SAGE II. As a result, the model calculations have a less pronounced drop in SWV in 2000 than HALOE. From the mid-2000s onward, models and observations agree reasonably, and some differences can be traced to problems in the temperature data. These results indicate that both SWV and temperature data may still suffer from artifacts that need to be resolved in order to answer the question whether the large-scale flow and temperature field is sufficient to explain water entering the stratosphere.

Fueglistaler, S.

EV Forecasting-Based Model Predictive Control for Distribution System Congestion Mitigation

The uncoordinated charging of electric vehicles (EVs) in time and space brings congestion issues to the distribution network. This paper proposes an EV charging demand forecasting-based model predictive control (MPC) method for distribution system congestion management. To effectively forecast the time-series EV station charging demand, a hybrid forecasting model that integrates the long short-term memory network (LSTM) and Transformer is proposed. The Transformer-LSTM model is trained using a one-year real historical charging dataset of EV stations to forecast future charging demand in 15-minute intervals. This informs the MPC for distribution network congestion management and minimization of PV curtailment. Numerical results carried out on the modified IEEE 123-bus distribution system demonstrate that the proposed method can effectively resolve line congestion issues through EV smart charging and PV curtailment while outperforming other benchmarks.

ADVANCED PROPULSION SYSTEMS,SOLAR ENERGY

A Scalable Real-Time Data Assimilation Framework for Predicting Turbulent Atmosphere Dynamics

AI-based foundation models like FourCastNet, GraphCast are revolutionizing weather and climate predictions but are not yet ready for operational use. Their limitation lies in the absence of a data assimilation system to incorporate real-time Earth system observations, crucial for accurately forecasting events like tropical cyclones. To overcome these obstacles, we introduce a generic real-time data assimilation framework and demonstrate its end-to-end performance on the Frontier supercomputer. This framework comprises two primary modules: an ensemble score filter (EnSF), which significantly outperforms the state-of-the-art data assimilation method, and a vision transformer-based surrogate capable of real-time adaptation through the integration of observational data. We demonstrate both the strong and weak scaling of our framework up to 1024 GPUs on the Exascale supercomputer, Frontier. Our results not only illustrate the framework's exceptional scalability on high-performance computing systems, but also demonstrate the importance of supercomputers in real-time data assimilation for weather and climate predictions.

Lu, Dan

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING

On the effectiveness of neural operators at zero-shot weather downscaling

Machine-learning (ML) methods have shown great potential for weather downscaling. These data-driven approaches provide a more efficient alternative for producing high-resolution weather datasets and forecasts compared to physics-based numerical simulations. Neural operators, which learn solution operators for a family of partial differential equations, have shown great success in scientific ML applications involving physics-driven datasets. Neural operators are grid-resolution-invariant and are often evaluated on higher grid resolutions than they are trained on, i.e., zero-shot super-resolution. Given their promising zero-shot super-resolution performance on dynamical systems emulation, we present a critical investigation of their zero-shot weather downscaling capabilities, which is when models are tasked with producing high-resolution outputs using higher upsampling factors than are seen during training. To this end, we create two realistic downscaling experiments with challenging upsampling factors (e.g., 8x and 15x) across data from different simulations: the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) and the Wind Integration National Dataset Toolkit. While neural operator-based downscaling models perform better than interpolation and a simple convolutional baseline, we show the surprising performance of an approach that combines a powerful transformer-based model with parameter-free interpolation at zero-shot weather downscaling. We find that this Swin-Transformer-based approach mostly outperforms models with neural operator layers in terms of average error metrics, whereas an Enhanced Super-Resolution Generative Adversarial Network-based approach is better than most models in terms of capturing the physics of the ground truth data. We suggest their use in future work as strong baselines.

17 WIND ENERGY

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani