Search NASA⌕ Search

SEARCH · Search NASA

Results for “preprocessing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Using Earth Observations to Analyze Vegetation Phenology and Climatology in Bhutan to Identify Forest Disturbance

Changes in climate in the Himalayan region cause variability in temperature, precipitation, and phenology. It can also impact the health of coniferous forest ecosystems including increased damage due to aggressive forest pests. Forest disturbance from bark beetle is a major concern in Bhutan, sometimes causing extensive tree mortality to pine and spruce forests. Climatological trends and changes in vegetation phenology were analyzed and incorporated into a tool in Google Earth Engine that identified patches of forest disturbance in Bhutan. Preprocessed phenology and meteorological data from the Advanced Very High-Resolution Radiometer (AVHRR) and Terra and Aqua Moderate Resolution Imaging Spectroradiometer (MODIS), along with Climate Hazards Center Infrared Precipitation with Station (CHIRPS), Famine Early Warning System Network Land Data Assimilation System (FLDAS), Sentinel-2 Multispectral Instrument (MSI), and Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper (ETM)+, and Landsat 8 Operational Land Imager (OLI) were used within the tool. Changes in temperature, precipitation, and phenology were analyzed throughout Bhutan, and forest disturbance caused by bark beetle was investigated in two districts of the country

Tashi Choden↗

DELTA: An Open-Source Framework to Simplify Machine Learning with Satellite Imagery

DELTA (Deep Earth Learning, Tools, and Analysis) is an open-source framework developed at NASA to simplify running and training machine learning (ML) models on satellite imagery. Users new to machine learning can run existing ML models on satellite imagery with minimal setup and configuration. For experienced ML users, DELTA helps simplify data engineering, preprocessing steps, and reduces the need for boilerplate code that needs written to make satellite imagery datasets palatable for machine learning. This lets data scientists focus on model development while DELTA handles the imagery manipulation. This presentation will demonstrate DELTA’s functionality and share some examples from an active project using it for flood mapping using imagery from multiple satellite sources.

deep learning↗

Post-Landing Major Element Quantification Using SuperCam Laser Induced Breakdown Spectroscopy

The SuperCam instrument on the PerseveranceMars 2020 rover uses a pulsed 1064 nm laser to ablate targets at a distance and conduct laser induced breakdown spectroscopy (LIBS) by analyzing the light from the resulting plasma. SuperCam LIBS spectra are preprocessed to remove ambient light, noise, and the continuum signal present in LIBS observations. Prior to quantification, spectra are masked to remove noisier spectrometer regions andspectra are normalized to minimize signal fluctuations and effectsof target distance.In some cases, the spectra are also standardized or binned prior to quantification. To determine quantitative elemental compositionsof diverse geologic materials at Jezero crater, Mars, we use a suite of 1198 laboratory spectra of 334 well-characterized reference samples. The samples were selected to span a wide range of compositions and include typical silicate rocks, pure minerals (e.g.,silicates, sulfates, carbonates, oxides),more unusual compositions (e.g.,Mn oreand sodalite), andreplicates of the sintered SuperCam calibration targets (SCCTs) onboardthe rover. For each major element (SiO2, TiO2, Al2O3, FeOT, MgO, CaO, Na2O, K2O), the database was subdivided into five“folds” with similar distributions of the element of interest. One fold was held out as an independent test set, and the remaining fourfolds were used to optimize multivariate regression models relating the spectrum to the composition. We considered a variety of models, and selected several for further investigation for each element, based primarily on the root mean squared error of prediction (RMSEP) on the test set, when analyzed at 3m. In cases with several models of comparable performance at 3 m, we incorporated the SCCT performance at different distances to choose the preferred model. Shortly after landing on Mars and collecting initial spectra of geologic targets, we selected one model per element. Subsequently, with additional data from geologic targets, some models were revised to ensure results that are more consistent with geochemical constraints. The calibration discussed here is a snapshot of an ongoing effort to deliver the most accurate chemical compositions with SuperCam LIBS.

Mars 2020↗

A Modern Load Relief Guidance Scheme for Space Launch Vehicles

Launch vehicle load relief algorithms are concerned with realizing a reduction of transient bending moments near maximum dynamic pressure. Traditional approaches to load relief typically use inner-loop acceleration feedback to reduce the wind-induced angle of attack. When implemented in the inner loop, load relief bandwidth is necessarily limited by the achievable stability margins, and when acceleration feedback is employed, by the uncertainty associated with structural modes that couple with the body-mounted accelerometer. The structure of inner loop load relief increases the dimensionality of the flight control gain and filter optimization problem. Most importantly, classical load relief laws do not take advantage of high-rate and high-accuracy GPS-aided inertial velocity data that is readily available from modern strap down IMUs. In this paper, a novel load relief guidance scheme is described that uses direct angle-of-attack feedback in a clever mechanization. The steering commands are determined by examining the wind-perturbed dynamics of a launch vehicle with respect to a gravity turn ascent trajectory. An angle of attack estimate is derived from GPS-aided inertial data and pre-launch range wind measurements, and it is shown that a reduction worst-case rigid-body loads can be realized without requiring air data. The algorithm also includes a high-rate navigation data preprocessing scheme that operates directly on the IMU delta-theta and delta-velocity measurements in order to produce a filtered acceleration estimate at the vehicle center of mass. The outer-loop guidance scheme simplifies the design process for the classical inner-loop autopilot. Algorithm performance is demonstrated using Monte Carlo analysis of a representative liquid booster in a production high fidelity launch vehicle simulation.

NESC↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

2022 Spring Internship Exit Presentation

As efforts of the National Aeronautics and Space Administration (NASA) and the Federal Aviation Administration (FAA) continue to digitize the air traffic management (ATM) domain, there is countless times of need for downstream natural language processing (NLP) tasks such as named entity recognition, text summarization, classification, and more. Although there are a plethora of open-sourced pre-trained transformer models in the NLP field such as BERT, RoBERTa, XLNet, and GPT-3, these models are trained on general corpora and perform poorly on domain-specific terminology and phraseology seen in ATM documents such as Notice to Airmen (NOTAMs) and Letters of Agreement (LoA). Our proposed research objective will be to first gather a large corpus of air traffic management related documents, orders, notices, books, technical papers, conference papers, articles, and other miscellaneous sources of text data from the FAA, NASA, and accredited conference and publication societies. After gathering this data, many steps will have to be taken to collate and preprocess the data into a format understandable by our test transformer models. Thirdly, we will set up training pipelines to train the RoBERTa model on its unsupervised training task masked language modelling (MLM) using resources provided by the NASA Advanced Supercomputing (NAS) facilities. Finally, these fine-tuned transformer models will be evaluated on their performance on down-stream NLP tasks as mentioned above, to show whether they will be effective when working with ATM related data or not. Once complete, this model could be made open-sourced on the HuggingFace website, where the rest of the ATM community can access and utilize this tool.

NLP↗

WET Water Resources: A Google Earth Engine Python API Tool to Automate Wetland Extent Mapping Using Radar Satellite Sensors for Wetland Management and Monitoring

Wetland ecosystems are annually or seasonally wet transition zones between land and water. They provide a range of ecosystem services such as water filtration, flood mitigation, and carbon sequestration, as well as hosting biodiversity hotspots. Although they fulfill fundamental physical and natural processes, wetland extent and health are threatened by anthropogenic influences related to urbanization, population increase, pollution, and climate change. Recognizing the need to quantitatively monitor changes in these recently threatened ecosystems in a timely and cost-effective way, we developed a Google Earth Engine (GEE) Python API tool for automated wetland extent mapping using optical and radar satellite sensors that can be applied globally. The tool will significantly improve wetland change analysis and monitoring as the optical and SAR data proves high resolution (5-10 m) imagery, and SAR data is unaffected by cloud cover and light availability (day vs. night), which are common limitations for other remotely sensed sensors. The tool utilizes Copernicus Sentinel-1 C-band and NISAR L-band synthetic aperture radar (SAR) imagery. During image preprocessing, we applied a MODIS snow mask product to mask global snow coverage, which would affect land classification sensitivity. Calibration and validation were conducted through a historical change and sensitivity analysis of the Sudd watershed located in central Sudan. The tool was the first of its kind, as it enables NISAR data processing through an open-source GEE repository, further expanding and improving the utility of NASA Earth observations and contributing to NASA Open Science initiatives. We anticipate the tool will be used by researchers and practitioners interested in wetland monitoring and management..

Lori Berberian↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning↗

WET Water Resources: A Google Earth Engine Python API Tool to Automate Wetland Extent Mapping Using Radar Satellite Sensors for Wetland Management and Monitoring

Wetland ecosystems are annually or seasonally wet transition zones between land and water. They provide a range of ecosystem services such as water filtration, flood mitigation, and carbon sequestration, as well as hosting biodiversity hotspots. Although they fulfill fundamental physical and natural processes, wetland extent and health are threatened by anthropogenic influences related to urbanization, population increase, pollution, and climate change. Recognizing the need to quantitatively monitor changes in these recently threatened ecosystems in a timely and cost-effective way, we developed a Google Earth Engine (GEE) Python API tool for automated wetland extent mapping using optical and radar satellite sensors that can be applied globally. The tool will significantly improve wetland change analysis and monitoring as SAR data provides high resolution (5-10 m) imagery, unaffected by cloud cover and light availability (day vs. night), common limitations for other remotely sensed sensors. The tool utilizes Copernicus Sentinel-1 C-band and NISAR L-band (once operational and available on the GEE repository) synthetic aperture radar (SAR) imagery. During image preprocessing, we applied a Terra Moderate Resolution Imaging Spectroradiometer (MODIS) snow product to determine regional snow coverage, which affects land classification sensitivity. Calibration and validation were conducted through a historical change and sensitivity analysis of the Sudd wetland located in central Sudan. The tool was the first of its kind, as it enables NISAR data processing through an open-source GEE repository, further expanding and improving the utility of NASA Earth observations and contributing to NASA Open Science initiatives. We anticipate the tool will be used by researchers and practitioners interested in wetland monitoring and management.

Inundation↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of trained machine-learning models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a Windows app that has been created to deploy trained machine-learning models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of machine-learning application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). Current version of the app focuses on the performance prediction of conventional turbofans. The app gets user input for a turbofan design, preprocesses the input data, and deploys trained machine-learning models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The machine-learning predictive models were built by employing supervised deep-learning algorithm to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these machine-learning models using the app shows that Aero-Engines AI is an easy-to-use and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage.

machine learning↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be easily expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning↗

Commercial Smallsat Data Acquisition Program On-ramp #2 Airbus U.S. Synthetic Aperture Radar (SAR) Evaluation Report

In 2017, NASA’s Earth Science Division (ESD) launched the Private-Sector Small Constellation Satellite Data Product Pilot, now referred to as the Commercial Smallsat Data Acquisition (CSDA) program. The objective of CSDA is to identify, evaluate, and acquire commercial remote sensing data that support NASA’s Earth science research and application activities. The Pilot successfully concluded in early 2020, when CSDA transitioned into a sustained program with on-ramping opportunities for new vendors as the industry emerges with new candidates and capabilities. In October 2019, a Request for Information (RFI) seeking capability statements from parties interested in providing data from spaceborne platforms was released for the CSDA on-ramp #2 evaluations. To be responsive to the RFI, the commercial satellite constellations had to consist of three or more operating spacecraft actively collecting data in a non-geostationary orbit with full latitudinal coverage and be U.S. companies. Two vendors responded to the RFI and were evaluated by a committee composed of NASA ESD leadership, program managers, and scientists. Both vendors satisfied the RFI requirements and were asked to respond to a Request for Proposal (RFP). After review of the proposals, NASA entered into a Blanket Purchase Agreement (BPA) with Airbus Defense and Space GEO, Inc. (Airbus) U.S. in September 2021 and with BlackSky Geospatial Solutions, Inc. (BlackSky) in November 2021. In this report, CSDA provides an evaluation of the usefulness of data provided by the Airbus U.S. Synthetic Aperture Radar (SAR) satellite constellation, consisting of TerraSAR-X (launched in 2007), TanDEM-X (launched in 2010), and PAZ (launched in 2018), for advancing NASA’s Earth system science research and applications. The evaluation of the BlackSky commercial data will be provided in a separate report. To conduct the Airbus evaluation, NASA’s ESD augmented 13 existing research projects that could potentially benefit from, and had the expertise to evaluate, the commercial data being considered for longer-term purchase. Investigators from NASA’s Research and Analysis Program science focus areas and from NASA’s Applied Sciences Program elements participated in the evaluation. A summary of the research areas evaluated by the Principal Investigator (PI) teams is presented in Figure 3. CSDA also funded a dedicated activity to evaluate the satellite data quality (calibration and geolocation) independently by assessing the accuracy of data from Airbus. Evaluation activities were carried out by the selected PIs from December 7, 2022, to December 7, 2023. Delivery of datasets requested by the researchers began in January 2023. The vendors were evaluated on the accessibility of data, accuracy and completeness of metadata, and promptness and quality of user support services. Datasets purchased during the evaluation have been archived by NASA and will be made available to current and future government-funded researchers in accordance with the End User License Agreement (EULA). This synthesis report distills and integrates the findings of research reports commissioned by NASA for the Airbus evaluation. This report also includes recommendations that inform the way ahead for the program. The scientific results from the evaluations demonstrated that the commercial data from Airbus were able to advance NASA research and applications. However, the PIs encountered limitations that diminished the usefulness of the data due to the amount of effort that was required to access, preprocess, and analyze these data. One significant issue encountered was the limited spatial and temporal coverage of the data in the Airbus archive that could be used to conduct time series analyses or assessments over large spatial scales. Overall, however, the utility and the quality of the evaluated data outweighed the difficulties encountered, and NASA has concluded that the Airbus SAR data would complement NASA’s existing Earth observation capabilities and Airbus U.S. would qualify to participate in the sustained phase of the program.

Batuhan Osmanoglu↗

Making NASA GES DISC Level 2 Data GIS Analysis Ready

There are many valuable data hosted by NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) for GIS applications in the areas of extreme weather events, climatic anomaly, and public health. However, using NASA Earth Science Data poses some challenges for GIS users. Many of these users are not experts in Earth Observation and have little knowledge about NASA's Earth science data. In the GIS community, GeoTiff is the most widely used raster format, whereas NASA's data is primarily in complex multidimensional netCDF and HDF formats. This complexity makes it difficult for GIS users, especially those who are unfamiliar with these formats. Although GIS software like ArcGIS has made progress in processing multidimensional netCDF data, certain issues still remain, particularly with level 2 data. In this study, we use TROPSpheric Monitoring instrument (TROPOMI) level 2 data as an example to demonstrate how to make such data GIS analysis ready. The process involves: 1. Creating a feature layer from the TROPOMI level 2 data. 2. Converting the feature layer to a gridded raster dataset. 3. Mosaicking gridded raster datasets into a raster dataset covering the entire desired extent. 4. Generating a symbology with a GIBS-specific style that aligns with the visual standards and requirements of GIBS. 5. Publishing image services. By performing these preprocessing and transformation steps, NASA level 2 data can be made compatible and ready for use within GIS software for various spatial analysis and visualization tasks.

Geographic Information System↗

Curating AI-Ready Datasets for Equity and Environmental Justice: A Data-Centric AI Case Study

An equitable and environmentally just community is essentialin order to avoid disproportionate burden borne by vulnerablecommunities. This need becomes pressing in the aftermathof an extreme event such as disaster or hazard when it is diffi-cult for the governing bodies to implement resource allocationas per the need. Artificial Intelligence (AI) algorithms canhelp surface Equity and Environmental Justice (EEJ) issueswhen trained on EEJ datasets. However, curating AI-readyEEJ training datasets is challenging due to differences in fac-tors such as heterogeneity, resolution, modality, and level ofexpertise in labeling. Additionally, EEJ issues involve sensi-tive information where uncertainties and errors could degradethe performance of AI algorithms. For eg. Error in seasonalcrop yield information can highly affect the prediction of an-nual crop yield. To address these challenges, Data-centricAI (DCAI) methods are employed, which enhance AI algo-rithm performance even with limited training samples. DCAIprioritizes data quality, thereby reducing the adverse effectsof uncertainties and errors during the model training process.This research proposes a novel dataset and benchmark for an-alyzing the effect of the Maui Wildfire of 2023 for Equityand Environmental Justice (EEJ) issues. The proposed datasetaligns with the concepts of DCAI such as annotation quality,data preprocessing, privacy, feature engineering, governanceand provenance. We firmly believe that the proposed datasetwould lay a foundation to implement robust and reliable mod-ern AI algorithms for addressing EEJ issues.

Paridhi Parajuli↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

Open Science Approach to Analyze Climate-Crop Relationships in the US Leveraging GES DISC and Galaxy Workflows

Understanding the intricate relationship between climate variability and agricultural production is crucial for ensuring food security. This study investigates the impact of climate parameters, such as temperature, precipitation, and soil moisture, on major US crop yields. Adopting an open science approach, the study analyzes the impact of climate on agricultural production in the United States. The Galaxy workflow engine serves as the primary tool for integrating climate data from the Goddard Earth Sciences Data and Information Services Center (GES DISC), retrieved via the Giovanni system, with yield statistics from the United States Department of Agriculture’s National Agricultural Statistics Service (USDA NASS). Extensions for reading, preprocessing, and analyzing external data have been developed, enabling the creation of workflows within the Galaxy platform. The development of a reproducible workflow allows for the calculation of seasonal climate averages, which are then assessed for their correlation with crop yields. This methodology ensures the replicability of the research, promoting transparency and collaboration in the scientific community. Correlational and regression analyses have been applied to different sub-zones and crops. The findings from this research offer valuable insights into the relationship between climate parameters and crop yields. These insights contribute to a deeper understanding of climate-crop relationships, providing a solid foundation for informed decision-making in the agricultural sector. The high correlation values indicate a significant relationship between climate parameters and crop yields, underscoring the importance of considering climate factors in agricultural planning and policymaking. This research also exemplifies the power of open science in advancing our understanding of complex environmental and agricultural phenomena. By leveraging open data and services, it provides a robust and replicable framework for future studies in this critical field.

Open science↗