Search NASA⌕ Search

SEARCH · Search NASA

Results for “Machine Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Observing Supraglacial Lakes Using Deep Learning and PlanetScope Imagery

Supraglacial lakes (SGL)s result from melt water accumulation in topographic depressions on the surface of glaciers. SGLs primarily affect glacial dynamics through a positive feedback loop in which the albedo-lowering effect of SGLs can escalate surface melt leading to increases in lake extent and depth, amplifying the afore mentioned albedo-lowering effect. The implications of accelerated glacial melt include increased sea level rise and modifications to ocean primary productivity. SGLs are critical indicators of surface melt and its downstream impacts and should be monitored efficiently. In situ observations and measurements of SGLs are time consuming, cost-prohibitive and difficult to scale. Earth observation data and machine learning enable scalable monitoring of SGLs through pattern detection and quantification of lake evolution over time [1]. This work presents a model developed by training a convolutional neural network with imagery and labels from NASA Operation IceBridge and predicting SGLs in high temporal and spatial resolution PlanetScope imagery.

Supraglacial lake↗

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael↗

Kaona: Deep Searching and Curating Safety Reporting Systems

Context: Several works in the literature have examined how safety narrative databases can be leveraged to share lessons learned. However, less attention has been given in augmenting existing processes of safety reporting systems. Aim: In this work, we introduce Kaona: An interface that weaves machine learning in existing aviation safety reporting systems activities. Method: We provide a use case of search, curation and newsletter writing to showcase how Kaona features build on existing processes and on its own to enhance information retrieval, curation and synthesis of narratives. Results: We created two instances of Kaona internally for evaluation, one using all public NASA's ASRS narratives and another using all public C3RS narratives. Data ranged from 1998 to 2024. Conclusion: Our tool provides a new way to explore safety narratives, serving to re-imagine how text databases can benefit of novel information retrieval mechanisms in the era of large language models.

asrs↗

Transcriptomics-based Machine Learning (ML) Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Building a Real-Time Predictive Flood Model for Improving Early Warning Systems in Ellicott City, Maryland

As flood events in the United States grow in frequency and intensity, the use of applied remote sensing analyses is increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP National Program partnered with the Howard County government in Maryland to investigate the use of machine learning for advanced flood risk detection, and to test the feasibility of integrating this approach into the county’s flood early warning system. To strengthen the efforts of the Howard County Office of Emergency Management (OEM), the project developed a prediction model capable of hindcasting the two severe flash flood events that devastated Ellicott City, and transitioned to an Long Short-Term Memory (LSTM) based sequence-to-sequence deep learning model with 8-hour forecast capability. The team combined data inputs from public sources including river and precipitation gauges, NASA and NOAA Earth observations, and numerical weather model products using scripts written in the Google Colaboratory Python scripting environment. In addition to designing the deep learning architecture, the team trained and tested the model, and evaluated its performance using the Nash-Sutcliffe Efficiency (NSE). The final product, called the Sequentially Trained Real-time EstimAted Model (STREAM), predicts stage height for the Hudson Branch gauge in Ellicott City using data products available in near real-time, including the High-Resolution Rapid Refresh (HRRR) model’s accumulated precipitation forecasts supplemented by stream gauge data from the OEM and the U.S. Geological Survey. STREAM was incorporated into an online dashboard in a user-friendly interface capable of triggering the alarms that initiate the OEM’s emergency response protocols up to 8 hours in advance of a predicted severe flood event. The project demonstrated the potential for the integration of open data and Earth observations into a flood risk forecasting tool capable of informing near real-time decision making.

Ryan Hammock↗

Towards a machine learning framework for acquiring and exploiting monitoring and diagnostic knowledge

In this paper we address the problem of detecting and diagnosing faults in physical systems, for which neither prior expertise for the task nor suitable system models are available. We propose an architecture that integrates the on-line acquisition and exploitation of monitoring and diagnostic knowledge. The focus of the paper is on the component of the architecture that discovers classes of behaviors with similar characteristics by observing a system in operation. We investigate a characterization of behaviors based on best fitting approximation models. An experimental prototype has been implemented to test it. We present preliminary results in diagnosing faults of the Reaction Control System of the Space Shuttle. The merits and limitations of the approach are identified and directions for future work are set.

Manganaris, Stefanos↗

Prediction of Weather Impacts on Airport Arrival Meter Fix Capacity

This paper introduces a data driven model for predicting airport arrival capacity with a look-ahead time 2-8 hour forecast. The model is suitable for air traffic flow management by explicitly investigating the impact of convective weather on airport arrival meter fix throughput. Estimation of the arrival airport capacity under arrival meter fix flow constraints due to severe weather is an important part of Air Traffic Management (ATM). Airport arrival capacity can be reduced if one or more airport arrival meter fixes are partially or completely blocked by convective weather. When the predicted airport arrival demands exceed the predicted available airport's arrival capacity for a sustained period, Ground Delay Program (GDP) operations will be triggered by ATM system. Serious imbalances between demand and capacity occur most frequently when the airport capacity is severely degraded due to either bad airport terminal surface weather or inclement convective weather around airport arrival fixes. A model that predicts the weather-impacted airport arrival meter fix throughput may help ATM personnel to plan GDP operations more efficiently. This paper identifies the characteristics of air traffic flow across arrival meter fixes at Newark Liberty International Airport (EWR). The proposed approach, based on machine-learning methods, is developed to predict the weather impacted EWR arrival Meter Fix (MF) throughput. Sector forecast coverage is used to envision the weather impact on airport arrival MF flow, and the validation is accomplished by using Convective Weather Avoidance Model (CWAM) 0.5 to 2-hour and Collaborative Convective Forecast Product (CCFP) 4 to 8-hour look-ahead forecast data for the period of April-September in 2014. Furthermore, the regression tree ensemble learning of random forests approach for translating a sector forecast coverage model to an EWR arrival meter fix throughput model is examined. The results suggest that ATM decision makers in charge of MF flow control and GDP planning may benefit from adopting the airport arrival meter capacity prediction models to estimate the inclement weather impacts.

Wang, Yao X.↗

Sampling Functions from Gaussian Processes and Structured Covariance Gaussian Networks

When learning aerodynamic models from data, it is critical to incorporate estimates of model uncertainty. This motivates the design of probabilistic aerodynamic databases which can be sampled to generate physically and statistically plausible aerodynamic models. In this talk we discuss how to sample deterministic functions from two different kinds of probabilistic models and demonstrate their use. First, Gaussian Process Regressors (GPRs) are a widely used probabilistic kernel-based model which can be thought of as Gaussian distributions over functions. GPRs are generally trained by maximizing the marginal likelihood of seeing the training data over the kernel parameter space. Sample functions are easily generated by drawing points from the Gaussian distribution at desired input points. However, when the points are not known ahead of time, the classical sampling approach is not possible since successive function samples will generate different function realizations. We present an approach for sampling consistent function evaluations from a GPR over multiple samples. Second, we describe a neural network architecture which learns a conditional Gaussian distribution by maximizing the marginal likelihood at each point in the input space. We then discuss and compare several options for generating sample functions which match this distribution. Finally, we demonstrate the use of these probabilistic aerodynamic models in an atmospheric reentry simulation.

Gaussian process regression↗

Building a Real-Time Flood Prediction Model for Improving Early Warning Systems in Ellicott City, Maryland

As flood events in the United States grow in frequency and intensity, the use of applied remote sensing analyses is increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP National Program partnered with the local government of Howard County, Maryland, to investigate the use of machine learning for advanced flood risk detection, and to test the feasibility of integrating this approach into the county’s flood early warning system. To strengthen the efforts of the Howard County Office of Emergency Management (OEM), the project developed a statistical model capable of hindcasting the two severe flash flood events that devastated Ellicott City and transitioned to a ‘Long Short-Term Memory’ based sequence-to-sequence deep learning model with 8-hour forecast capability. The team combined data inputs from public sources including river and precipitation gauges, NASA and NOAA Earth observations, and numerical weather model products using scripts written in the Google Colaboratory Python scripting environment. In addition to designing the deep learning architecture, the team trained and tested the model, and evaluated its performance using Nash-Sutcliffe Efficiency. The final product, the Sequentially Trained Real-time EstimAted Model (STREAM) predicts stage height for the Hudson Branch gauge in Ellicott City using data products available in near real-time, including the High-Resolution Rapid Refresh model’s accumulated precipitation forecasts supplemented by stream gauge data from the OEM and the U.S. Geological Survey. STREAM was incorporated into an online dashboard in a user-friendly interface capable of triggering the alarms that initiate emergency response protocols up to 8 hours in advance of a predicted severe flood event. The project demonstrated the potential for the integration of open data and Earth observations into a flood risk forecasting tool capable of informing near real-time decision making.

NASA DEVELOP↗

Heuristic Area Cost Estimation for Observational Coverage Schedulers

This paper presents a comparison of heuris- tics used to estimate the amount of time it would take for a spacecraft to image an area using Boustrophedon decomposition (Choset and Pignon 1998). Machine learning tech- niques are used to characterize algorithmic performance of coverage algorithms. It is shown that an ordinary least-squares linear model is among the most accurate in a set of constant and linear order regression models both in terms of memory consumption and schedule duration. These are demonstrated using the ASPEN planning system (Fukunaga et al. 1997) on the Eagle Eye domain.

Knight, Russell↗

Exploring Applications of Machine Learning for Wildfire Monitoring and Detection using Unmanned Aerial Vehicles

Wildfires are increasing in frequency and severity around the world, including the United States. The losses caused by wildfires could be mitigated if high-risk areas, hotspots, and flare-ups could be monitored continuously, such as through the use of Unmanned Aerial Vehicles (UAVs). This paper documents exploratory efforts using machine learning to determine efficient flight paths for UAVs and to detect wildfires using image classification. On path planning, three machine learning techniques—Genetic Algorithm, Simulated Annealing, and Dynamic Programming—were explored. Genetic Algorithm was found to be an effective approach for path planning for wildfire monitoring and surveillance by UAVs. For a scenario of 25 locations in a circular arrangement, the algorithm was able to return the optimal path. The accuracy and execution time was found to be sensitive to the algorithm hyperparameters selected, which was especially evident in scenarios with hundreds or thousands of locations. Simulated Annealing was also found to be an effective approach for UAV path planning, with a major benefit of avoiding getting trapped in local minima and being straightforward to implement. Like Genetic Algorithm, the performance of Simulated Annealing was also found to be sensitive to the algorithm hyperparameters selected. By comparison, Dynamic Programming guarantees optimality for any number of locations, but it was found to be less practical in terms of execution time for scenarios with more than about a couple dozen locations. On wildfire detection, image classification using deep learning with a convolutional neural network was explored. Transfer learning was found to be a useful technique to efficiently train deep learning models. Also, it was determined that GPU processing can increase training speed by an order of magnitude, which enables significantly faster development. For a validation test set of 500 images, there were only two false negatives and zero false positives. These results demonstrate that detecting wildfires in static cameras using machine learning is feasible and establish a baseline for using images captured by UAVs in flight for wildfire detection.

Wildfire management↗

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term. The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission (TRMM) Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters↗

Ellicott City Disasters II: Enhancing a Statistical Flood Risk Model to Continue Improving Early Warning Systems and Public Safety in Ellicott City, Maryland

As flooding events in the United States grow in frequency and intensity, the use of technological advancements and applied science are increasingly necessary for effective flood monitoring and warning systems. The NASA DEVELOP Ellicott City Disasters II project investigated the use of machine learning for applications in flood risk detection to support the improvement of early warning systems. To strengthen the efforts of the Howard County Office of Emergency Management (OEM) in building a more robust flood monitoring system, the project improved the original statistical flood risk model, FLuME (Flood Learning Model Environment), programmed by the first DEVELOP term The enhancements incorporated an additional six years of precipitation and soil moisture data from the North American Land Data Assimilation System (NLDAS), modeled using Aqua Advanced Microwave Scanning Radiometer for EOS and Tropical Rainfall Measuring Mission TRMM Microwave Imager. These Earth observations were supplemented by stream gauge data from the OEM and the US Geological Survey. The resultant flood risk model FLASH (Flood Learning Environment and Severity Assessment Hub) was trained to evaluate input variables and predict stage height in Ellicott City in real time. The addition of an advanced deep learning framework known as long short-term memory improved the model’s ability to capture relationships between variables. To assess the effectiveness of the new model, FLASH produced a model efficiency metric of 0.99, a significant improvement over the 0.85 value produced by the previous model. The project assisted the OEM in pursuing the integration of open data and NASA Earth observations into a threat matrix capable of informing near real-time decision making.

Disasters↗

A Recursive Multi-step Machine Learning Approach for Airport Configuration Prediction

Airport configuration selection is a complex decision-making process that involves several operational and human factors. In this paper we propose a novel recursive multi-step machine learning (ML) approach to predict airport configuration. The multi-step approach guarantees stability of the predicted configuration by taking as input the configuration predicted at the previous time step. The features of the proposed model include weather data, future arrival and departure counts and current configuration. Due to the importance of arrival and departure counts in predicting the airport configuration, arrival counts are calculated using landing time predictions selected from physics-based landing time predictions available in FAA System Wide Information Management data feeds for each flight. The selection rules were developed and refined to select the most accurate time for different phases of flight. The proposed model predicts the airport configurations up to 6 hours ahead. In this paper we show the predictive performance of the proposed model for six major US airports, including Charlotte Douglas International Airport (CLT), Dallas/Fort Worth International Airport (DFW), John F. Kennedy International Airport (JFK), Newark Liberty International Airport (EWR), LaGuardia Airport (LGA) and Dallas Love Field Airport (DAL). We trained and evaluated models on 2019 and 2020 data in order to study the effect of the pandemic and how changes in traffic patterns affected the performance of the proposed model. Results are compared with a baseline assuming no airport configuration changes. In our results for DFW, we obtained a prediction accuracy of 89.3% for 3 hours ahead prediction, and 82.8% for 6 hours ahead when applied on 2019 data.

machine learning↗

A Recursive Multi-step Machine Learning Approach for Airport Configuration Prediction

Airport configuration selection is a complex decision-making process that involves several operational and human factors. In this paper we propose a novel recursive multi-step machine learning (ML) approach to predict airport configuration. The multi-step approach guarantees stability of the predicted configuration by taking as input the configuration predicted at the previous time step. The features of the proposed model include weather data, future arrival and departure counts and current configuration. Due to the importance of arrival and departure counts in predicting the airport configuration, arrival counts are calculated using landing time predictions selected from physics-based landing time predictions available in FAA System Wide Information Management data feeds for each flight. The selection rules were developed and refined to select the most accurate time for different phases of flight. The proposed model predicts the airport configurations up to 6 hours ahead. In this paper we show the predictive performance of the proposed model for six major US airports, including Charlotte Douglas International Airport (CLT), Dallas/Fort Worth International Airport (DFW), John F. Kennedy International Airport (JFK), Newark Liberty International Airport (EWR), LaGuardia Airport (LGA) and Dallas Love Field Airport (DAL). We trained and evaluated models on 2019 and 2020 data in order to study the effect of the pandemic and how changes in traffic patterns affected the performance of the proposed model. Results are compared with a baseline assuming no airport configuration changes. In our results for DFW, we obtained a prediction accuracy of 89.3% for 3 hours ahead prediction, and 82.8% for 6 hours ahead when applied on 2019 data.

machine learning↗

In-Situ and Remote-Sensing Data Fusion Using Machine Learning Techniques to Infer Urban and Fire Related Pollution Plumes

Airmass type characterization is key in understanding the relative contribution of various emission sources to atmospheric composition and air quality and can be useful in bottom-up model validation and emission inventories. However, classification of pollution plumes from space is often not trivial. Sub-orbital campaigns, such as SEAC4RS (Studies of Emissions, Atmospheric Composition, Clouds and Climate Coupling by Regional Surveys) give us a unique opportunity to study atmospheric composition in detail, by using a vast suite of in-situ instruments for the detection of trace gases and aerosols. These measurements allow identification of spatial and temporal atmospheric composition changes due to various pollution plumes resulting from urban, biogenic and smoke emissions. Nevertheless, to transfer the knowledge gathered from such campaigns into a global spatial and temporal context, there is a need to develop workflow that can be applicable to measurements from space. In this work we rely on sub-orbital in-situ and total column remote sensing measurements of various pollution plumes taken aboard the NASA DC-8 during 2013 SEAC4RS campaign, linking them through a neural-network (NN) algorithm to allow inference of pollution plume types by input of columnar aerosol and trace-gas measurements. In particular, we use the 4STAR (Spectrometer for Sky-Scanning, Sun-Tracking Atmospheric Research) airborne measurements of wavelength dependent aerosol optical depth (AOD), particle size proxies, O3, NO2 and water vapor to classify different pollution plumes. Our method relies on assigning a-priori ground-truth labeling to the various plumes, which include urban pollution, different fire types (i.e. forest and agriculture) and fire stage (i.e. fresh and aged) using cluster analysis of aerosol and trace-gases in-situ and auxiliary (e.g. trajectory) data and the training of a NN scheme to fit the best prediction parameters using 4STAR measurements as input. We explore our misclassification rates as related to our ground-truth labels, and with multi-layered pollution plume cases. The next step in our analysis is to optimize parameter selection for a scheme that can be applied to space-borne aerosol and trace-gas observation platforms such as OMI, and future geostationary satellites such as TEMPO and GEO-CAPE.

Neural-network↗