Search NASA⌕ Search

SEARCH · Search NASA

Results for “Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Predicting Antarctic Net Snow Accumulation at the Kilometer Scale and Its Impact on Observed Height Changes

Sub-grid-scale processes occurring at or near the surface of an ice sheet have a potentially large impact on local and integrated net accumulation of snow via redistribution and sublimation. Given observational complexity, they are either ignored or parameterized over large-length scales. Here, we train random forest (RF) models to predict variability in net accumulation over the Antarctic Ice Sheet using atmospheric variables and topographic characteristics as predictors at 1 km resolution. Observations of net snow accumulation from both in situ and airborne radar data provide the input observable targets needed to train the RF models. We find that local net accumulation deviates by as much as 172% of the atmospheric model mean. The correlation in space between the predicted net accumulation variability and satellite-derived surface-height change indicates that surface processes operate differently through time, driven largely by the seasonal anomalies in snow accumulation.

Antarctic↗

Utilizing Earth Observations to Model Probable Coastal Wetland Extent, Sea-Level Rise Inundation Risk, and Assess Impacts on Historic Hawaiian Lands

Climate induced sea-level rise poses a risk to coastal areas on the Island of Hawai’i, and many of the island’s historic cultural lands are in danger of becoming overtaken by wetlands or inundation. In partnership with the County of Hawai’i, State of Hawai’i Department of Land and Natural Resources, and Arizona State University, NASA DEVELOP mapped wetland extent and short-term sea-level rise inundation risk. We utilized Earth observations over a 10-year span (2013 – 2022) that included the NASA MEaSUREs Gridded Sea Surface Height Anomalies and MEaSUREs Group for High Resolution Sea Surface Temperature datasets, United States Geological Survey (USGS) Hawaii Digital Elevation Models (DEM), and in situ tidal gauge data. Flood risk index values were acquired for 5 known Hawai’i flood events between 2019 – 2021 from the Global Flood Mapper tool on Google Earth Engine. We used a random forest model to predict short-term sea-level rise inundation risk along the entire coast of Hawai’i. Current wetland extents and probabilistic locations of new wetlands were modeled with the most recently available data from PlanetScope Surface Reflectance optical imagery (2022), USGS 3D Elevation Program (3DEP) 10m DEM (2020), temperature and precipitation data from the Hawai’i Climate Atlas, and soils data from the Hawai’i Soil Atlas (2014) using the Wetland Intrinsic Potential tool. Results indicated locations that had the highest probability of wetland creation. The end products aimed to help the partners prioritize efforts to meeting regulation requirements for wetlands protection, evaluate the inundation risk to historical features, and support decision-making for their Shoreline Setback and Climate Adaption plans.

Lisa Tanh↗

Designing and Evaluating NASA SPoRT Center’s DustTracker-AI Model for Detecting Dust in NASA/NOAA Geostationary Satellite Imagery

- Near real-time identification of airborne dust in satellite imagery is important for mitigating the adverse effects of dust storms on human activities. - False color Red-Green-Blue (RGB) imagery has been used for dust detection, but it has limitations and can be difficult to interpret. - The NASA Short-term Research and Transition (SPoRT) center has developed a night-time dust detection random forest (NT-DustTracker-AI, Berndt et al. 2021) model using NASA/NOAA Geostationary Operational Environmental Satellite-16 (GOES-16) Advanced Baseline Imager (ABI) infrared imagery as inputs. - The SPoRT center has partnered with the NOAA National Weather Service to evaluate the model for use in weather forecasting operations, and preliminary results have been positive. - The SPoRT center has expanded the model to cover both day and night, continuing to use infrared imagery as inputs. This new model, known as DustTracker-AI, has shown good agreement with available dust observations

Robert A. Junod↗

Bryce Canyon Water Resources: Monitoring Vegetation Health and Water Availability in Bryce Canyon National Park for Drought Stress Mitigation Planning

Bryce Canyon National Park is home to groundwater-dependent ecosystems (GDEs) that are threatened by a multidecadal drought and increased groundwater extraction due to a spike in tourism. These ecosystems contain unique species that are only found in areas where near-surface groundwater is present, such as aspen groves and fens. These species contribute to the high biodiversity found in Bryce Canyon, which boosts an ecosystem’s productivity and the services it provides to the park. Unfortunately, many of these GDEs are too small to identify with traditional Earth observation platforms and are difficult to physically reach for monitoring purposes. This project partnered with the National Park Service to identify springs and seeps as a proxy for GDEs within Bryce Canyon from 2013–2022. Furthermore, this project tested the feasibility of various methods to detect and monitor springs and seeps and therefore facilitate the partner’s efforts to conserve these ecologically valuable GDEs in Bryce Canyon. The team mapped groundwater discharge with high resolution National Agriculture Imagery Program (NAIP) and assessed park vegetation trends with Landsat 8 Operational Land Imager (OLI) and PlanetScope imagery. In-situ precipitation data and the Western Land Data Assimilation System (WLDAS) were used to produce time series of climatic variables. Seeps and spring locations were predicted using random forest classification and maximum entropy machine learning models.

Groundwater dependent ecosystems↗

Evaluating Meteorological Dust Events and Machine-Learning Based Dust Identification in Geostationary Satellite Imagery

NASA scientists in the Short-term Prediction Research and Transition Center (SPoRT) developed a physically-based machine learning approach to identify dust in satellite imagery with a focus on night-time dust detection (Berndt et al. 201; DustTracker-AI). NASA/NOAA Geostationary Environmental Operational Satellite-16 (GOES-16) imagery was used for training and model inputs. The training, testing and validation data set consists of 28 events in the Southwest United States, capturing dust and null events in the region from 2018-2020.With 83 distinct images and millions of pixels a random forest model was trained and validated, correctly labeling 85% of dust pixels.For the first time, the model was run in near-real time production during the spring of 2022 and dust probability visualizations were made available to NOAA National Weather Service (NWS) forecasters to assess its utility for dust forecasting. Results indicated the model helped increase the confidence in the presence of dust and enabled dust tracking for a longer period of time into the night-time hours. Forecaster assessment and running the model in near real-time allowed for the team to determine the types of events missed, captured, and false alarms. To gain additional context on model performance,the SPoRT team sought to gather more detailed information on the training database(e.g., meteorological characteristics and drivers). The goal of this project was to identify the meteorological drivers for the dust events and create a database which synthesized information from observations, forecaster discussions, and analyses pertaining to the dust events to understand the types of events currently used to train the model. A more detailed meteorological synopsis was created for each dust event in the training, testing, and validation datasets. Following the completion of the database and documentation, the classification details revealed that 88% of the dust events were synoptically driven while mesoscale events were less prevalent in model datasets. Meteorological conditions found such as mixing layer depth and wind velocity had mean values of 645mb and 21kt respectively.With conditions of deep mixed layers and moderate to strong surface winds a mesoscale thunderstorm outflow event was considered and subsequently added to the model training data set to test the impact of additional mesoscale training data. The model was retrained and then qualitatively tested on a sample thunderstorm outflow case that the original model was unable to identify. Preliminary results showed potential that the addition of more mesoscale events included in the training data could help to better identify indistinct and localized dust events.

Connor Welch↗

A Comprehensive Machine Learning Study to Classify Precipitation Type over Land from Global Precipitation Measurement Microwave Imager (GPM-GMI) Measurements

Precipitation type is a key parameter used for better retrieval of precipitation characteristics as well as to understand the cloud–convection–precipitation coupling processes. Ice crystals and water droplets inherently exhibit different characteristics in different precipitation regimes (e.g., convection, stratiform), which reflect on satellite remote sensing measurements that help us distinguish them. The Global Precipitation Measurement (GPM) Core Observatory’s microwave imager (GMI) and dual-frequency precipitation radar (DPR) together provide ample information on global precipitation characteristics. As an active sensor, the DPR provides an accurate precipitation type assignment, while passive sensors such as the GMI are traditionally only used for empirical understanding of precipitation regimes. Using collocated precipitation type flags from the DPR as the “truth”, this paper employs machine learning (ML) models to train and test the predictability and accuracy of using passive GMI-only observations together with ancillary information from a reanalysis and GMI surface emissivity retrieval products. Out of six ML models, four simple ones (support vector machine, neural network, random forest, and gradient boosting) and the 1-D convolutional neural network (CNN) model are identified to produce 90–94% prediction accuracy globally for five types of precipitation (convective, stratiform, mixture, no precipitation, and other precipitation), which is much more robust than previous similar effort. One novelty of this work is to introduce data augmentation (subsampling and bootstrapping) to handle extremely unbalanced samples in each category. A careful evaluation of the impact matrices demonstrates that the polarization difference (PD), brightness temperature (Tc) and surface emissivity at high-frequency channels dominate the decision process, which is consistent with the physical understanding of polarized microwave radiative transfer over different surface types, as well as in snow and liquid clouds with different microphysical properties. Furthermore, the view-angle dependency artifact that the DPR’s precipitation flag bears with does not propagate into the conical-viewing GMI retrievals. This work provides a new and promising way for future physics-based ML retrieval algorithm development.

machine learning/artificial intelligence↗

Fine particulate concentrations over East Asia derived from aerosols measured by the Advanced Himawari Imager using machine learning

Fine particulate matter with a diameter below 2.5 μm (PM 2.5 ) is deleterious to the cardiovascular and respiratory systems. It is often difficult to assess the effects of PM 2.5 on human health over regions with limited ground monitoring sites, especially in East Asia. As an alternative, we estimated near-surface PM 2.5 concentrations by analyzing Advanced Himawari Imager (AHI) Yonsei Aerosol Retrieval (YAER) products. This study incorporates daytime data for East Asia covering the Korean Peninsula, China, Japan, Southeast Asia, and southern Mongolia. We collocated AHI YAER product pixels with meteorological, land-cover, and other ancillary data for the period from March 2018 to February 2019. To estimate PM 2.5 concentrations over wide areas spanning many countries displaying various relationships between aerosol optical depth and PM 2.5 , monthly models were developed by considering both the spatial and temporal characteristics of ground-based PM 2.5 measurements. Random forest machine learning model estimated ground-level mass concentrations of PM 2.5 ; subsequent 10-fold cross validation (CV) yielded a CV R 2 value of 0.81 and a CV root mean squared error (RMSE) of 12.3 μg m -3 . We investigated the spatial pattern of PM 2.5 concentrations over multiple countries and seasonal variation in PM 2.5 concentrations. Diurnal variation of a severe PM 2.5 event in the Korean Peninsula was investigated as a case study. The model captured the extremely heterogeneous spatial distribution of PM 2.5 concentrations peaked around local noon. To measure the capability of the developed model to estimate PM 2.5 concentrations in areas with few in-situ data, its predictive performance was evaluated using a dataset independent of the training process with an R 2 of 0.60 and RMSE of 8.18 μg m −3 . This study demonstrates the potential for satellite-based PM 2.5 estimation for areas with insufficient measuring stations.

Pm2.5↗

Assessing Change in Aspen Extent in Northern Yellowstone National Park

The trophic cascade among wolves, elk, and aspen has influenced the landscape of Yellowstone National Park. Aspen promote greater biodiversity and have been indirectly affected by the 1926 removal and 1995 reintroduction of wolves in the park. Partnered with Yellowstone National Park, Utah State University, and the University of Wisconsin–Stevens Point, NASA DEVELOP analyzed the change in aspen stand extent from 1954 to 2021 over the elk wintering range that occurs in the northern part of Yellowstone and southern Montana. Focusing on 113 stands corresponding with belt transects monitored annually since 1999, the team georeferenced and digitized 1954 historical aerial imagery to determine aspen stand extent. From 1986 to 2019, the team processed Landsat 5 Thematic Mapper (TM) and Sentinel-2 Multispectral Instrument (MSI) imagery for the entire elk wintering range using a random forest model to classify landcover types. Outputs were refined using a phenological approach that distinguishes between deciduous and evergreen landcover by differencing summer and fall vegetation indices. Finally, the team conducted a similar analysis for 2021, utilizing both Landsat 8 Operational Land Imager (OLI) and Sentinel-2 MSI imagery. The team again focused on the stands associated with belt transects for this analysis to compare the beginning and end of the study period. Preliminary results indicate a slight decline over time. These results expand the understanding of the role of wolves on aspen in the Yellowstone ecosystem and inform future rewilding decisions within the park and beyond.

Vanessa Bailey↗

Using Machine Learning to Infer Material Properties of Debris Fragments from X-ray Images in the DebriSat Project

The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.

Saik Anam Siam↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential in mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail walking and US-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986 to 2022. Using data from Landsat 5 and 8, Sentinel-2, the National Agriculture Imagery Program, and PlanetScope, the team computed vegetation indices including the Normalized Difference Vegetation Index, Normalized Difference Moisture Index, Modified Soil Adjusted Vegetation Index 2, Enhanced Vegetation Index, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on the vegetation indices and spectral bands before running k-means++ clustering and random forest classification algorithms. Between all datasets, we found the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use the end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants. The NPS’s collection of ground data for 2022–2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive spread over time.

Carson Schuetze↗

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. It affects the efficiency of the National Airspace System (NAS) and both surface and airspace operations can benefit from better understanding future runway configurations. In this paper, we present a comprehensive implementation of predictive models for runway configuration estimation from large volumes of historical data. Specifically, operational data from two full years (2018 and 2019) is collected, analyzed, and fused together to build the data product used in this work. The data set differs from prior work in the field in terms of its scope, resolution, and variety of factors collected and considered. Meteorological data is collected from two different sources – current weather conditions from METAR (Meteorological Terminal Aviation Routine Weather Report) and forecast weather conditions from Localized Aviation MOS Program (LAMP). Operational data from the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) related to scheduled and actual number of arrivals and departures, average taxi times, etc. are collected. NASA’s Sherlock Data Warehouse is used to identify critical information such as go-arounds, and other events that might impact RCM decision-making. All data is collected and aggregated over 15-minute intervals throughout the two years. This provides a resolution like the timescales that might be necessary for runway configuration management decision-making. A variety of supervised learning algorithms are tested including Support Vector Machine, Random Forest, Gradient Boosting, etc. including tuning of the model hyperparameters. The modeling process is applied and presented on two representative U.S. airports – Charlotte Douglas International Airport (KCLT) and Denver International Airport (KDEN). The two airports present different levels of complexity in terms of the total number of configurations used and provide a balanced perspective on the generalizability of the developed approach to other airports in the NAS. Initial results are promising (F1 score of 0.91 at KCLT and 0.83 at KDEN) for data in the test set. The final paper will contain a comprehensive comparison between different models and model building strategies as well as further refined results. Most important predictors for each airport will be identified along with a discussion and recommendations on adapting the framework to other scenarios.

Tejas G Puranik↗

What’s That Supposed to Mean? Capturing Micro-Behaviors in Teams

Future long-duration space exploration (LDSE) crews will require extensive coordination, cooperation, and team functioning as they face a myriad of challenges rooted in both taskwork and teamwork (Bell et al., 2015; Landon et al., 2018). While exposed to extreme conditions, crew members must navigate living and working together in prolonged confinement. Moreover, astronaut teams are becoming increasingly diverse, introducing significant variability in team composition. This increasing diversity, alongside traditional constraints of LDSE, introduces additional challenges into effective team functioning. To date, most methods for capturing team functioning rely on self-report measures. Such measures are prone to several limitations, including but not limited to social desirability bias, halo effect, and leniency effects (Trull & Ebner-Priemer, 2013), which skew data and limit nuanced understandings of phenomena at play. Self-report measures broadly capture team functioning, lending the nature of such methods to identifying underlying “macro”-behaviors (i.e., behaviors that are long-standing and last over time). However, team functioning is far more complex than a series of macro-behaviors, rendering reliance on self-report data deficient for accurate measurement. Recent research demonstrates the potential of alternative methods for capturing team functioning, such as speech and physiological data (Chaffin et al., 2017; Murray & Oertel, 2018). Consequently, these methods are more suitable for capturing micro-behaviors: brief, often unconscious expressions that affect the extent to which an individual feels included by others around them (Paletz et al., 2013). Micro-behaviors can be further classified into micro-aggressions (i.e., subtle, negative exchanges; Keller & Galgay, 2010) or micro-affirmations (i.e., subtle, positive exchanges; Kyte et al. 2020), both of which influence team functioning. Due to the subtle nature of micro-behaviors, contextual factors have a significant impact when determining if it is aggressive or affirmative. Additionally, several iterations of micro-behaviors can have lingering effects on team interactions. For example, the use of “mm-hmm” by a crew member can function as both a micro-affirmation and micro-aggression. Specifically, it can be indication of active listening (i.e., micro-affirmation) or as an expression of annoyance (i.e., aggression) depending on the context in which it occurs. Auditory features (e.g., tone, frequency) can help delineate between the two forms; however, the contextual factors (e.g., previous interactions between team members, crew demographics) add a layer of complexity that render auditory features alone as insufficient to capture micro-behaviors. Consequently, this paper seeks to provide a novel approach in which multi-modal data (i.e., auditory features and contextual features) are used in a random-forest model to better identify distinguishing characteristics between micro-affirmations and micro-aggressions. In turn, detected micro-behaviors are used to predict team performance, thereby demonstrating the value of capturing micro-behaviors as supplemental data to macro-behaviors.

Sydney Begerowski↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Bandelier Ecological Conservation: Mapping Invasive Species Along the Rio Grande Corridor in Bandelier National Monument

The Southwest U.S. has experienced a growth of invasive riparian species, specifically Elaeagnus angustifolia (Russian olive), Tamarix ramosissima (saltcedar), and Ulmus pumila (Siberian elm), which alter local soil chemistry and outcompete native species. Locating these exotic species is critical for ecological conservation; however, field identification can be resource intensive. NASA DEVELOP partnered with the National Park Service (NPS) at Bandelier National Monument (BAND) to assess the feasibility of using Earth observation data to map invasive species along the Rio Grande corridor of the park. The team used Landsat 8 OLI, Sentinel-2 MSI, and ISS DESIS imagery to compute principal components based on spectral bands, vegetation indices, and terrain indices. Using the first five principal components, the team created classification maps using both a k-means classification algorithm and a random forest algorithm to differentiate between native and non-native species. The team derived maps for the three invasive riparian species in the region for the last five years. The team found that invasive species covered 33% of the park's river corridor in 2023, and the invasive species extent has increased by 5.7% from 2019 to 2023. The methods will serve as a guide for aiding historic and present invasive species identification in riparian regions, and the NPS staff at BAND will use the results to inform local mitigation practices and advocate for invasive species removal.

Evan Barrett↗

Using Machine Learning to Infer Material Properties of Debris Fragments from X-ray Images in the DebriSat Project

The DebriSat project is a collaboration effort with the NASA Orbital Debris Program Office, the U.S. Space Force Space Systems Command Center, The Aerospace Corporation, and the University of Florida. To date, over 200,000 fragments from this ground-based, hypervelocity impact experiment have been collected, and processing is underway to determine their physical characteristics, such as material, shape, color, characteristic length, and average cross-sectional area. The x-ray process is primarily used to identify the location of the fragments and estimated size for extraction, so that these physical characteristics can be assessed. This paper proposes a machine learning-based approach to characterize materials from x-ray images of debris fragments embedded in soft-catch foam used in the DebriSat project. The novel methodology discussed in this paper will highlight the use of x-ray imagery data to characterize these fragments without extraction or a human-in-the-loop. Both supervised and unsupervised machine learning techniques are utilized with this approach to infer the physical parameters of the fragments embedded in the soft-catch foam panels used in the impact experiment based on x-ray images of the foam panels. Additionally, 3D reconstructions of the extracted fragments are created with images taken from two different angles using the structure from motion (SfM) method. The characteristic lengths and shape from the 3D reconstruction, alongside the physical characteristics of the debris, are used in the inference of the material type. To develop and test the approach, a dataset of x-ray images of debris fragments of varying sizes and materials is collected. Supervised learning methods such as convolutional neural networks (CNNs), support vector machines (SVM), decision trees, and random forest classifiers are used due to the high-dimensional feature spaces of the debris and nonlinear decision boundaries for material categorization. Given the limited pre-labeled data of embedded debris materials smaller than 10 mm, unsupervised machine learning techniques such as clustering algorithms and autoencoders are used, in addition to supervised learning methods. The clustering algorithms group similar fragments together based on their physical properties, and autoencoders reduce the dimensionality of the x ray images and extract relevant features. The performance of the proposed approach's is analyzed using a range of statistical methods, including confusion matrices, receiver operating characteristic curves, and precision-recall curves. The results are compared with those obtained using a baseline approach that relies on manual identification and classification of debris fragments. To evaluate the effectiveness of different machine learning methods, statistical tests such as t-tests, ANOVA, and cross-validation are performed, comparing the performance of CNNs, SVMs, clustering algorithms, and autoencoders. Additional analysis needs to be conducted to identify any sources of bias or variability that may affect the results, such as variations in imaging conditions or fragmentation patterns. Other topics explored are limitations, refinements, and the potential use of semi-supervised learning techniques, such as self-training to label unlabeled datasets and co-training using x-ray images taken from two different angles as two different models.

Saik Anam Siam↗

Mapping Surface Vapor Pressure Deficits From Geostationary Satellites for Fire Weather Monitoring

The increase in the wildfires were observed globally in accordance with global warming, and to real- time monitoring of wildfire risk in broad scale is demanded for wildfire management to prevent the spread of wildfires. Scientists invented a lot of indices to assess the wildfire risk. Vapor Pressure Deficit (VPD) is one of the most important meteorological components for those indices. Compared to other components of fire weather indices, VPD can change quickly from lower risk to higher risk even in sub-hourly. Therefore, real-time fire risk monitoring requires high-resolution and high- temporal VPD spatial map. Here, we developed VPD estimation method using the GOES Advanced Baseline Imager (ABI) data. Unlike the polar-orbital satellite data, the ABI can observe target region every 10 minutes, so that we can estimate VPD for fire weather in real-time. The method used to estimate VPD is same with the algorithm of NASA Earth Exchange Gridded Daily Meteorology (NEX- GDM), which estimate meteorological variables from ground weather observation and spatial variables based on random forest (RF). We calculated RF importance to select bands of ABI as input of the model. To validate our results, we compared the spatial pattern of our VPD data with the Real- Time Mesoscale Analysis (RTMA) data over the conterminous USA. We sought possibility of applying our method to the region where no real-time high-resolution weather data is available, such as South America. The developed method can produce real-time high-resolution high-frequent VPD data in the continental scale. The derived data from GOES ABI could contribute to improve the fire weather monitoring and lead to prevent wildfires.

Hirofumi Hashimoto↗

Capitol Reef Ecological Conservation: Mapping Vegetation Functional Groups to Inform Invasive Vegetation Management, Ecological Conservation and Restoration in Capitol Reef National Park

Invasive exotic plant (IEP) species have been found within the park boundaries of Capitol Reef National Park (CARE) in Utah. Currently, remotely sensed datasets such as the Rangeland Analysis Platform (RAP) from the United States Department of Agriculture (USDA) have been used to investigate IEP species within the park, but validation of the national RAP program is necessary for informing decisions at a local scale. CARE seeks a remote monitoring solution that can precisely target managerial efforts within the park’s challenging terrain and hard-to-reach locations. To fulfill this objective, we harnessed Landsat 8 Operational Land Imager (OLI) imagery and leveraged Random Forest (RF) modeling to generate classification maps characterizing vegetation functional groups for 2013 and 2022 within the park. Subsequently, the Land Change Modeler (LCM) in Idrisi TerrSet facilitated the production of a predicted classification map for 2033. The team also devised an annual grass probability map to accentuate areas impacted by exotic grasses. A comparative assessment between the RF classification map and the RAP map for 2022 revealed an overall agreement of 47.41%, with disparities primarily arising from differences in bare soil and shrub areas. Significantly, the 2022 RF-generated classification map showcased an impressive overall accuracy of 92.17%. In short, the probability map, the land cover change detection spanning 2013 to 2022, and the forecasting of observed trends into the future aids in the evaluation of invasive plant impacts and facilitation of CARE’s preparedness for potential ecological disturbances. Notably, in comparison to the RAP, the RF classification method generates functional group maps that are more representative of the study area.

Vanchy Li↗

Document Classification Techniques for Aviation Letters of Agreement

Often when working with technical documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques, including Transfer Learning, for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

Aayushi Batra↗