Search NASA⌕ Search

SEARCH · Search NASA

Results for “random forest classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Mapping Inundation from Hurricane Florence (2018) with L-Band Synthetic Aperture Radar, Commercial Imagery, and Ancillary Data via Random Forest Classification

Mapping the extent of floodwaters following extreme rainfall aids in the distribution of resources, recovery efforts, and damage assessment practices. Development of a land cover classification system focused on mapping inundation after major hurricane events using synthetic aperture radar (SAR) data could allow for the production of near-real-time inundation mapping, enabling government and emergency response entities to get a preliminary idea of a developing situation. In response to Hurricane Florence of 2018, NASA JPL collected numerous swaths of quad-pol L-band SAR data with the Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) instrument observing the record-setting river stages across North and South Carolina. The resulting fully-polarized SAR images allow for mapping of inundation extent at a high spatial resolution with a unique advantage over optical imaging stemming from the sensor’s ability to penetrate cloud cover and dense vegetation. This study seeks to determine how accurately maps of inundation can be generated from L-band SAR imagery through Random Forest classification. Once the extent of water and inundated vegetation is classified, cleanup operations are performed using fuzzy logic to reduce false detections. Estimates of water extent are then combined with datasets describing the distribution of population, buildings, and roads throughout the domain to evaluate societal impacts. Results from the Hurricane Florence case study will be discussed along with the limitations of available validation data for assessment of the classifier’s accuracy.

Alexander Melancon↗

Reference Shapefiles and Pre-trained Random Forest Classification Models for Detecting Aufeis on the North Slope of Alaska in Landsat Imagery

This dataset provides shapefiles and trained machine learning models used for aufeis detection at four sites on the North Slope of Alaska. It includes reference data for evaluating Landsat-based detection methods, supporting research on remote sensing approaches for identifying aufeis. The ReferenceData folder contains ArcGIS shapefiles of semi-automated land cover classifications for 217 Landsat Collection 2 images, categorizing pixels into six classes: aufeis, snow, ground, none, water, and cloud. The SiteBuffers.zip file includes 10-kilometer buffer shapefiles defining regions of interest around four aufeis fields (Canning21, FH1, Firth, and Kuparuk), used to test three detection techniques. Additionally, the TrainedRFModels folder contains six pre-trained Scikit-Learn Random Forest classifiers (100 trees, max depth = 30) designed to predict aufeis presence in Landsat Collection 2 Surface Reflectance images using Red, Blue, SWIR2, NDVI, and NDWI bands. This dataset supports the development and validation of remote sensing methods for mapping aufeis in Arctic environments.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Seed classification with random forest models

Premise: To improve forest conservation monitoring, we developed a protocol to automatically count and identify the seeds of plant species with minimal resource requirements, making the process more efficient and less dependent on human operators. Methods and Results: Seeds from six North American conifer tree species were separated from leaf litter and imaged on a flatbed scanner. In the most successful species-classification approach, an ImageJ macro automatically extracted measurements for random forest classification in the software R. The method allows for good classification accuracy, and the same process can be used to train the model on other species. Conclusions: This protocol is an adaptable tool for efficient and consistent identification of seed species or potentially other objects. Automated seed classification is efficient and inexpensive, making it a practical solution that enhances the feasibility of large-scale monitoring projects in conservation biology.

59 BASIC BIOLOGICAL SCIENCES↗

Invasion in the Niger Delta: Remote Sensing of Mangrove Conversion to Invasive Nypa fruticans from 2015-2020

Invasive species are a leading threat to biodiversity worldwide. Nypa palm ( Nypa fruticans ) has emerged as the predominant invasive species in the Niger Delta region of Nigeria. While endemic mangroves have high rates of carbon sequestration, stabilize coastlines, and protect biodiversity, Nypa does not provide these services outside its native region of Southeast Asia. Oil exploration and urbanization in this region also exacerbates mangrove loss and Nypa spread. As Nypa is difficult to distinguish from endemic mangrove species in remotely sensed data, estimates of mangrove and ecosystem services losses in Nigeria are highly uncertain. Here, we analyze multisensor satellite data with machine learning to quantify the rapid expansion of Nypa from 2015-2020 in Nigeria. Using Landsat imagery and random forest classification, we quantify total potential Nypa extent in Nigeria in 2019. We then produced a Nypa extent map using iterative combinations of Sentinel-1 SAR, Sentinel-2 MSI, and ALOS PALSAR. Random forest classifications using SAR data from ALOS and Sentinel-1 were best suited for mapping Nypa extent with similar accuracies (78% and 75% respectively). Based on data availability and accuracy, we focused our change analysis on Sentinel-1 SAR. Our results show ~28,000 ha of mangroves were converted to Nypa in Nigeria by 2020 and covered a larger extent than endemic mangroves, compounding the effect of the existing degradation and deforestation in the region. We also compared forest height and complexity estimates from GEDI (Global Ecosystem Dynamics Investigation) LiDAR to further distinguish between endemic mangroves and Nypa in three dimensions. Nypa structural variability, measured by top-of-canopy height, vegetation cover, plant area index, and foliage height diversity, was lower than that of mangroves. At current rates of Nypa expansion, the entire area of study would be invaded by Nypa by 2028, with potentially detrimental consequences to the ecosystem services provided by mangroves.

GEE↗

Learning From User Behavior: A Survey-Assist Algorithm for Longitudinal Mobility Data Collection

GPS-based travel surveys are widely used in mobility studies to gather crucial qualitative data, like purpose, transportation mode and replaced mode. However, survey response still poses a burden to users, especially in long-term mobility studies, leading to response fatigue. We explore a survey-assist strategy to ease this burden by a novel, user-level modeling approach that leverages past responses from each user to predict responses for new trips, without relying on external data sources like GIS data. We investigate three main algorithms for predicting responses: (i) clustering trips and extrapolating responses for similar trips, (ii) using random forest classification, and (iii) clustering that uses a hybrid algorithm to determine spatial structure, which is then fed as input to a classic random forest classifier. The clustering approach can flexibly predict responses for even complex qualitative survey questions; it achieved F-scores of 65%. The random forest pipeline uses architecture that restricts it to predicting three predetermined survey questions: trip purpose, mode, and replaced mode. However, it achieved F-scores of 78%. While the survey-assist approach has been implemented by several proprietary systems, to our knowledge, this is the first exploration in the academic literature. It follows that this is also the first rigorous evaluation of multiple algorithms that can implement the approach. The evaluation uses a large scale, publicly available, longitudinal dataset consisting of ~ 92k trips from 235 users over a period of roughly one and a half years. With this approach, travel surveys can be pre-filled with the predicted responses for each trip, thus streamlining the survey process for users. Combined with an active learning system that requests user input on low-confidence predictions, models can be updated and improved over time to better support the long-term collection of longitudinal qualitative data.

clustering↗

Classifying Forest Type in the National Forest Inventory Context with Airborne Hyperspectral and Lidar Data

Forest structure and composition regulate a range of ecosystem services, including biodiversity, water and nutrient cycling, and wood volume for resource extraction. Forest type is an important metric measured in the US Forest Service Forest Inventory and Analysis (FIA) program, the national forest inventory of the USA. Forest type information can be used to quantify carbon and other forest resources within specific domains to support ecological analysis and forest management decisions, such as managing for disease and pests. In this study, we developed a methodology that uses a combination of airborne hyperspectral and lidar data to map FIA-defined forest type between sparsely sampled FIA plot data collected in interior Alaska. To determine the best classification algorithm and remote sensing data for this task, five classification algorithms were tested with six different combinations of raw hyperspectral data, hyperspectral vegetation indices, and lidar-derived canopy and topography metrics. Models were trained using forest type information from 632 FIA subplots collected in interior Alaska. Of the thirty model and input combinations tested, the random forest classification algorithm with hyperspectral vegetation indices and lidar-derived topography and canopy height metrics had the highest accuracy (78% overall accuracy). This study supports random forest as a powerful classifier for natural resource data. It also demonstrates the benefits from combining both structural (lidar) and spectral (imagery) data for forest type classification.

random forest↗

Integrating Cloud-Based Workflows in Continental-Scale Cropland Extent Classification

Accurate information on cropland spatial distribution is required for global-scale assessments and agricultural land use policies. Cloud computing platforms such as Google Earth Engine (GEE) provide unprecedented opportunities for large-scale classifications of Landsat data. We developed a novel method to fuse pixel-based random forest classification of continental-scale Landsat data on GEE and an object-based segmentation approach known as recursive hierarchical segmentation (RHSeg). Using our fusion method, we produced a continental-scale cropland extent map for North America at 30m spatial resolution for the nominal year 2010. The total cropland area for North America was estimated at 275.18 million hectares (Mha). The overall accuracies of the map are>90% across the continent. This map also compares well with the United States Department of Agriculture (USDA) cropland data layer (CDL), Agriculture and Agri-food Canada (AAFC) annual crop inventory (ACI), and the Mexican government agency Servicio de Informacion Agroalimentaria y Pesquera (SIAP)'s agricultural boundaries. Furthermore, our map compared well with sub-country statistics including state-wise and county-wise cropland statistics in regression models resulting in R2 > 0.84. This key contribution paves the way for more detailed products such as crop intensity, crop type, and crop irrigation, and provides a method for creating high-resolution cropland extent maps for other countries where spatial information about croplands are not as prevalent.

Massey, Richard↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Mapping Inundation from Hurricane Florence (2018) with L-Band Synthetic Aperture Radar, Commercial Imagery, and Ancillary Data via Machine Learning Classification

During and after flooding events, mapping the extent of floodwaters aids in the distribution of resources, recovery efforts, and damage assessment practices. Development of a land cover classification system focused on mapping inundation after major hurricane events using synthetic aperture radar (SAR) data could allow for the production of near-real-time inundation mapping, enabling government and emergency response entities to get a preliminary idea of a developing situation. Complimentary optical and SAR images from domestic and foreign entities are brought together through activations of the International Charter: Space and Major Disasters to support response efforts, from true-color, near-infrared, and thermal remote sensing data obtained by NASA, NOAA, and international satellites to the collection of high-resolution true color aerial photography by NOAA and the National Geodetic Survey. In response to Hurricane Florence of 2018, NASA JPL collected numerous swaths of quad-pol L-band SAR data with the Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) instrument observing the record-setting river stages across North and South Carolina. The resulting fully-polarized SAR images allow for mapping of inundation extent at a high spatial resolution with a unique advantage over optical imaging stemming from the sensor’s ability to penetrate cloud cover and dense vegetation. In this study, true-color NOAA aerial and commercial satellite imagery are used in conjunction with four UAVSAR data swaths centered on the Lumberton and Cape Fear River basins in southeastern North Carolina to develop a Random Forest classification model focused on mapping open water and floodwater otherwise obscured by vegetation or lingering cloud cover. Ancillary building footprint, transportation route, and population data will also be incorporated into the classification scheme to estimate the societal impacts of flooding based on the proximity of features to detected inundation. Preliminary results from the Hurricane Florence case study will be discussed in addition to the limitations of available validation data for assessment of the classifier’s accuracy.

Alexander M Melancon↗

Hot, cold, or just right? An infrared biometric sensor to improve occupant comfort and reduce overcooling in buildings via closed-loop control

To improve occupant comfort and save energy in buildings, we have developed a closed-loop air conditioning (AC) sensor-controller that predicts occupant thermal sensation from the thermographic measurement of skin temperature distribution, then uses this information to reduce overcooling (cooling-energy overuse that discomforts occupants) by regulating AC output. Taking measures to protect privacy, it combines thermal-infrared (TIR) and color (visible spectrum) cameras with machine vision to measure the skin-surface temperature profile. Since the human thermoregulation system uses skin blood flow to maintain thermoneutrality, the distribution of skin temperature can be used to predict warm, neutral, and cool thermal states. We conducted a series of human-subject thermal-sensation trials in cold-to-hot environments, measuring skin temperatures and recording thermal sensation votes. We then trained random-forest classification machine-learning models (classifiers) to estimate thermal sensation from skin temperatures or skin-temperature differences. The estimated thermal sensation was input to a proportional integral (PI) control algorithm for the AC, targeting a sensation level between neutral and warm. Our sensor-controller includes a sensor assembly, server software, and client software. The server software orients the cameras and transmits images to the client software, which in turn assesses occupant skin temperature distribution, estimates occupant thermal sensation, and controls AC operation. A demonstration conducted in a conference room in an office building near Houston, TX showed that our system reduced overcooling, decreasing AC load by 42% when the room was occupied while improving occupant comfort (fraction of “comfortable” votes) by 15 percentage points.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Transitioning from Simulation to Reality: Applying Chatter Detection Models to Real-World Machining Data

Chatter, a self-excited vibration phenomenon, is a critical challenge in high-speed machining operations, affecting tool life, product surface quality, and overall process efficiency. While machine learning models trained on simulated data have shown promise in detecting chatter, their real-world applicability remains uncertain due to discrepancies between simulated and actual machining environments. The primary goal of this study is to bridge the gap between simulation-based machine learning models and real-world applications by developing and validating a Random Forest-based chatter detection system. This research focuses on improving manufacturing efficiency through reliable chatter detection by integrating Operational Modal Analysis (OMA), Receptance Coupling Substructure Analysis (RCSA), and Transfer Learning (TL). The study applies a Random Forest classification model trained on over 140,000 simulated machining datasets, incorporating techniques like Operational Modal Analysis (OMA), Receptance Coupling Substructure Analysis (RCSA), and Transfer Learning (TL) to adapt the model for real-world operational data. The model is validated against 1600 real-world machining datasets, achieving an accuracy of 86.1%, with strong precision and recall scores. The results demonstrate the model’s robustness and potential for practical implementation in industrial settings, highlighting challenges such as sensor noise and variability in machining conditions. This work advances the use of predictive analytics in machining processes, offering a data-driven solution to improve manufacturing efficiency through more reliable chatter detection.

42 ENGINEERING↗

Mapping Species Composition of Forests and Tree Plantations in Northeastern Costa Rica with an Integration of Hyperspectral and Multitemporal Landsat Imagery

An efficient means to map tree plantations is needed to detect tropical land use change and evaluate reforestation projects. To analyze recent tree plantation expansion in northeastern Costa Rica, we examined the potential of combining moderate-resolution hyperspectral imagery (2005 HyMap mosaic) with multitemporal, multispectral data (Landsat) to accurately classify (1) general forest types and (2) tree plantations by species composition. Following a linear discriminant analysis to reduce data dimensionality, we compared four Random Forest classification models: hyperspectral data (HD) alone; HD plus interannual spectral metrics; HD plus a multitemporal forest regrowth classification; and all three models combined. The fourth, combined model achieved overall accuracy of 88.5%. Adding multitemporal data significantly improved classification accuracy (p less than 0.0001) of all forest types, although the effect on tree plantation accuracy was modest. The hyperspectral data alone classified six species of tree plantations with 75% to 93% producer's accuracy; adding multitemporal spectral data increased accuracy only for two species with dense canopies. Non-native tree species had higher classification accuracy overall and made up the majority of tree plantations in this landscape. Our results indicate that combining occasionally acquired hyperspectral data with widely available multitemporal satellite imagery enhances mapping and monitoring of reforestation in tropical landscapes.

hyperspectral fusion↗

Bryce Canyon Water Resources: Monitoring Vegetation Health and Water Availability in Bryce Canyon National Park for Drought Stress Mitigation Planning

Bryce Canyon National Park is home to groundwater-dependent ecosystems (GDEs) that are threatened by a multidecadal drought and increased groundwater extraction due to a spike in tourism. These ecosystems contain unique species that are only found in areas where near-surface groundwater is present, such as aspen groves and fens. These species contribute to the high biodiversity found in Bryce Canyon, which boosts an ecosystem’s productivity and the services it provides to the park. Unfortunately, many of these GDEs are too small to identify with traditional Earth observation platforms and are difficult to physically reach for monitoring purposes. This project partnered with the National Park Service to identify springs and seeps as a proxy for GDEs within Bryce Canyon from 2013–2022. Furthermore, this project tested the feasibility of various methods to detect and monitor springs and seeps and therefore facilitate the partner’s efforts to conserve these ecologically valuable GDEs in Bryce Canyon. The team mapped groundwater discharge with high resolution National Agriculture Imagery Program (NAIP) and assessed park vegetation trends with Landsat 8 Operational Land Imager (OLI) and PlanetScope imagery. In-situ precipitation data and the Western Land Data Assimilation System (WLDAS) were used to produce time series of climatic variables. Seeps and spring locations were predicted using random forest classification and maximum entropy machine learning models.

Groundwater dependent ecosystems↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential in mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail walking and US-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986 to 2022. Using data from Landsat 5 and 8, Sentinel-2, the National Agriculture Imagery Program, and PlanetScope, the team computed vegetation indices including the Normalized Difference Vegetation Index, Normalized Difference Moisture Index, Modified Soil Adjusted Vegetation Index 2, Enhanced Vegetation Index, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on the vegetation indices and spectral bands before running k-means++ clustering and random forest classification algorithms. Between all datasets, we found the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use the end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants. The NPS’s collection of ground data for 2022–2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive spread over time.

Carson Schuetze↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗