Search NASA⌕ Search

SEARCH · Search NASA

Results for “random forest regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

First Measurements of Ambient PM2.5 in Kinshasa, Democratic Republic of Congo and Brazzaville, Republic of Congo Using Field-calibrated Low-cost Sensors

Estimates of air pollution mortality in sub-Saharan Africa are limited by a lack of surface observations of fine particulate matter (PM2.5). Despite being large metropolises, Kinshasa, Democratic Republic of the Congo (DRC), population 14.3 million, and Brazzaville, Republic of the Congo (ROC), population 2.4 million, have no reference air pollution monitors at the time of writing. Recently, a few reference monitors have been deployed in other parts of sub-Saharan Africa, including Kampala, Uganda. A low-cost PurpleAir PM2.5 monitor was collocated next to the Kampala US Embassy BAM-1020 (Met One Beta Attenuation Monitor) starting in August 2019. Raw PurpleAir data are strongly correlated with the BAM (r(exp 2) = 0.88), but have a mean absolute error of approximately 14 μg/cu.m. Two calibration models, multiple linear regression and a random forest approach, decrease mean absolute error (MAE) from 14.3 μg/cu.m to 3.4 µg/cu.m or less and improve the the r(exp 2) from 0.88 to 0.96. Given )the similarity in climate and emissions in Kampala, we apply the collocated field correction factors to four PurpleAir sensors in Kinshasa, DRC and one in neighboring Brazzaville, ROC deployed beginning April 2018. Annual average PM2.5 for 2019 in Kinshasa is estimated at 43.5 µg/cu.m, more than 4 times higher than WHO Interim Target 1 of 10 µg/cu.m. Surface PM2.5 and aerosol optical depth were each about 40% lower during the 2020 COVID19 lockdown period compared to the same time period in 2019, which cannot be explained by changes in meteorology or wildfire emissions alone. Our results highlight the need for clean air solutions implementation in the Congo.

Celeste McFarlane↗

Comparison Study of Machine Learning Techniques to Predict Flight Energy Consumption for Advanced Air Mobility

This paper addresses the need to predict the flight energy consumption of aerial vehicles in the presence of wind using machine learning techniques. The presented work is critical to achieving sustainable and efficient operations for Advanced Air Mobility (AAM) and to evaluating the readiness of the ground-supporting energy infrastructure, e.g., electric grid and AAM portals. The flight energy consumption is described using the "energy per meter" (EPM) metric. We present a comparison study of influential machine learning techniques in predicting EPM using real-world flight test data. We presented new results of using the Decision Tree, Random Forest, and linear regression techniques, along with our previous results using the Recurrent Neural Network and Feed Forward Neural Network techniques. The comparison results show that the Linear Regression method outperforms other methods on the basis of the Mean Squared Error and error variance.

Machine Learning↗

Bhutan Agriculture: Developing a Crop Mask for Rice and Creating a Data Collection Protocol Utilizing Remotely Sensed Data in Bhutan

Rice cultivation in Bhutan has been increasingly threatened by deteriorating soil health and outbreaks of diseases and pests associated with the global change in climate patterns. Field surveys, which the national government of Bhutan has relied on to monitor remote agricultural lands, are becoming increasingly overwhelmed by growing threats to agricultural health. To address these concerns, NASA DEVELOP partnered with the Department of Agriculture of Bhutan, the Bhutan Foundation, and the Ugyen Wangchuck Institute of Conservation and Environmental Research (UWICER) and worked to increase the government of Bhutan’s agricultural monitoring capacity. Utilizing Earth observations including Landsat 8 Operational Land Imager (OLI), Sentinel-1 C-band Synthetic Aperture Radar (C-SAR), Shuttle Radar Topography Mission (SRTM), and Planet imagery, the DEVELOP team worked with NASA SERVIR and created a sampling protocol to identify rice plantations and supplement field surveys for more efficient agriculture monitoring. The analysis focused on districts Paro, Punakha, Samtse, Sarpang, Trongsa, Zhemgang, Wangdue Phodrang, and Samdrup Jongkhar in the year 2020 during the period of transplantation (June) to harvesting of rice (November). The team provided the partners with a sampling protocol for integrating NASA Earth observations into their crop monitoring methods, as well as a crop mask for rice identification and to aid crop management. The crop mask for rice was developed using the Random Forest (RF) classifier for the eight districts of Bhutan. Visually, the random forest model has proved to be more accurate and precise than the classification and Regression Tree model. Statistically, the Random Forest model was 91.8% accurate in identifying rice in Bhutan.

Yeshey Seldon↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse↗

Taxi-Out Time Prediction for Departures at Charlotte Airport Using Machine Learning Techniques

Predicting the taxi-out times of departures accurately is important for improving airport efficiency and takeoff time predictability. In this paper, we attempt to apply machine learning techniques to actual traffic data at Charlotte Douglas International Airport for taxi-out time prediction. To find the key factors affecting aircraft taxi times, surface surveillance data is first analyzed. From this data analysis, several variables, including terminal concourse, spot, runway, departure fix and weight class, are selected for taxi time prediction. Then, various machine learning methods such as linear regression, support vector machines, k-nearest neighbors, random forest, and neural networks model are applied to actual flight data. Different traffic flow and weather conditions at Charlotte airport are also taken into account for more accurate prediction. The taxi-out time prediction results show that linear regression and random forest techniques can provide the most accurate prediction in terms of root-mean-square errors. We also discuss the operational complexity and uncertainties that make it difficult to predict the taxi times accurately.

Safe and efficient surface operations↗

Online LIBS–ML Framework for Dynamic Characterization of Heterogeneous Waste-Derived Gasification Feedstocks

LIBS−ML framework for real time feedstock characterization during continuous conveyor transport Heterogeneous waste derived feedstocks (e.g., waste coal, biomass and blends) introduce rapid variability in heating value and ash chemistry that affect gasifier operation, yet conventional laboratory characterization techniques are too slow to support proactive control. To address this gap, this study reports on an online, in situ, dynamic characterization framework that couple’s laser-induced breakdown spectroscopy (LIBS) with leakage safe machine learning (ML) regression to deliver real time, decision quality predictions of gasifier relevant properties. A controlled sample matrix spanning two different waste coals, two different biomasses, and engineered blends under two particle size conditions were constructed and benchmarked using standardized laboratory analyses for proximate/ultimate properties and ash composition. LIBS spectra were acquired dynamically as material flowed on a conveyor belt, using high energy 1064 nm laser ablation and shot averaging to improve repeatability and precision. Supervised regression models (multi layer perceptron (MLP) /artificial neural network (ANN), random forest (RF), and support vector regression (SVR)) and an optimized weighted ensemble were trained on emission line feature sets using nested cross validation with Bayesian hyperparameter tuning and validated against an independent hold out set. The proposed LIBS−ML workflow achieves near laboratory predictive fidelity across parametric targets (including higher heating value (HHV), ash content, fixed carbon, sulfur, major ash forming oxides, and initial deformation temperature (IDT)), with the weighted ensemble providing a robust default predictor under dynamic measurement conditions. These results demonstrate a practical pathway for real time feedstock characterization that can enable feedforward adjustments and more resilient gasifier operation for variable quality waste derived fuels.

Biomass↗

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Using Machine Learning to Predict Cloud Turbulent Entrainment–Mixing Processes

Different turbulent entrainment–mixing mechanisms between clouds and environment are essential to cloud–related processes; however, accurate representation of entrainment–mixing in weather/climate models still poses a challenge. This study exploits the use of machine learning (ML) to address this challenge. Four ML (Light Gradient Boosting Machine [LGB], eXtreme Gradient Boosting, Random Forest, and Support Vector Regression) are examined and compared. It is found that LGB performs best, and thus is selected to understand the impact of entrainment–mixing on microphysics using simulation data from Explicit Mixing Parcel Model. Compared with traditional parameterizations, the trained LGB provides more accurate microphysical properties (number concentration and cloud droplet spectral dispersion). The partial dependences of predicted microphysics on features exhibit a strong alignment with physical mechanisms and expectations, as determined by the interpreting method, thus overcoming the limitations of the “black box” scheme. The underlying mechanisms are that the smaller number concentration and larger spectral dispersion correspond to more inhomogeneous entrainment–mixing. Specifically, number concentration after entrainment–mixing is positively correlated with adiabatic number concentration and liquid water content affected by entrainment–mixing, and inversely correlated with adiabatic volume mean radius. Spectral dispersion after entrainment–mixing is negatively correlated with liquid water content affected by entrainment–mixing, turbulent dissipation rate and relative humidity of entrained air. Sensitivity analysis further suggests that number concentration is mainly determined by cloud microphysical properties whereas spectral dispersion is influenced by both cloud microphysical properties and environmental variables. The results indicate that the LGB scheme has the potential to enhance the representation of entrainment–mixing in weather/climate models.

54 ENVIRONMENTAL SCIENCES↗

Machine-z: Rapid Machine-Learned Redshift Indicator for Swift Gamma-Ray Bursts

Studies of high-redshift gamma-ray bursts (GRBs) provide important information about the early Universe such as the rates of stellar collapsars and mergers, the metallicity content, constraints on the re-ionization period, and probes of the Hubble expansion. Rapid selection of high-z candidates from GRB samples reported in real time by dedicated space missions such as Swift is the key to identifying the most distant bursts before the optical afterglow becomes too dim to warrant a good spectrum. Here, we introduce 'machine-z', a redshift prediction algorithm and a 'high-z' classifier for Swift GRBs based on machine learning. Our method relies exclusively on canonical data commonly available within the first few hours after the GRB trigger. Using a sample of 284 bursts with measured redshifts, we trained a randomized ensemble of decision trees (random forest) to perform both regression and classification. Cross-validated performance studies show that the correlation coefficient between machine-z predictions and the true redshift is nearly 0.6. At the same time, our high-z classifier can achieve 80 per cent recall of true high-redshift bursts, while incurring a false positive rate of 20 per cent. With 40 per cent false positive rate the classifier can achieve approximately 100 per cent recall. The most reliable selection of high-redshift GRBs is obtained by combining predictions from both the high-z classifier and the machine-z regressor.

gamma-ray burst: general↗

Evaluation of the Planetary Boundary Layer Height From ERA5 Reanalysis With MOSAiC Observations Over the Arctic Ocean

The planetary boundary layer height (PBLH) is a crucial indicator reflecting the region of the atmosphere characterized by continuous turbulence. Here, we use radiosonde and surface meteorological observations (4–7 times per day, year-round measurements) during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition to derive the PBLH (PBLH MOSAiC ), and further evaluate the PBLH from the ERA5 reanalysis (PBLH ERA5 ). Comparisons between PBLH MOSAiC and PBLH ERA5 from different perspectives reveal that: (a) The overestimation of PBLH ERA5 when the sea ice concentration is >90% is significant with the centered root mean squared error reaching up to 201 m; (b) The difference between the two products is notably pronounced in cold seasons, while it is comparatively diminished in warm seasons; (c) In neutral boundary layers, differences in PBLH ERA5 are larger compared with stable and convective boundary layers. In addition, the analysis of error sources indicates that the bias of PBLH ERA5 is sensitive to the bias of vertical thermal structure and wind speed profiles in ERA5 data sets in all conditions. Finally, we find a Random Forest model effectively reduces the bias of PBLH ERA5 with the index of agreement reaching up to 0.71 in the test data set, while a multiple linear regression demonstrates comparable performance to the Random Forest model.

54 ENVIRONMENTAL SCIENCES↗

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integration of LIBS with Machine Learning for Real-Time Monitoring of Feedstock in H 2 Gasification Applications

This project, funded by the U.S. Department of Energy (DOE) – Office of Fossil Energy under Award Number DE-FE0032177, aimed to assess the feasibility of an integrated Laser-Induced Breakdown Spectroscopy (LIBS) system with advanced machine learning (ML) models for real-time characterization and potential control of hydrogen gasifiers running on waste materials as feedstocks. This was a multidisciplinary effort that encompassed the acquisition and standardized analysis of individual and blended feedstocks—comprising biomass, coal waste, and plastic waste, followed by the development of a dynamic LIBS bench system for material sample analysis and development of predictive ML models. Comprehensive laboratory testing enabled the creation of a robust elemental dataset that served as the foundation for ML model training. Techniques such as Random Forest, Gradient Boosting, Support Vector Regression, and Neural Networks were employed to predict key feedstock properties, including higher heating value (HHV), moisture content, thermal conductivity, and ash composition with high accuracy. The results were validated against experimental data and demonstrated strong potential for real-time application in gasifier control systems. The project concluded with a study on the integration of the LIBS+ML approach for gasifier control and a techno-economic analysis of the implementation of the approach into hydrogen (H 2 ) gasification systems. Dissemination of results was carried out at a DOE meeting. This work establishes a scalable framework for automated, in-line feedstock quality assessment, offering significant implications for process optimization and emissions reduction in hydrogen production.

01 COAL, LIGNITE, AND PEAT↗

Next-Level Energy Management in Manufacturing: Facility-Level Energy Digital Twin Framework Based on Machine Learning and Automated Data Collection

This research introduces an energy prediction framework at the facility level supported by automated data collection and machine learning models. It investigates whether reducing the prediction time scale allows for applying more complex machine learning techniques and if those techniques improve the prediction accuracy. The primary advantages of this framework lie in its automation of the energy prediction process and its provision of real-time energy data suitable for use in energy dashboards or digital twins. A sitewide dataset was created by combining 15 min energy and daily production data of five shops—assembly, battery, body (electric), body (gas), and paint—from a globally recognized electric vehicle manufacturer. Various machine learning models were evaluated on daily, weekly, and monthly datasets, including, in increasingly complex order: naïve, simple linear regression, net regularized generalized linear regression, principal component regression, k-nearest neighbor, random forest, and Bayesian regularized neural network. Compared to the current state-of-the-art energy consumption prediction for the industrial facility level, this research investigates more complex models and smaller time intervals for higher accuracy. The findings revealed that the more complex monthly models require a minimum of a year and a half of data to operate, while weekly models demand a year of data to achieve improved accuracy. Daily models can operate with only six months of data but exhibit poor performance due to reduced prediction accuracy of production. Key challenges identified include access to reliable, high-quality energy and production data and the initial demand for human labor.

digital twin↗

Prediction of Weather Impacts on Airport Arrival Meter Fix Capacity

This paper introduces a data driven model for predicting airport arrival capacity with a look-ahead time 2-8 hour forecast. The model is suitable for air traffic flow management by explicitly investigating the impact of convective weather on airport arrival meter fix throughput. Estimation of the arrival airport capacity under arrival meter fix flow constraints due to severe weather is an important part of Air Traffic Management (ATM). Airport arrival capacity can be reduced if one or more airport arrival meter fixes are partially or completely blocked by convective weather. When the predicted airport arrival demands exceed the predicted available airport's arrival capacity for a sustained period, Ground Delay Program (GDP) operations will be triggered by ATM system. Serious imbalances between demand and capacity occur most frequently when the airport capacity is severely degraded due to either bad airport terminal surface weather or inclement convective weather around airport arrival fixes. A model that predicts the weather-impacted airport arrival meter fix throughput may help ATM personnel to plan GDP operations more efficiently. This paper identifies the characteristics of air traffic flow across arrival meter fixes at Newark Liberty International Airport (EWR). The proposed approach, based on machine-learning methods, is developed to predict the weather impacted EWR arrival Meter Fix (MF) throughput. Sector forecast coverage is used to envision the weather impact on airport arrival MF flow, and the validation is accomplished by using Convective Weather Avoidance Model (CWAM) 0.5 to 2-hour and Collaborative Convective Forecast Product (CCFP) 4 to 8-hour look-ahead forecast data for the period of April-September in 2014. Furthermore, the regression tree ensemble learning of random forests approach for translating a sector forecast coverage model to an EWR arrival meter fix throughput model is examined. The results suggest that ATM decision makers in charge of MF flow control and GDP planning may benefit from adopting the airport arrival meter capacity prediction models to estimate the inclement weather impacts.

Wang, Yao X.↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗