Search NASA⌕ Search

SEARCH · Search NASA

Results for “support vector regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Spacebased Estimation of Moisture Transport in Marine Atmosphere Using Support Vector Regression

An improved algorithm is developed based on support vector regression (SVR) to estimate horizonal water vapor transport integrated through the depth of the atmosphere ((Theta)) over the global ocean from observations of surface wind-stress vector by QuikSCAT, cloud drift wind vector derived from the Multi-angle Imaging SpectroRadiometer (MISR) and geostationary satellites, and precipitable water from the Special Sensor Microwave/Imager (SSM/I). The statistical relation is established between the input parameters (the surface wind stress, the 850 mb wind, the precipitable water, time and location) and the target data ((Theta) calculated from rawinsondes and reanalysis of numerical weather prediction model). The results are validated with independent daily rawinsonde observations, monthly mean reanalysis data, and through regional water balance. This study clearly demonstrates the improvement of (Theta) derived from satellite data using SVR over previous data sets based on linear regression and neural network. The SVR methodology reduces both mean bias and standard deviation comparedwith rawinsonde observations. It agrees better with observations from synoptic to seasonal time scales, and compare more favorably with the reanalysis data on seasonal variations. Only the SVR result can achieve the water balance over South America. The rationale of the advantage by SVR method and the impact of adding the upper level wind will also be discussed.

Support vector regression↗

Application of Support Vector Regression to Derive Crater Depth/Diameter From Satellite Images

Through the study of impact crater shapes, one can draw important conclusions about the nature and evolution of planetary surfaces [e.g., 1-4].In particular, studying the depth (d) to diameter (D)ratio (d/D) of a population of impact craters, in combination with crater count statistics, can yield valuable insights regarding rates of erosion and burial[5]. Motivated by the great abundance of available planetary surface image data, the goal of this project is to develop an efficient way to estimate d/D from satellite images of impact craters for which stereo information is not available [6]. We set out to develop and train a machine learning algorithm to extract d/D from a dataset of synthetic impact crater images for which model d/D is known. The applications of machine learning to planetary science are numerous and diverse [7], including automatic planetary surface mapping [8] and the detection of impact craters [9]. Our algorithm makes use of Support Vector Regression (SVR), which is a type of Support Vector Machine (SVM) [10, 11].SVMs are a branch of supervised machine learning valued for their straightforward implementation and versatility in solving both classification and regression problems. In regression analysis, an SVR algorithm produces a hyperplane function to fit the training data points, as well as an ε-tube that surrounds the hyperplane. Tunable hyperparameters include the width of the ε-tube (ε) and the amount an algorithm is penalized for points which fall outside the ε-tube.

L R Chin↗

New Data-Driven Estimation of Terrestrial CO2 Fluxes in Asia Using a Standardized Database of Eddy Covariance Measurements, Remote Sensing Data, and Support Vector Regression

The lack of a standardized database of eddy covariance observations has been an obstacle for data-driven estimation of terrestrial carbon dioxide fluxes in Asia. In this study, we developed such a standardized database using 54 sites from various databases by applying consistent postprocessing for data-driven estimation of gross primary productivity (GPP) and net ecosystem carbon dioxide exchange (NEE). Data-driven estimation was conducted by using a machine learning algorithm: support vector regression (SVR), with remote sensing data for 2000 to 2015 period. Site-level evaluation of the estimated carbon dioxide fluxes shows that although performance varies in different vegetation and climate classifications, GPP and NEE at 8 days are reproduced (e.g., r (exp 2) =0.73 and 0.42 for 8 day GPP and NEE). Evaluation of spatially estimated GPP with Global Ozone Monitoring Experiment 2 sensor-based Sun-induced chlorophyll fluorescence shows that monthly GPP variations at subcontinental scale were reproduced by SVR (r (exp 2)=1.00, 0.94, 0.91, and 0.89 for Siberia, East Asia, South Asia, and Southeast Asia, respectively). Evaluation of spatially estimated NEE with net atmosphere-land carbon dioxide fluxes of Greenhouse Gases Observing Satellite (GOSAT) Level 4A product shows that monthly variations of these data were consistent in Siberia and East Asia; meanwhile, inconsistency was found in South Asia and Southeast Asia. Furthermore, differences in the land carbon dioxide fluxes from SVR-NEE and GOSAT Level 4A were partially explained by accounting for the differences in the definition of land carbon dioxide fluxes. These data-driven estimates can provide a new opportunity to assess carbon dioxide fluxes in Asia and evaluate and constrain terrestrial ecosystem models.

chlorophyll fluorescence↗

Exploring Flooded Fraction Prediction through Machine Learning Models Focusing on Medical Infrastructure in the Southeast U.S. Coastal Areas

Rising sea levels due to climate change increasingly threaten medical infrastructure through flooding. This study develops machine learning models to predict flood exposure for 11,508 medical facilities in the southeastern coastal regions of the United States by integrating datasets including meteorological, hydrological, topographic, and geological data, the Natural Risk Index, and historical flood records from NASA, HIFLD, and FEMA. Six regression models, namely Linear Regression, Support Vector Regression, Random Forest, k-Nearest Neighbors, XGBoost, and Artificial Neural Networks, are trained using 16 explanatory variables identified through literature review and correlation analysis. Data preprocessing employs the SMOGN for class imbalance and Winsorization for outliers. Model performance is evaluated using MAE, MSE, and RMSE, with Random Forest and XGBoost models achieving the highest performance (MSE of 2.58e-5 and 3.69e-5, respectively). This multifactorial approach allows the models to capture complex flood-influencing relationships, enhancing adaptability and performance across geographic regions. Future work focuses on expanding across the U.S. and developing a near real-time flood monitoring system.

Jihoon Chung↗

Taxi-Out Time Prediction for Departures at Charlotte Airport Using Machine Learning Techniques

Predicting the taxi-out times of departures accurately is important for improving airport efficiency and takeoff time predictability. In this paper, we attempt to apply machine learning techniques to actual traffic data at Charlotte Douglas International Airport for taxi-out time prediction. To find the key factors affecting aircraft taxi times, surface surveillance data is first analyzed. From this data analysis, several variables, including terminal concourse, spot, runway, departure fix and weight class, are selected for taxi time prediction. Then, various machine learning methods such as linear regression, support vector machines, k-nearest neighbors, random forest, and neural networks model are applied to actual flight data. Different traffic flow and weather conditions at Charlotte airport are also taken into account for more accurate prediction. The taxi-out time prediction results show that linear regression and random forest techniques can provide the most accurate prediction in terms of root-mean-square errors. We also discuss the operational complexity and uncertainties that make it difficult to predict the taxi times accurately.

Safe and efficient surface operations↗

Analyzing Double Delays at Newark Liberty International Airport

When weather or congestion impacts the National Airspace System, multiple different Traffic Management Initiatives can be implemented, sometimes with unintended consequences. One particular inefficiency that is commonly identified is in the interaction between Ground Delay Programs (GDPs) and time based metering of internal departures, or TMA scheduling. Internal departures under TMA scheduling can take large GDP delays, followed by large TMA scheduling delays, because they cannot be easily fitted into the overhead stream. In this paper we examine the causes of these double delays through an analysis of arrival operations at Newark Liberty International Airport (EWR) from June to August 2010. Depending on how the double delay is defined between 0.3 percent and 0.8 percent of arrivals at EWR experienced double delays in this period. However, this represents between 21 percent and 62 percent of all internal departures in GDP and TMA scheduling. A deep dive into the data reveals that two causes of high internal departure scheduling delays are upstream flights making up time between their estimated departure clearance times (EDCTs) and entry into time based metering, which undermines the sequencing and spacing underlying the flight EDCTs, and high demand on TMA, when TMA airborne metering delays are high. Data mining methods (currently) including logistic regression, support vector machines and K-nearest neighbors are used to predict the occurrence of double delays and high internal departure scheduling delays with accuracies up to 0.68. So far, key indicators of double delay and high internal departure scheduling delay are TMA virtual runway queue size, and the degree to which estimated runway demand based on TMA estimated times of arrival has changed relative to the estimated runway demand based on EDCTs. However, more analysis is needed to confirm this.

traffic management advisor↗

Space-Borne Cloud-Native Satellite-Derived Bathymetry (SDB) Models Using ICESat-2 And Sentinel-2

Shallow nearshore coastal waters provide a wealth of societal, economic and ecosystem services, yet their topographic structure is poorly mapped due to a reliance upon expensive and time intensive methods. Space‐borne bathymetric mapping has helped address these issues, but has remained largely dependent upon in situ measurements. Here we fuse ICESat‐2 lidar data with Sentinel‐2 optical imagery, within the Google Earth Engine cloud platform, to create openly available spatially continuous high‐resolution bathymetric maps at regional‐to‐national scales in Florida, Crete and Bermuda. ICESat‐2 bathymetric classified photons are used to train three Satellite Derived Bathymetry (SDB) methods, including Lyzenga, Stumpf and Support Vector Regression algorithms. For each study site the Lyzenga algorithm yielded the lowest RMSE (approx. 10‐15%) when compared with validation data. We demonstrate a means of using ICESat‐2 for both model calibration and validation, thus cementing a pathway for fully space‐borne estimates of nearshore bathymetry in shallow, clear water environments.

N. Thomas↗

Comparative Analysis of Empirical and Machine Learning Models for Chla Extraction Using Sentinel-2 and Landsat OLI Data: Opportunities, Limitations, and Challenges

Remote retrieval of near-surface chlorophyll-a (Chla) concentration in small inland waters is challenging due to substantial optical interferences of various water constituents and uncertainties in the atmospheric correction (AC) process. Although various algorithms have been developed to estimate Chla from moderate-resolution terrestrial missions (∼10–60 m), the production of both accurate distribution maps and time series of Chla has proven challenging, limiting the use of remote analyses for lake monitoring. Here, we develop a support vector regression (SVR) model, which uses satellite-derived remote-sensing reflectance spectra () from Sentinel-2 and Landsat-8 images as input for Chla retrieval in a representative eutrophic prairie lake, Buffalo Pound Lake (BPL), Saskatchewan, Canada. Validated against in situ Chla from seven ice-free seasons (N ∼ 200; 2014–2020), the SVR model outperformed both locally tuned, -fed empirical models (Normalized Difference Chlorophyll Index, 2- and 3-band, and OC3) and Mixture Density Networks (MDNs) by 15–65%, while exhibiting comparable performance to a locally trained MDN, with an error of ∼35%. Comparison of Chla retrieval models, AC processors (iCOR, ACOLITE), and radiometric products (Rayleigh-corrected, surface, and top-of-atmosphere reflectance) showed that the best Chla maps and optimal time series (up to 100 mg m−3) were produced using a coupled SVR-iCOR system.

algal blooms↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Estimating Dust and Water Ice Content of the Martian Atmosphere From THEMIS Data

Researchers at JPL and Arizona State University conducted a comparative study of three candidate algorithms for estimating components of the Martian atmosphere, using raw (uncalibrated) data collected by the Thermal Emission Imaging System (THEMIS). THEMIS is an instrument onboard the Mars Odyssey spacecraft that acquires image data in five visible and nine infrared (IR) wavelength bands. The algorithms under study used data collected from eight of the nine IR bands to estimate the dust and water ice content of the atmosphere. Such an algorithm could be used in onboard data processing to trigger other algorithms that search for features of scientific interest and to reduce the volume of data transmitted to Earth. The algorithms studied were based on regression models. In the study, the optical depths estimated by these algorithms were compared with optical depths estimated in ground-based processing using fully calibrated data from both THEMIS and the Thermal Emission Spectrometer (TES). TES is an instrument onboard the Mars Global Surveyor spacecraft that also observes the planet at infrared wavelengths, but at a lower spatial resolution than THEMIS does. Of the algorithms studied, the one that performed best was based on a Gaussian Support Vector Machine regression model. The test results indicated that this algorithm, operating on the raw data, had error rates that were within the uncertainty associated with the estimates obtained by the groundbased analysis of the fully calibrated data. This level of fidelity demonstrates that these algorithms are sufficiently accurate for use in an onboard setting.

Bandfield, Joshua↗

Influence of Global Climate on Freshwater Changes in Africa’s Largest Endorheic Basin Using Multi-Scaled Indicators

The poor investments in gauge measurements for hydro-climatic research in Africa has necessitated the need to investigate how decision makers can leverage on sophisticated spaceborne measurements to improve knowledge on surface water hydrology that can feed directly into water accounting processes and risk assessment from extreme droughts and its impacts. To demonstrate such potential, a suite of satellite earth observations (Sentinel-2, altimetry, Landsat, GRACE, and TRMM) and model data are combined with the standardized precipitation evapotranspiration index to assess the impacts of global climate on freshwater dynamics over the LCB (Lake Chad basin), Africa’s largest endorheic basin. As shown in the results of this study, the significant relationship of climate modes (AMO; r = 0.68 and 0.59; and AMM; r = 0.2 and 0.47) with drought patterns in the LCB highlights the evidence of global climate influence in the region. The significant declines in drought extents and their intensities (2004 - 2015) over LCB coincide with the rise in surface water extent of the Lake Chad during the same period. Change detection analysis of open water features in the southern pool of Lake Chad during the 2015 - 2019 period shows that on the average, only 28.4% of inundated areas within the vicinity of the Lake persisted during the period. While the association of terrestrial water storage (TWS) with model-derived surface water storage (SWS) is strongest (r = 0.89) in the catchments that provide the most nourishment to the Lake Chad, the relationship of rainfall (2002 - 2017) with TWS (r = 0.85), model TWS (r = 0.87) and SWS (r = 0.88) confirm that the LCB’s hydrology is predominantly climate-driven. This notion is further reinforced as the predicted SWS over the LCB using a support vector machine regression scheme was found to be strongly correlated (r = 0.95 at = 0.05) with observed SWS.

Sentinel-2↗

The Added Value of SMAP Soil Moisture in Crop Yield Forecasting Over Argentina

Argentina is one of the major producers and exporter of soybeans, corn, and wheat to the world market; therefore, the accurate and timely forecasting of those crops yield is crucial to national crop management and global food security. Previous studies have mainly focused on developing forecasting models for a specific crop type and location using a single source of data (e.g., vegetation indices), thus providing little insight into the forecasting models' performance on different crop types and regions. Besides, these models are based on traditional statistical regression algorithms, while more advanced machine learning approaches have not been explored. This study investigated the estimation of crop yields of three major crops (corn, soybean, and winter wheat) using Multiple Linear Regression (MLR) and Support Vector Machine (SVM), over major growing provinces in Argentina. Our models were trained and evaluated on data from 2015 to 2020, where three remote sensing products (Normalized difference vegetation index (NDVI), SMAP soil moisture, and MODIS evapotranspiration) were used as predictors. Our results indicated that accurate crop yield forecasts using the developed regression models could be made one to two months before harvest. The MLR and SVM model performance varied among different crop types, where soybean and corn exhibited better predictability compare to the wheat. In most cases, the SVM outperformed the multiple linear regression model due to its ability to capture the nonlinear and complex features of the crop-production process. The forecasted model that combines data from multiple sources outperformed single-source satellite data. The highest accuracy was obtained when the three data sources were all considered in the model development. Results also indicated that the inclusion of SMAP soil moisture improved crop yield forecasting in most provinces, and the most significant improvements occurred in the drier region.

Nazmus Shams Sazib↗

Algorithmic Detection of Elemental Biosignatures

Machine learning models that classify a sample as indicative or non-indicative of life could play an important role in life-detection missions. Their predictions result from agnostic algorithms and thereby add redundancy to judgements resulting from human expertise. Additionally, their important features can reveal the most informative measurements within the operational constraints of a life-detection mission. The Ladder of Life Detection (Neveu 2018) identifies the need for an understanding of how combinations of multiple biosignatures affect overall confidence. The present work provides a starting point to answer this need, and future work will expand the data types to obtain even more predictive combinations of features. Elemental abundance was chosen as a starting set of features due to its availability in diverse sample types, which are needed to train a generalizable model. A standardized dataset was collected, including 35 non-indicative, e.g., lunar rock, basalt; 19 indicative mixed, e.g., seawater, agricultural soil; 46 indicative non-alive, e.g., coal, chalk; and 10 indicative alive, e.g., biofilm, bacteria. This dataset could be valuable for complementary biosignature research. The samples were standardized to the same limit of detection of a simulated mission scenario. Four classification models were used: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), and Gaussian naïve Bayes (GNB). To obtain feature importances, KNN was run on three principal components of the training data and LR and SVM were run with L1 and L2 regularization. The performances and feature importances of the six model variants on 40:60 train to validation ratios were assessed with Monte Carlo simulations. ROC AUC and mean accuracy scores ranged between 82% - 94%, with sensitivity greater than specificity. For indicative of life predictors, all models had C and Ca as strong and Cl as medium; a majority of models had N, K, and P as medium. For non-indicative of life predictors, all models had Si as strong, and a majority of models had Mg, Al, and Ti as medium. Varied elements were Fe (slightly non-indicative), H (slightly indicative), O (widely varied), Na, Mn, and S. These results serve as a proof of concept and suggest important elemental signals beyond merely the CHNOPS of Earth-based life.

Algorithmic↗

Algorithmic Classification of Raman Spectra Biosignatures: Improving Life Detection Confidence

“Agnostic” biosignatures – indicators of life (or the absence of life), independent of a particular biochemistry – are increasingly considered a high standard for life detection. The Ladder of Life Detection (2018) called for investigating how combinations of independent and different potential biosignatures affect confidence. To address this gap, statistical classification of elemental abundances, isotopic fractionation, and reflectance spectroscopy (VNIR) has been implemented. Raman spectroscopy, highly desirable due to its wide availability, has the potential to improve this predictive power. This work implemented biosignature classification algorithms on Raman data alone, in preparation for combination with the other data types. Raman spectroscopy data was collected from published databases and papers as part of a manually curated dataset of “indicative” and “non-indicative of life” samples. These currently include 61 non-indicative samples (meteorites, magnetite); 3 indicative living samples (bacteria); 20 indicative non-living samples (chalk, bone); and 12 indicative mixed (with non-indicative material) samples (soil, microbial mats). Laboratory work is ongoing to characterize additional samples, particularly a greater breadth of mixed systems. Spectra were interpolated, filtered with the Savitzsky-Golay filter, and de-noised. For a preliminary examination, agnostic features were manually extracted including mean intensity, number of peaks, and mean peak width. Different peak prominences and filtering polynomials were used to refine features. Classification algorithms were implemented: k-nearest neighbors (KNN), logistic regression (LR), linear support vector machines (SVM), random forest (RF), Gaussian naïve bayes (GNB). Lastly, Monte Carlo simulations on 1,000 50%-train-test-splits were used to validate classification performance and feature significance. The preliminary feature set achieved its highest AUC of 0.52 with LR, with no strongly discriminatory features. Work to improve feature extraction, such as through deep learning with back propagation, is planned. In future work, the Raman data will be combined with the other data types, and potentially new data types such as enantiomeric excess. This project was partially supported through the NASA Ames Project EXcellence (APEX) incubator program.

Astrobiology↗

Estimation of Snow Mass Information via Assimilation of C-Band Synthetic Aperture Radar Backscatter Observations Into an Advanced and Surface Model

This study assimilated Sentinel-1 C-band backscatter observations over snow-covered terrain into the Noah-Multiparameterization land surface model using support vector machine (SVM) regression and an ensemble Kalman filter to improve the modeled terrestrial snow mass estimates. The data assimilation (DA) experiment was conducted across Western Colorado from September 2016 to August 2017. As part of the DA experiments, the impact of a rule-based update was evaluated by comparing snow water equivalent (SWE) estimates via DA (with [ DAv1 ] and without [ DAv2 ] the rule-based update) against SNOTEL SWE measurements. Results confirmed that rule-based update helped minimize SVM controllability issues, and in turn, improved the accuracy of SWE estimates relative to both open loop (OL) and DAv2 . Comparison of SWE estimates from Sentinel-1 DAv1 against SNOTEL SWE revealed that 75% of stations showed improvements in bias and correlation coefficient relative to the OL. Assimilated SWE estimates also showed statistical improvements during both the snow accumulation and snow ablation periods. However, unbiased root mean square error showed a slight increase during the snow ablation period due to the large variability in the electromagnetic response of C-band backscatter over deep and/or wet snow. Improvement of the SWE estimates also resulted in improving river discharge estimates compared to in situ measurements. River discharge using Sentinel-1 DAv1 improved the Nash–Sutcliffe efficiency at all available stations. These results suggest that physically constrained SVM can serve as an efficient observation operator for snow mass DA through explicit consideration of the first-order C-band scattering mechanisms over different terrestrial snow conditions.

Jongmin Park↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗