Search NASASearch

SEARCH · Search NASA

Results for “Support vector machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Feature Extraction and Selection Strategies for Automated Target Recognition

Several feature extraction and selection methods for an existing automatic target recognition (ATR) system using JPLs Grayscale Optical Correlator (GOC) and Optimal Trade-Off Maximum Average Correlation Height (OT-MACH) filter were tested using MATLAB. The ATR system is composed of three stages: a cursory region of-interest (ROI) search using the GOC and OT-MACH filter, a feature extraction and selection stage, and a final classification stage. Feature extraction and selection concerns transforming potential target data into more useful forms as well as selecting important subsets of that data which may aide in detection and classification. The strategies tested were built around two popular extraction methods: Principal Component Analysis (PCA) and Independent Component Analysis (ICA). Performance was measured based on the classification accuracy and free-response receiver operating characteristic (FROC) output of a support vector machine(SVM) and a neural net (NN) classifier.

computer vision

Advances in Spectral-Spatial Classification of Hyperspectral Images

Recent advances in spectral-spatial classification of hyperspectral images are presented in this paper. Several techniques are investigated for combining both spatial and spectral information. Spatial information is extracted at the object (set of pixels) level rather than at the conventional pixel level. Mathematical morphology is first used to derive the morphological profile of the image, which includes characteristics about the size, orientation, and contrast of the spatial structures present in the image. Then, the morphological neighborhood is defined and used to derive additional features for classification. Classification is performed with support vector machines (SVMs) using the available spectral information and the extracted spatial information. Spatial postprocessing is next investigated to build more homogeneous and spatially consistent thematic maps. To that end, three presegmentation techniques are applied to define regions that are used to regularize the preliminary pixel-wise thematic map. Finally, a multiple-classifier (MC) system is defined to produce relevant markers that are exploited to segment the hyperspectral image with the minimum spanning forest algorithm. Experimental results conducted on three real hyperspectral images with different spatial and spectral resolutions and corresponding to various contexts are presented. They highlight the importance of spectral–spatial strategies for the accurate classification of hyperspectral images and validate the proposed methods.

hyperspectral image

Analyzing Double Delays at Newark Liberty International Airport

When weather or congestion impacts the National Airspace System, multiple different Traffic Management Initiatives can be implemented, sometimes with unintended consequences. One particular inefficiency that is commonly identified is in the interaction between Ground Delay Programs (GDPs) and time based metering of internal departures, or TMA scheduling. Internal departures under TMA scheduling can take large GDP delays, followed by large TMA scheduling delays, because they cannot be easily fitted into the overhead stream. In this paper we examine the causes of these double delays through an analysis of arrival operations at Newark Liberty International Airport (EWR) from June to August 2010. Depending on how the double delay is defined between 0.3 percent and 0.8 percent of arrivals at EWR experienced double delays in this period. However, this represents between 21 percent and 62 percent of all internal departures in GDP and TMA scheduling. A deep dive into the data reveals that two causes of high internal departure scheduling delays are upstream flights making up time between their estimated departure clearance times (EDCTs) and entry into time based metering, which undermines the sequencing and spacing underlying the flight EDCTs, and high demand on TMA, when TMA airborne metering delays are high. Data mining methods (currently) including logistic regression, support vector machines and K-nearest neighbors are used to predict the occurrence of double delays and high internal departure scheduling delays with accuracies up to 0.68. So far, key indicators of double delay and high internal departure scheduling delay are TMA virtual runway queue size, and the degree to which estimated runway demand based on TMA estimated times of arrival has changed relative to the estimated runway demand based on EDCTs. However, more analysis is needed to confirm this.

traffic management advisor

Nominal 30-M Cropland Extent Map of Continental Africa by Integrating Pixel-Based and Object-Based Algorithms Using Sentinel-2 and Landsat-8 Data on Google Earth Engine

A satellite-derived cropland extent map at high spatial resolution (30-m or better) is a must for food and water security analysis. Precise and accurate global cropland extent maps, indicating cropland and non-cropland areas, is a starting point to develop high-level products such as crop watering methods (irrigated or rainfed), cropping intensities (e.g., single, double, or continuous cropping), crop types, cropland fallows, as well as assessment of cropland productivity (productivity per unit of land), and crop water productivity (productivity per unit of water). Uncertainties associated with the cropland extent map have cascading effects on all higher-level cropland products. However, precise and accurate cropland extent maps at high spatial resolution over large areas (e.g., continents or the globe) are challenging to produce due to the small-holder dominant agricultural systems like those found in most of Africa and Asia. Cloud-based Geospatial computing platforms and multi-date, multi-sensor satellite image inventories on Google Earth Engine offer opportunities for mapping croplands with precision and accuracy over large areas that satisfy the requirements of broad range of applications. Such maps are expected to provide highly significant improvements compared to existing products, which tend to be coarser in resolution, and often fail to capture fragmented small-holder farms especially in regions with high dynamic change within and across years. To overcome these limitations, in this research we present an approach for cropland extent mapping at high spatial resolution (30-m or better) using the 10-day, 10 to 20-m, Sentinel-2 data in combination with 16-day, 30-m, Landsat-8 data on Google Earth Engine (GEE). First, nominal 30-m resolution satellite imagery composites were created from 36,924 scenes of Sentinel-2 and Landsat-8 images for the entire African continent in 2015-2016. These composites were generated using a median-mosaic of five bands (blue, green, red, near-infrared, NDVI) during each of the two periods (period 1: January-June 2016 and period 2: July-December 2015) plus a 30-m slope layer derived from the Shuttle Radar Topographic Mission (SRTM) elevation dataset. Second, we selected Cropland/Non-cropland training samples (sample size 9791) from various sources in GEE to create pixel-based classifications. As supervised classification algorithm, Random Forest (RF) was used as the primary classifier because of its efficiency, and when over-fitting issues of RF happened due to the noise of input training data, Support Vector Machine (SVM) was applied to compensate for such defects in specific areas. Third, the Recursive Hierarchical Segmentation (RHSeg) algorithm was employed to generate an object-oriented segmentation layer based on spectral and spatial properties from the same input data. This layer was merged with the pixel-based classification to improve segmentation accuracy. Accuracies of the merged 30-m crop extent product were computed using an error matrix approach in which 1754 independent validation samples were used. In addition, a comparison was performed with other available cropland maps as well as with LULC maps to show spatial similarity. Finally, the cropland area results derived from the map were compared with UN FAO statistics. The independent accuracy assessment showed a weighted overall accuracy of 94, with a producers accuracy of 85.9 (or omission error of 14.1), and users accuracy of 68.5 (commission error of 31.5) for the cropland class. The total net cropland area (TNCA) of Africa was estimated as 313 Mha for the nominal year 2015.

Cropland mapping; cropland areas; 30-m; Landsat-8;

Exploring Spatiotemporal Relations Between Soil Moisture, Precipitation, and Streamflow for a Large Set of Watersheds Using Google Earth Engine

An understanding of streamflow variability and its response to changes in climate conditions is essential for water resource planning and management practices that will help to mitigate the impacts of extreme events such as floods and droughts on agriculture and other human activities. This study investigated the relationship between precipitation, soil moisture, and streamflow over a wide range of watersheds across the United States using Google Earth Engine (GEE). The correlation analyses disclosed a strong association between precipitation, soil moisture, and streamflow, however, soil moisture was found to have a higher correlation with the streamflow relative to precipitation. Results indicated different strength of the association depends on the watershed classes and lag times assessments. The perennial watersheds showed higher coherence compared to intermittent watersheds. Previous month precipitation and soil moisture have a stronger influence on the current month streamflow, particularly in the snow-dominated watersheds. Monthly streamflow forecasting models were developed using an autoregressive integrated moving average (ARIMA) and support vector machine (SVM). The results showed that the SVM model generally performed better than the ARIMA model. Overall streamflow forecasting model performance varied considerably among watershed classes, and perennial watersheds tend to exhibit better predictably compared to intermittent watersheds due to lower streamflow variability. The SVM models with precipitation and streamflow inputs performed better than those with streamflow input only. Results indicated that the inclusion of antecedent root-zone soil moisture improved the streamflow forecasting in most of the watersheds, and the largest improvements occurred in the intermittent watersheds. In conclusion, this work demonstrated that knowing the relationship between precipitation, soil moisture, and streamflow in different watershed classes will enhance the understanding of the hydrologic process and can be effectively utilized in improving streamflow forecasting for better satellite-based water resource management strategies.

Nazmus Sazib

Influence of Global Climate on Freshwater Changes in Africa’s Largest Endorheic Basin Using Multi-Scaled Indicators

The poor investments in gauge measurements for hydro-climatic research in Africa has necessitated the need to investigate how decision makers can leverage on sophisticated spaceborne measurements to improve knowledge on surface water hydrology that can feed directly into water accounting processes and risk assessment from extreme droughts and its impacts. To demonstrate such potential, a suite of satellite earth observations (Sentinel-2, altimetry, Landsat, GRACE, and TRMM) and model data are combined with the standardized precipitation evapotranspiration index to assess the impacts of global climate on freshwater dynamics over the LCB (Lake Chad basin), Africa’s largest endorheic basin. As shown in the results of this study, the significant relationship of climate modes (AMO; r = 0.68 and 0.59; and AMM; r = 0.2 and 0.47) with drought patterns in the LCB highlights the evidence of global climate influence in the region. The significant declines in drought extents and their intensities (2004 - 2015) over LCB coincide with the rise in surface water extent of the Lake Chad during the same period. Change detection analysis of open water features in the southern pool of Lake Chad during the 2015 - 2019 period shows that on the average, only 28.4% of inundated areas within the vicinity of the Lake persisted during the period. While the association of terrestrial water storage (TWS) with model-derived surface water storage (SWS) is strongest (r = 0.89) in the catchments that provide the most nourishment to the Lake Chad, the relationship of rainfall (2002 - 2017) with TWS (r = 0.85), model TWS (r = 0.87) and SWS (r = 0.88) confirm that the LCB’s hydrology is predominantly climate-driven. This notion is further reinforced as the predicted SWS over the LCB using a support vector machine regression scheme was found to be strongly correlated (r = 0.95 at = 0.05) with observed SWS.

Sentinel-2

Passive Microwave Brightness Temperature Assimilation to Improve Snow Mass Estimation across Complex Terrain in Pakistan, Afghanistan, and Tajikistan

An ensemble Kalman filter is used to assimilate Advanced Microwave Scanning Radiometer-2 (AMSR2) observations of passive microwave (PMW) brightness temperatures (spectral differences, ΔT b ) into land surface model estimates of snow mass over northwestern high mountain Asia (HMA). Trained support vector machines serve as the observation operator and map the geophysical modeled variables into ΔT b space within the assimilation framework. Evaluation of the assimilation routine is carried out through comparison of assimilated snow mass estimates with an in situ dataset. The assimilation framework helps improve the land surface model estimates through PMW ΔT b assimilation, particularly in terms of decreasing the domain-wide bias. The assimilation framework proved more effective during the (dry) snow accumulation season and decreased the bias and root-mean-square error (RMSE) in snow mass estimates at 76% and 58% of the comparative pixels, respectively. During the snow ablation season, the PMW brightness temperature signal contained less information related to snow mass due to the presence of other concurrent geophysical features that effectively serve as noise during the snow mass update. The utilization of PMW ΔT b for accurate snow mass estimation in complex terrain such as HMA is dependent on a multitude of factors for optimal results; however, it does add utility to the land surface model if the relevant pitfalls are taken into consideration prior to the state variable update.

Jawairia Ahmad

Estimation of Snow Mass Information via Assimilation of C-Band Synthetic Aperture Radar Backscatter Observations Into an Advanced and Surface Model

This study assimilated Sentinel-1 C-band backscatter observations over snow-covered terrain into the Noah-Multiparameterization land surface model using support vector machine (SVM) regression and an ensemble Kalman filter to improve the modeled terrestrial snow mass estimates. The data assimilation (DA) experiment was conducted across Western Colorado from September 2016 to August 2017. As part of the DA experiments, the impact of a rule-based update was evaluated by comparing snow water equivalent (SWE) estimates via DA (with [ DAv1 ] and without [ DAv2 ] the rule-based update) against SNOTEL SWE measurements. Results confirmed that rule-based update helped minimize SVM controllability issues, and in turn, improved the accuracy of SWE estimates relative to both open loop (OL) and DAv2 . Comparison of SWE estimates from Sentinel-1 DAv1 against SNOTEL SWE revealed that 75% of stations showed improvements in bias and correlation coefficient relative to the OL. Assimilated SWE estimates also showed statistical improvements during both the snow accumulation and snow ablation periods. However, unbiased root mean square error showed a slight increase during the snow ablation period due to the large variability in the electromagnetic response of C-band backscatter over deep and/or wet snow. Improvement of the SWE estimates also resulted in improving river discharge estimates compared to in situ measurements. River discharge using Sentinel-1 DAv1 improved the Nash–Sutcliffe efficiency at all available stations. These results suggest that physically constrained SVM can serve as an efficient observation operator for snow mass DA through explicit consideration of the first-order C-band scattering mechanisms over different terrestrial snow conditions.

Jongmin Park

A Census of Young Stellar Objects in Two Line-of-Sight Star-Forming Regions Toward IRAS 22147+5948 in the Outer Galaxy

Context. Star formation in the outer Galaxy, namely, outside of the Solar circle, has not been extensively studied in part due to the low CO brightness of the molecular clouds linked with the negative metallicity gradient. Recent infrared surveys provide an overview of dust emission in large sections of the Galaxy, but they suffer from cloud confusion and poor spatial resolution at far-infrared wavelengths. Aims. We aim to develop a methodology to identify and classify young stellar objects (YSOs) in star-forming regions in the outer Galaxy and use it to resolve a long-standing disparity in terms of the distance and evolutionary status of IRAS 22147+5948. Methods. We used a support vector machine learning algorithm to complement standard color–color and color–magnitude diagrams in our search for YSOs in the IRAS 22147 region, based on publicly available data from the Spitzer Mapping of the Outer Galaxy survey. The agglomerative hierarchical clustering algorithm was used to identify clusters. Then the physical properties of individual YSOs were calculated. The distances were determined using CO 1–0 from the Five College Radio Astronomy Observatory survey. Results. We identified 13 Class I and 13 Class II YSO candidates using the color–color diagrams, along with an additional 2 and 21 sources, respectively, using the applied machine learning techniques. The spectral energy distributions of 23 sources were modeled with a star and a passive disk, corresponding to Class II objects. The models of three sources include envelopes that are typical for Class I objects. The objects were grouped into two clusters located at a distance of 2:2 kpc and 5 clusters at 5:6 kpc. The spatial extent of CO, radio continuum, and dust emission confirms the origin of YSOs in two distinct star-forming regions along a similar line of sight. Conclusions. The outer Galaxy may serve as a unique laboratory for exploring star formation across environments, on the condition that complementary methods and ancillary data are used to properly account for cloud confusion and distance uncertainties.

Agata Karska

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. It affects the efficiency of the National Airspace System (NAS) and both surface and airspace operations can benefit from better understanding future runway configurations. In this paper, we present a comprehensive implementation of predictive models for runway configuration estimation from large volumes of historical data. Specifically, operational data from two full years (2018 and 2019) is collected, analyzed, and fused together to build the data product used in this work. The data set differs from prior work in the field in terms of its scope, resolution, and variety of factors collected and considered. Meteorological data is collected from two different sources – current weather conditions from METAR (Meteorological Terminal Aviation Routine Weather Report) and forecast weather conditions from Localized Aviation MOS Program (LAMP). Operational data from the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) related to scheduled and actual number of arrivals and departures, average taxi times, etc. are collected. NASA’s Sherlock Data Warehouse is used to identify critical information such as go-arounds, and other events that might impact RCM decision-making. All data is collected and aggregated over 15-minute intervals throughout the two years. This provides a resolution like the timescales that might be necessary for runway configuration management decision-making. A variety of supervised learning algorithms are tested including Support Vector Machine, Random Forest, Gradient Boosting, etc. including tuning of the model hyperparameters. The modeling process is applied and presented on two representative U.S. airports – Charlotte Douglas International Airport (KCLT) and Denver International Airport (KDEN). The two airports present different levels of complexity in terms of the total number of configurations used and provide a balanced perspective on the generalizability of the developed approach to other airports in the NAS. Initial results are promising (F1 score of 0.91 at KCLT and 0.83 at KDEN) for data in the test set. The final paper will contain a comprehensive comparison between different models and model building strategies as well as further refined results. Most important predictors for each airport will be identified along with a discussion and recommendations on adapting the framework to other scenarios.

Tejas G Puranik

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley

Document Classification Techniques for Aviation Letters of Agreement

Often when working with technical documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques, including Transfer Learning, for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

Aayushi Batra

Enabling Intelligent Data Downlink Prioritization of In-Situ Observations through Generalizable and Computationally Inexpensive Anomaly Detection

High-fidelity measurements of magnetic fields and other observed properties, such as energetic particle fluxes, are a necessary component to our understanding of the highly dynamic near-Earth space environment. As our desire to study smaller-scale phenomena such as shocks and dipolorizations has increased, we have been driven to take and telemeter measurements at higher cadences. Unfortunately, many missions are unable to downlink all their captured data due to the well-known data transmission bottleneck at the DSN. These missions must then prioritize their high-cadence data such that the most scientifically useful intervals are transmitted. One simple prioritization technique uses the spacecraft position to telemeter data from only the region of interest. Although easy to implement, this method does not leverage the available scientific data and can omit intervals of useful scientific data when they lie outside the region of interest. The Magnetospheric Multiscale Mission (MMS) uses mission-specific parameterization of several data products to automatically prioritize scientifically useful intervals. Then, MMS verifies the automatically selected intervals by having a domain expert manually select intervals for downlink. The overall complexity required by this technique make it prohibitive for deployment on low-cost platforms (i.e., CubeSats) or on future missions featuring large constellations of satellites such as the Geospace Dynamics Constellation (GDC). We present preliminary results for a simple, generic, and data-driven method of downlink prioritization for magnetic field (and other) measurements. Specifically, Principal Components Analysis (PCA) and One-Class Support Vector Machines (OC-SVMs) are used to detect intervals containing anomalous activity, which can then be prioritized for subsequent downlink. The computational simplicity of this algorithm makes it an excellent candidate for implementation on spaceflight hardware, as well as provide generalizability to a broad range of missions and data products. Initial analysis of this technique has been performed using magnetic field measurements from the Magnetospheric Multiscale Mission and CASSIOP, where it automatically identified scientifically interesting intervals containing Alfvén waves and EMIC activity.

Matthew G. Finley

Document Classification Techniques for Aviation Letters of Agreement

Often when working with historic air traffic management (ATM) documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

ATM

Document Classification Techniques for Aviation Letters of Agreement

Often when working with historic air traffic management (ATM) documents, it is helpful to classify them into specific categories. In this paper, we conduct a thorough review of natural language processing techniques to perform this classification task on Letters of Agreement (LOAs), technical aviation documents outlining rules for utilizing US airspace. We evaluate multiple techniques for representing the text in the documents as embeddings: unigram and bigram Term Frequency Inverse Document Frequency (TFIDF), Word2Vec, Doc2Vec, GloVe and RoBERTa. We investigate a wide range of classification models: K-Nearest Neighbors, Random Forest, Support Vector Machines (SVM), Logistic Regression, Naive Bayes, Feed-Forward Neural Network, Convolutional Neural Networks (CNNs) and Long-Short Term Memory (LSTM). By comparing the different methods, we found the best overall approach for our task was to use unigram TFIDF representations with SVM while also gaining insight into how the other methodologies performed on a small technical datasets.

ATM

GRB Progenitor Classification from Gamma-Ray Burst Prompt and Afterglow Observations

Using an established classification technique, we leverage standard observations and analyses to predict the progenitors of gamma-ray bursts (GRBs). This technique, utilizing support vector machine (SVM) statistics, provides a more nuanced prediction than the previous two-component Gaussian mixture in duration of the prompt gamma-ray emission. Based on further covariance testing from Fermi-GBM, Swift-BAT, and Swift-XRT data, we find that our classification based only on prompt emission properties gives perspective on the recent evidence that mergers and collapsars exist in both “long” and “short” GRB populations.

P Nuessle

Active Learning with Irrelevant Examples

Active learning algorithms attempt to accelerate the learning process by requesting labels for the most informative items first. In real-world problems, however, there may exist unlabeled items that are irrelevant to the user's classification goals. Queries about these points slow down learning because they provide no information about the problem of interest. We have observed that when irrelevant items are present, active learning can perform worse than random selection, requiring more time (queries) to achieve the same level of accuracy. Therefore, we propose a novel approach, Relevance Bias, in which the active learner combines its default selection heuristic with the output of a simultaneously trained relevance classifier to favor items that are likely to be both informative and relevant. In our experiments on a real-world problem and two benchmark datasets, the Relevance Bias approach significantly improved the learning rate of three different active learning approaches.

machine learning

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION