Search NASA⌕ Search

SEARCH · Search NASA

Results for “support vector machine (SVM)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Novel Machine Learning Algorithm for Planetary Boundary Layer Height Estimation Using AERI Measurement Data

Accurately determining the height of the planetary boundary layer (PBL) is important since it can affect the climate, weather, and air quality. Ground-based infrared hyperspectral remote sensing is an effective way to obtain this parameter. Compared with radiosonde measurements, its temporal resolution is much higher. In this study, a method to retrieve the PBL height (PBLH) from the ground-based infrared hyperspectral radiance data is proposed based on machine learning. In this method, the channels that are sensitive to temperature and humidity profiles are selected as the feature vectors, and the PBLHs derived from radiosonde are taken as the true values. The support vector machine (SVM) is applied to train and test the data set, and the parameters are optimized in the process. The data set collected at the Atmospheric Radiation Measurement (ARM) program Southern Great Plains (SGP) from 2012 to 2015 is analyzed. The instruments used in this letter include Atmospheric Emitted Radiance Interferometer (AERI), Vaisala CL31 ceilometer, and radiosonde. It shows that the root mean square error (RMSE) between the PBLHs calculated by the proposed method using AERI data and those from radiosonde data can be within 370 m, and the square correlation coefficient (SCC) is greater than 0.7. Compared with the PBLHs derived from the ceilometer, it can be found that the new method is more stable and less affected by clouds.

54 ENVIRONMENTAL SCIENCES↗

Line Faults Classification Using Machine Learning on Three Phase Voltages Extracted from Large Dataset of PMU Measurements

An end-to-end supervised learning method is developed to classify transmission line faults in a twoyear field-recorded dataset that includes synchronized measurements of three-phase voltages recorded by 38 Phasor Measurement Units (PMU) sparsely located in in the US Western Grid interconnection. Statistical analysis is performed to extract features from this large dataset to train Support Vector Machine (SVM), Random Forest (RF), and eXtreme Gradient Boosting (XGBoost) classifiers initially. The training further leverages a simulated dataset from a synthetic grid with 12 PMUs to increase the number of faults of types infrequently seen in the field-recorded dataset. Training the classification models with the combined dataset resulted in a classification accuracy of 97.7%. This is a significant improvement over 89.7% to 92.5% accuracy obtained by relying on the field-recorded dataset alone.

47 OTHER INSTRUMENTATION↗

A Novel Machine Learning Algorithm for Cloud Detection Using AERI Measurement Data

Infrared hyperspectral remote sensing has been widely used in the field of meteorology. Many scientists have carried out research on inversion methods of meteorological elements such as thermodynamic profile, boundary layer height, cloud base height, etc. In this study, a method based on machine learning for cloud detection using ground-based infrared hyperspectral radiation data is proposed. The features of outliers, the cloudy and cloud-free data of Atmospheric Emitted Radiance Interferometer (AERI) radiation are extracted. The “reference values” of cloudy and cloud-free are determined based on the observation data of Vaisala CL31 ceilometer within the time range of 8 min before the corresponding time of AERI. A support vector machine (SVM) algorithm is used for training. The dataset comes from the Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site and North Slope Alaska (NSA) site from 2015 to 2017, and the ARM West Antarctic Radiation Experiment (AWARE) site in 2016 is also analyzed. The instruments used in this paper include AERI, ceilometer, etc. The experimental results reveal that the agreement of cloud detection results between the proposed algorithm and ceilometer is about 93% at each site. However, for high clouds or optically thin clouds, the agreement will decrease.

47 OTHER INSTRUMENTATION↗

Protecting Customer Privacy Through Distributed Energy Resource Anonymization

Due to their stochastic nature, the increase of Renewable Energy Resources (RERs) as a primary source of energy for power grids creates challenges regarding the reliability and resilience of the system. In order to combat these obstacles, expansion of Distributed Energy Resources (DERs) and their participation in Demand Response (DR) programs is necessary. Widespread participation requires prioritizing customer privacy and addressing concerns that may arise regarding communication between DERs and the Grid Service Provider (GSP). This paper discusses the use of flow reservation resources to split the operating cycles of DER load profiles into unique phases. The splitting of phases increases anonymization of the DERs by making it more difficult to determine the individual characteristics of the device. We discuss an example of this using simulated DER load profile data and examine the resulting effectiveness by using a machine learning algorithm for classification, called Support Vector Machine (SVM).

Distributed Energy Resource, Anonymization, Renewa↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Advancing Artificial Intelligence with Liquid Argon Neutrino Experiments (Technical Report)

The grant allowed two main contributions: 1) The development of a first successful demonstration of the employment of Optimal Transport in liquid argon time projection chamber neutrino detectors. Optimal Transport, used in other contexts and specifically with LHC calorimetric data, was adapted to address a key particle identification challenge in LArTPCs: the separation of pi0 backgrounds from single-electrons produced in charged-current electron neutrino interactions. The work, leveraging ML methods such as k-nearest-neighbor (kNN) and support-vector-machine (SVM), showed an increase in background rejection of a factor of two or more. Work is now ongoing to incorporate this development in physics analyses for LArTPC experiments and more broadly expand the use of OT in LArTPC detectors including DUNE. This work was done in collaboration with the phenomenology group led by Nathaniel Craig at UCSB. 2) The deployment of NuGraph2, a graph neural network developed for LArTPC reconstruction, in the MicroBooNE experiment. NuGraph2 uses novel graph-neural-network methods on the rather simple LArTPC inputs of reconstructed hits, greatly simplifying the workflow compared to the use of waveform or signal-deconvolved wire ROIs. The network performed particle classification and was shown to address many challenging problems in LArTPC imaging including track-shower separation and the identification of protons and charged pions from primary muons. Our group collaborated with Giuseppe Cerati (FNAL scientist) who is one of the core developers of NuGraph2 to integrate this tool in MicroBooNE’s analysis framework. This consisted in tow key contributions: a) Studying performance on real data, which came with several months of iterations because the MC-trained version of the network was found to show significant bias that our group investigated and addressed. b) Integrating the output hit labeling of NuGraph2 into the existing particle tracking and shower reconstruction code. As a result of this work led by our team NuGraph2 is now enabling a suite of new analyses which benefit from enhanced capabilities and thus broader physics reach. The grant supported primarily the salary of UCSB graduate student Chuyue “Michaelia” Fang as well as partial summer salary support for PI Caratelli. Some funds were used for travel by Michaelia to ML related schools and conferences.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unraveling the Correlation between Raman and Photoluminescence in Monolayer MoS 2 through Machine‐Learning Models

Abstract 2D transition metal dichalcogenides (TMDCs) with intense and tunable photoluminescence (PL) have opened up new opportunities for optoelectronic and photonic applications such as light‐emitting diodes, photodetectors, and single‐photon emitters. Among the standard characterization tools for 2D materials, Raman spectroscopy stands out as a fast and non‐destructive technique capable of probing material's crystallinity and perturbations such as doping and strain. However, a comprehensive understanding of the correlation between photoluminescence and Raman spectra in monolayer MoS 2 remains elusive due to its highly nonlinear nature. Here, the connections between PL signatures and Raman modes are systematically explored, providing comprehensive insights into the physical mechanisms correlating PL and Raman features. This study's analysis further disentangles the strain and doping contributions from the Raman spectra through machine‐learning models. First, a dense convolutional network (DenseNet) to predict PL maps by spatial Raman maps is deployed. Moreover, a gradient boosted trees model (XGBoost) with Shapley additive explanation (SHAP) to bridge the impact of individual Raman features in PL features is applied. Last, a support vector machine (SVM) to project PL features on Raman frequencies is adopted. This work may serve as a methodology for applying machine learning to characterizations of 2D materials.

Lu, Ang‐Yu↗

Pyrrole‐Imine Macrocycle: Self‐Organizing Cross‐Reactive Anion Receptor and Sensor

Self-organizing macrocyclic receptor-sensors for phosphorus oxyanions, phosphates, and phosphonates comprising imine moieties were prepared by condensation of dipyrrolylmethane dicarbaldehyde with diethylene triamine. The incorporation of flexible ethylene moieties endows the macrocycle with unprecedented flexibility and ability to accommodate numerous phosphorus oxyanions from orthophosphate to large anions such as ATP or phosphonate glyphosate. The anion binding was elucidated by NMR titrations, low-temperature NMR, and NOESY NMR. The incorporation of dansyl fluorophore enables sensing of anions using the fluorescence signal, whereas the changes in fluorescence intensity, width of the fluorescence band, and position of the maxima are analyte-specific and useful in recognition and identification of eleven different P-oxyanions in water. The affinity (K assoc ) for Na + salts was H 2 PO 4 − ≈ Methylphosphonate > H 2 P 2 O 7 2− > Phenylphosphonate- > Glyphosate 2− > AMP 2− > ADP 2− > ATP 2− . Interestingly, phosphonates, including methylphosphonate and glyphosate anions, were also found to display a strong affinity (K assoc ∼10 6 M −1 ) while halides, nitrate, carbonates, or hydrogen sulfate did not show a significant affinity. The determined fluorescence spectral parameters were used to classify the 12 analytes (11 anions and water) using Linear Discriminant Analysis (LDA). Quantification was performed using LDA and Support Vector Machine (SVM), and the phosphonate concentrations in unknown samples were determined with an error of 3.5% or lower.

anions↗

Risk assessment of engineering diseases of embankment–bridge transition section for railway in permafrost regions

Abstract The embankment–bridge transition section (EBTS) is one of the zones where railway diseases occur frequently in permafrost regions. Disease risk assessment of EBTSs can provide guidance for maintenance. In this study, considering the engineering geological conditions, climate characteristics, and embankment structure types along the Qinghai–Tibet Railway (QTR) as well as based on the disease inventory of the QTR from 2010 to 2019, the logistic regression (LR), support vector machine (SVM), and combination‐weight‐based gay relation analysis (GRA) were used for disease risk assessment of the EBTSs along the QTR in permafrost regions. The results indicate that the LR and SVM models have a better capability for EBTS disease prediction than the GRA model, and the SVM model can select more disease samples in relatively larger regions than the LR model. Based on the SVM and LR models, the risk level of EBTSs is divided into four classes: low‐ (29.9%), moderate‐ (39.6%), high‐ (22.1%), and very high (8.4%) risk. Finally, we selected 272 EBTSs in high‐ and very‐high‐risk classes for key observation during the maintenance of the QTR in permafrost regions. This study provides a reference for the risk assessment of railways built in permafrost regions using data‐driven methods.

Zhang, Saize↗

Predicting Search Task Difficulty through a Discrete‐Time Action Log Representation on Spectrum Kernel

ABSTRACT Predicting perceived difficulty on a web search task is an open problem in the interactive information retrieval field. A common approach to tackle it, is through features obtained from full search sessions, which are then used to train classification models. In this poster we attempt to predict perceived task difficulty at different stages of the search process. To do so, we use the spectrum kernel for support vector machine (SVM) classification. Our preliminary results suggest that by using behavioral data from the first query segment, it is possible to provide timely classifications of whether a search task is perceived as hard or easy.

Gacitúa, Daniel↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

Localized keyhole pore prediction during laser powder bed fusion via multimodal process monitoring and X-ray radiography

Systematic fault detection and control during laser powder bed fusion (L-PBF) has been a long-standing objective for system manufacturers and researchers in the additive manufacturing (AM) industry. This manuscript investigates a data fusion approach for detection of keyhole porosity formation during laser irradiation of Ti-6Al-4V substrates by concurrent recording of thermally induced optical emission measured using both off-axis and coaxial photodiode sensors, and acoustic emission. Subsurface defect formation was monitored via high-speed synchrotron X-ray imaging at 20,000 frames per second, enabling temporal registration of keyhole pore formation events to the monitoring signals at a resolution of 50 µs. We developed data fusion machine learning (ML) models for localized prediction of keyhole pore formation at various time scales ranging from 0.5 ms to 2 ms. The signal segments were featurized using two independent approaches: (1) power spectral density (PSD) and (2) highly comparative time series analysis (HCTSA) framework. The extracted features from different sensor modalities were fused together to construct a multimodal feature space and sequential feature selection was used to determine the most informative features for training the ML models. The predictive performance was evaluated for three classifying algorithms: Support Vector Machine (SVM), K-Nearest Neighbor (KNN), and Gaussian Naive Bayes (GNB). As a result, pore formation events were predicted with up to 0.95 F1-score, 1.0 recall and 0.94 accuracy. The most heavily weighted features indicate that model performance is chiefly governed by the acoustic monitoring signal, with a secondary contribution from the optical emission sensors.

36 MATERIALS SCIENCE↗

High bias machine learning for antineutrino-based safeguards for small reactors

The statistical methods used for antineutrino detection will need to be improved to effectively monitor the inventory of next-generation nuclear reactors. In this sensitivity study, we evaluate machine learning models compared to previously used statistical approaches to identify diversion scenarios in a simulated Advanced Fast Reactor (AFR)-100. A chi-square goodness-of-fit technique, which individually compares the simulated antineutrino yields to the expected antineutrino yield, resulted in precise but low diversion detection probability. Various support vector machine (SVM) models were applied with diverse training datasets to evaluate the robustness of the method towards unexpected or “unseen” diversion scenarios. Furthermore, our results indicate that while the SVM models significantly improved the detection probability of near-field antineutrino-based safeguards, up to a probability of ~0.04, for the simulated small reactor, the detection system still needs improvements to reach the 0.2 detection limit established by the International Atomic Energy Agency.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Applying NIR and MIR spectroscopy for C and soil property prediction in northern cold-region ecosystems. Which approach works better?

Here, developing reliable predictions of soil attributes is necessary to understand northern cold-region climate-soil feedback. Calibration models using near-infrared (NIR) and mid-infrared (MIR) spectroscopy were developed to predict eight commonly measured soil properties for 119 soil samples representing a range of vegetation types, parent materials, and soil types spanning >23° of latitude from southeast Alaska to the Canadian high Arctic. In order to obtain a more accurate prediction, this study compared the performance of linear and non-linear calibration techniques, including lasso regression (Lasso), support vector machine (SVM), random forest (RF) and classic partial least squares (PLS) to predict different soil properties of these soils. Comparing the four models, we noticed that their performance was quite similar for MIR overall, while NIR achieved better results with a PLS model for our dataset. PLS coupled with MIR showed a better performance for soil parameters, such as total organic carbon (TOC), total nitrogen (TN), cation exchange capacity (CEC) and clay (R-squared of 0.9, 0.81, 0.80, and 0.84) when compared with NIR (R-squared of 0.85, 0.72, 0.81 and 0.68). However, using either MIR or NIR spectroscopy, PLS predictions for bulk density (BD) and sand content were not accurate. The variable importance analysis based on the PLS model successfully estimated the relative contribution of wavelengths influencing soil property predictions most. Overall, TOC, TN, CEC and clay mineral predictions are closely related to the occurrence of specific spectral bands in the MIR region. For example, wavelengths at 2978 and 1761 cm -1 for TOC and TN, as well as at 3064 cm -1 for CEC, were selected as the most influential predictor variables. We demonstrated that MIR spectroscopy is a powerful tool for more extensive monitoring in soils of the northern cold climate region; however, NIR could be utilized for rapid estimates when the highest accuracy is not essential.

54 ENVIRONMENTAL SCIENCES↗