Search NASASearch

SEARCH · Search NASA

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Community‐Level Metabolic Shifts Following Land Use Change in the Amazon Rainforest Identified by a Supervised Machine Leaning Approach

ABSTRACT The Amazon rainforest has been subjected to high rates of deforestation, mostly for pasturelands, over the last few decades. This change in plant cover is known to alter the soil microbiome and the functions it mediates, but the genomic changes underlying this response are still unresolved. In this study, we used a combination of deep shotgun metagenomics complemented by a supervised machine learning approach to compare the metabolic strategies of tropical soil microbial communities in pristine forests and long‐term established pastures in the Amazon. Machine learning‐derived metagenome analysis indicated that microbial community structures (bacteria, archaea and viruses) and the composition of protein‐coding genes were distinct in each plant cover type environment. Forest and pasture soils had different genomic diversities for the above three taxonomic groups, characterised by their protein‐coding genes. These differences in metagenome profiles in soils under forests and pastures suggest that metabolic strategies related to carbohydrate and energy metabolisms were altered at community level. Changes were also consistent with known modifications to the C and N cycles caused by long‐term shifts in aboveground vegetation and were also associated with several soil physicochemical properties known to change with land use, such as the C/N ratio, soil temperature and exchangeable acidity. In addition, our analysis reveals that these alterations in land use can also result in changes to the composition and diversity of the soil DNA virome. Collectively, our study indicates that soil microbial communities shift their overall metabolic strategies, driven by genomic alterations observed in pristine forests and long‐term established pastures with implications for the C and N cycles.

carbon and nitrogen cycles

ldrd_virus_work

This is a Python code base that takes openly-available genetic information on known viruses and performs supervised machine learning and feature importance analysis on the relationship of the viral genomes to the competence to infect humans or bind to a specific host cell receptor.

Reddy, Tyler [LANL]

Spectral Data Fusion From Handheld Laser-Induced Breakdown Spectroscopy (LIBS) and X-ray Fluorescence (XRF) Analyzers for Improved Detection of Cerium in a Simulated Dispersal Accident

Here, this work implements a mid-level data fusion methodology on spectral data from handheld X-ray fluorescence and laser-induced breakdown spectroscopy analyzers to quantify plutonium surrogate (CeO 2 ) contamination in soil samples for the first time. Spectral data from each analyzer were used independently to train supervised machine learning regressions to predict Ce concentration. Fused features from both data sets were then used to train the same models, comparing prediction performance by evaluating model precision and sensitivity. Fusing principal component scores from the two sensors yielded an order of magnitude improvement in precision and sensitivity of predictions made with an artificial neural network, compared to predictions made by models trained on independent sensor data. As a result, a boosted ensemble trained on the fused spectral features yielded an ideal predictor with root-mean-squared error on the order of 10 –6 and calculated limit of detection order 10 –5 wt %.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Real-time neutron multiplicity and source localization for criticality safety during fuel debris removal

Advancing neutron detection and analysis techniques for complex radiation environments is an ongoing focus in nuclear instrumentation and monitoring. This proposal presents research and development of a generalized real-time neutron monitoring and analysis system, applicable to any detector capable of producing time-tagged neutron count data. While the work is demonstrated using the Neutron Multiplication Analysis Detector (NoMAD), a modular 15-tube helium-3 (He-3) array, due to its availability, spatial resolution, and flexible deployment, the methods developed are extensible to other systems, including organic scintillators and fast digital detectors. This research investigates two complementary analytical techniques for real-time characterization of neutron emitting sources: neutron multiplicity estimation based on the Hage-Cifarelli formalism and spatial localization using supervised machine learning applied to spatial count rate patterns. These methods are designed to operate under dynamic, evolving conditions such as fuel debris retrieval or reactor startup, where neutron-emitting material geometries may be partially unknown or changing over time. By integrating statistical neutron emission data with spatial localization, this research aims to develop and evaluate methods for real time neutron monitoring, source characterization, and material verification. Key contributions include implementation of a low-latency data pipeline for continuous neutron multiplicity analysis, development and validation of machine learning models for spatial inference, and experimental evaluation of system performance under variable measurement conditions. The outcomes are intended to support applications in nuclear safeguards, verification, emergency response, and reactor startup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

A Model Based Approach to Extract Health Information from Textual Data

In current nuclear power plants (NPPs) a large amount of condition-based data is being generated and stored to assess and monitor component health and performance. The format of this data can be either numeric (e.g., pump vibration data) or textual (e.g., condition report which assess component health). While assessing component health from numeric data can be performed with a large variety of methods, the extraction of information from textual data still remains a challenge. Natural language processing (NLP) methods are starting to be deployed in current NPPs mainly to filter out incident reports (IRs) that are not safety related by employing supervised machine learning methods. However, these methods do not really provide the quantitative information that might be contained in IRs. This paper presents an approach to extract information from textual data (e.g., from IRs, maintenance reports) that is based on NLP data analytics methods coupled with model-based system engineer (MBSE) models. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence; such analysis includes: part of speech (POS) tagging (i.e., identification of grammatic elements of each string - e.g., nouns, verbs), named entity recognition (i.e., identification of text entities - e.g., names, dates, events), and relation extraction (e.g., coreference resolution). On the other hand, semantic analysis is designed to analyze the logic structure of a sentence. Through a specific set of rules, our methods can identify whether a sentence contains health information of a component (e.g., degraded performance, anomaly behavior) or the causal relationship between two events (i.e., a cause-effect pair). An innovative element of our approach is that semantic analysis relies on MBSE models to identify links between textual elements. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. This paper presents in detail how the integration of NLP methods and MBSE models is performed. Few analysis examples focusing on centrifugal pumps are presented.

97 - MATHEMATICS AND COMPUTING

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics

A machine-learning approach to measure 3D sample properties from 2D Transmission Electron Microscopy images

Transmission Electron Microscopy (TEM) is a powerful tool for the characterization of materials at the nanoscale; however, its inherent two-dimensional (2D) nature poses significant challenges to accurately measure three-dimensional (3D) properties. We introduce a supervised machine-learning model that predicts 3D structural information, such as sample thickness and curvature, from a series of conventional 2D TEM images. The model, a U-Net convolutional neural network, is trained on a large synthetic dataset generated from dynamical diffraction simulations that model TEM’s complex, nonlinear image formation, accounting for sample thickness and curvature. This physically realistic framework enables exploration of a broad parameter space impractical to sample experimentally. We demonstrate that the trained model has accurate predictions for experimental single-crystal silicon samples, achieving performance comparable to established measurement techniques. This work highlights the critical role of robust, simulation-based training in overcoming the limitations of real-world imaging artifacts and inconsistent sample geometries. By integrating machine learning with numerical simulations, we offer an efficient and scalable framework for quantitative TEM analysis, paving the way for more sophisticated 3D characterization of complex materials.

Dynamical diffraction

Machine Learning for Multipactor Susceptibility Prediction in Planar RF Gaps

Multipactor discharge is a nonlinear electron avalanche that limits the performance of high-power radio-frequency (RF) and vacuum electronic devices. Predicting multipactor susceptibility traditionally relies on Monte Carlo or particle-in-cell (PIC) simulations, which become computationally expensive for large parametric studies. In this work, we present a supervised machine-learning (ML) framework for prediction of multipactor susceptibility in a two-surface planar geometry. The models are trained using high-fidelity PIC simulation generated susceptibility data and learn the relationship between operational parameters, geometry, and material-dependent secondary electron emission properties. The proposed approach enables rapid reconstruction of susceptibility charts while preserving the physical structure of multipactor growth regions.

43 PARTICLE ACCELERATORS

Physics-guided dual implicit neural representations for source separation

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions, such as background and signal distortions, that can obscure the physically relevant information of interest. To address this, we have developed a self-supervised machine-learning approach for source separation using a dual implicit neural representation framework that jointly trains two neural networks: one for approximating distortions of the physical signal of interest and the other for learning the effective background contribution. Our method learns directly from the raw data by minimizing a reconstruction-based loss function without requiring labeled data or pre-defined dictionaries. We demonstrate the effectiveness of our framework by considering a challenging case study involving large-scale simulated, as well as experimental, momentum-energy-dependent inelastic neutron scattering data in a four-dimensional parameter space, characterized by heterogeneous background contributions and unknown distortions to the target signal. The method is found to successfully separate physically meaningful signals from a complex or structured background even when the signal characteristics vary across all four dimensions of the parameter space. An analytical approach that informs the choice of the regularization parameter is presented. Our method offers a versatile framework for addressing source separation problems across diverse domains, ranging from superimposed signals in astronomical measurements to structural features in biomedical image reconstructions.

47 OTHER INSTRUMENTATION

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics

Improving neutrino oscillation measurements through event classification

Precise neutrino energy reconstruction is essential for next-generation long-baseline oscillation experiments, yet current methods remain limited by large uncertainties in neutrino-nucleus interaction modeling. Even so, it is well established that different interaction channels produce systematically varying amounts of missing energy and therefore yield different reconstruction performance–information that standard calorimetric approaches do not exploit. We introduce a strategy that incorporates this structure by classifying events according to their underlying interaction type prior to energy reconstruction. Using supervised machine-learning techniques trained on labeled generator events, we leverage intrinsic kinematic differences among quasielastic scattering, meson-exchange current, resonance production, and deep-inelastic scattering processes. A cross-generator testing framework demonstrates that this classification approach is robust to microphysics mismodeling and, when applied to a simulated DUNE 𝜈 𝜇 disappearance analysis, yields improved accuracy and sensitivity at the 10%–20% level. These results highlight a practical path toward reducing reconstruction-driven systematics in future oscillation measurements.

Ellis, Sebastian A. R. [King's College, London (Un

Machine learning-powered data cleaning for LEGEND: a semi-supervised approach using affinity propagation and support vector machines

Neutrinoless double-beta decay ($0\nu\beta\beta$) is a rare nuclear process that, if observed, will provide insight into the nature of neutrinos and help explain the matter-antimatter asymmetry in the Universe. The large enriched germanium experiment for neutrinoless double-beta decay (LEGEND) will operate in two phases to search for $0\nu\beta\beta$. The first (second) stage will employ 200 (1000) kg of High-Purity Germanium (HPGe) enriched in 76 Ge to achieve a half-life sensitivity of 10 27 (10 28 ) years. In this study, we present a semi-supervised data-driven approach to remove non-physical events captured by HPGe detectors powered by a novel artificial intelligence model. We utilize affinity propagation to cluster waveform signals based on their shape and a support vector machine to classify them into different categories. We train, optimize, and test our model on data taken from a natural abundance HPGe detector installed in the Full Chain Test experimental stand at the University of North Carolina at Chapel Hill. We demonstrate that our model yields a maximum sacrifice of physics events of $0.024 ^{+0.004}_{-0.003} \%$ after data cleaning. Our model is being used to accelerate data cleaning development for LEGEND-200 and will serve to improve data cleaning procedures for LEGEND-1000.

artificial intelligence

Co-Firing Switchgrass and Waste Coal in A Power Plant: A Techno-Economic and Life Cycle Evaluation for The Ohio River Valley (SWITCH) (Final Technical Report for Ohio State/FE0032204)

Abandoned coal mine lands (AMLs) represent one of the most persistent environmental challenges in the United States. Prior to the enactment of the Surface Mining Control and Reclamation Act (SMCRA) in 1977, coal mining operations were not legally required to reclaim disturbed lands, leaving behind approximately 500,000 AML sites nationwide. These sites pose severe environmental and health risks, including acid mine drainage, soil and water contamination, and spontaneous combustion of waste coal piles. Millions of Americans live within one mile of these AMLs, underscoring the urgency of remediation. Traditional reclamation practices, such as planting cool-season grasses, often fail to fully restore ecological function or leverage the economic potential of these lands. This project addressed these challenges by developing integrated strategies for resource recovery, land reclamation, and sustainable energy production. This project evaluated an integrated strategy to convert this liability into an opportunity by recovering waste coal and co-firing it with switchgrass (Panicum virgatum L.) cultivated on reclaimed or marginal AML areas in existing coal-fired power plants. Switchgrass not only provides a renewable feedstock but also aids in land reclamation and carbon sequestration. 1) Remote Sensing and Machine Learning for Waste Coal Identification Using Sentinel-2 satellite imagery and supervised classification, we applied four machine learning models to detect historical waste coal piles. Random Forest achieved the highest accuracy (precision: 86%, recall: 77%). Time-series analysis revealed gradual vegetation recovery since 1986, indicating natural reclamation processes in historical sites, while active mining areas showed ongoing disturbance. This workflow enables scalable monitoring and prioritization of reclamation efforts. 2) UAS-Based Stockpile Volume Estimation To quantify recoverable waste coal, we evaluated Unmanned Aerial Systems (UAS) equipped with Light Detection and Ranging (LiDAR) and multispectral sensors. Structure-from-Motion (SfM) photogrammetry combined with interpolated Digital Terrain Models (DTMs) achieved strong agreement with LiDAR reference volumes (Root Mean Square Error (RMSE) ≈147 m 3 , Mean Absolute Percentage Error (MAPE) ≈2%). Sensitivity analysis confirmed that spatial resolution significantly influences accuracy, emphasizing the need for high-resolution data for precise volume estimation. This approach offers a scalable, cost-effective, and accurate alternative to conventional ground-based surveys. 3) Switchgrass Cultivation for Bioenergy and Water Quality Improvement We assessed the hydrological and water quality impacts of converting AMLs to switchgrass production areas using the Soil and Water Assessment Tool (SWAT). Results showed that converting 10% of the watershed area into the switchgrass production zone reduced streamflow by 3.1%, total suspended solids by 18.1%, total nitrogen by 7.6%, and total phosphorus by 6.2%, while achieving biomass yields of 8.6–9.2 metric tons per hectare. These findings highlight switchgrass as a dual-benefit strategy for land reclamation and bioenergy feedstock production. 4) Integrated Co-Firing and CCS for Carbon-Negative Power Generation We modeled co-firing scenarios using the Power Plant Flexible Model (PPFM) to evaluate plant efficiency, greenhouse gas (GHG) emissions, and levelized cost of electricity (LCOE). Without carbon capture and storage (CCS), increasing switchgrass co-firing ratios reduced LCOE from $\$$150/MWh at 0% biomass to $\$$110/MWh at full substitution. Under CCS, costs remained higher (~$\$$250/MWh at 0% biomass) but decreased to $\$$200/MWh at 100% biomass, while enabling net-zero or carbon-negative electricity due to switchgrass sequestration benefits. Although CCS introduced efficiency penalties, pairing it with biomass co-firing offset these impacts and maximized climate benefits. Overall, optimizing co-firing ratios between 60-100%, supported by reliable logistics and storage strategies, emerged as a practical pathway to balance affordability, sustainability, and net-zero or negative GHG emissions while promoting productive reuse of AMLs.

01 COAL, LIGNITE, AND PEAT

Sim-to-real supervised domain adaptation for radioisotope identification

Machine learning has the potential to improve the speed and reliability of radioisotope identification using gamma spectroscopy. However, meticulously labeling an experimental dataset for training is often prohibitively expensive, while training models purely on synthetic data is risky due to the domain gap between simulated and experimental measurements. In this research, we demonstrate that supervised domain adaptation can substantially improve the performance of radioisotope identification models by transferring knowledge between synthetic and experimental data domains. We consider two domain adaptation scenarios: (1) a simulation-to-simulation adaptation, where we perform multi-label proportion estimation using simulated high-purity germanium detectors, and (2) a simulation-to-experimental adaptation, where we perform multi-class, single-label classification using measured spectra from handheld lanthanum bromide (LaBr) and sodium iodide (NaI) detectors. We begin by pretraining a spectral classifier on synthetic data using a custom transformer-based neural network. After subsequent fine-tuning on just 64 labeled experimental spectra, we achieve a test accuracy of 96% in the sim-to-real scenario with a LaBr detector, far surpassing a synthetic-only baseline model (75%) and a model trained from scratch (80%) on the same 64 spectra. Furthermore, we demonstrate that domain-adapted models learn more human-interpretable features than experiment-only baseline models. Overall, our results highlight the potential for supervised domain adaptation techniques to bridge the sim-to-real gap in radioisotope identification, enabling the development of accurate and explainable classifiers even in real-world scenarios where access to experimental data is limited.

Lalor, Peter W.

Incorporating Physical Priors into Weakly Supervised Anomaly Detection

We propose a new machine-learning-based anomaly detection strategy for comparing data with a background-only reference (a form of weak supervision). The sensitivity of previous strategies degrades significantly when the signal is too rare or there are many unhelpful features. Our prior-assisted weak supervision (PAWS) method incorporates information from a class of signal models to significantly enhance the search sensitivity of weakly supervised approaches. As long as the true signal is in the prespecified class, PAWS matches the sensitivity of a dedicated, fully supervised method without specifying the exact parameters ahead of time. On the benchmark LHC Olympics anomaly detection dataset, our mix of semisupervised and weakly supervised learning is able to extend the sensitivity over previous methods by a factor of 10 in cross section. Furthermore, if we add irrelevant (noise) dimensions to the inputs, classical methods degrade by another factor of 10 in cross section while PAWS remains insensitive to noise. This new approach could be applied in a number of scenarios and pushes the frontier of sensitivity between completely model-agnostic approaches and fully model-specific searches.

artificial neural networks

Self-Supervised and Interpretable Anomaly Detection Using Network Transformers

Machine learning and deep neural networks (DNNs) have been proposed as a tool to identify anomalies in computer network communications. However, due the obfuscated nature of off-the-shelf machine learning models, their output often does not provide enough information to isolate the source of the anomaly to take corrective measures. In this article, we introduce the network transformer (NeT), a DNN model for anomaly detection that incorporates the graph structure of the communication network in order to improve interpretability. Further, the presented approach has the following advantages: first, enhanced interpretability by incorporating the graph structure of computer networks; second, provides a hierarchical set of features that enables analysis at different levels of granularity; second, self-supervised training that does not require labeled data. The NeT model was evaluated on a set of anomalous scenarios executed in a real industrial control system. The presented approach successfully identified the anomalies, the devices affected, and the specific connections causing the anomalies, providing a data-driven hierarchical approach to analyze the behavior of a cyber network.

97 MATHEMATICS AND COMPUTING