Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Hyperspectral Sounder Spectral Fingerprinting: Using Machine Learning Techniques to Enhance Model-Based Physical Inversion

Different retrieval algorithms have been developed to process top-of-atmosphere (TOA) spectral radiance data provided by hyperspectral infrared sounder missions. Those algorithms are either optimal estimation method (OEM) based schemes with radiative transfer calculation involved in the retrieval process, or machine learning based methods that allow ultra-efficient data procession but lack of radiometric consistency validation based on the directly measured information. Combining both approaches leverages their respective technical advantages, leading to more accurate results. This study introduces a hyperspectral sounder fingerprinting algorithm to explore this hybrid approach. This approach involves the use of a spectral information-based classification method to identify an reference geophysical state and the corresponding radiative kernel. This enables the efficient retrieval of geophysical variables of interest through a radiative kernel-based linear inversion procedure. The fingerprinting method has been applied to analyze a decade-long hyperspectral sounder data record.

Wan Wu↗

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Generative models for simulation of KamLAND-Zen

Abstract The next generation of searches for neutrinoless double beta decay ($$0 \nu \beta \beta $$ 0 ν β β ) are poised to answer deep questions on the nature of neutrinos and the source of the Universe’s matter–antimatter asymmetry. They will be looking for event rates of less than one event per ton of instrumented isotope per year. To claim discovery, accurate and efficient simulations of detector events that mimic$$0 \nu \beta \beta $$ 0 ν β β is critical. Traditional Monte Carlo (MC) simulations can be supplemented by machine-learning-based generative models. This work describes the performance of generative models that we designed for monolithic liquid scintillator detectors like KamLAND to produce accurate simulation data without a predefined physics model. We present their current ability to recover low-level features and perform interpolation. In the future, the results of these generative models can be used to improve event classification and background rejection by providing high-quality abundant generated data.

Physics↗

GRinding Automated Classification Engine

This work is an ML-driven framework for automated surface analysis of microscopy images. We create a training dataset by imaging stainless steel samples to benchmark four developed deep neural network architectures. These models, based on a YOLOv8n-cls backend, integrate image features and process metadata using various fusion methods to distinguish between acceptable and unacceptable surface finishes. This code is associated with publication "Classifying Alloy Surface Preparation Quality with Metadata-Infused Machine Learning for Rapid Alloy Discovery" for project APEX LDRD-ER (25-ERD-039)

Gongora, AldairE [Lawrence Livermore National Labo↗

Machine learning for a Toolkit for Image Mining

A prototype user environment is described that enables a user with very limited computer skills to collaborate with a computer algorithm to develop search tools (agents) that can be used for image analysis, creating metadata for tagging images, searching for images in an image database on the basis of image content, or as a component of computer vision algorithms. Agents are learned in an ongoing, two-way dialogue between the user and the algorithm. The user points to mistakes made in classification. The algorithm, in response, attempts to discover which image attributes are discriminating between objects of interest and clutter. It then builds a candidate agent and applies it to an input image, producing an 'interest' image highlighting features that are consistent with the set of objects and clutter indicated by the user. The dialogue repeats until the user is satisfied. The prototype environment, called the Toolkit for Image Mining (TIM) is currently capable of learning spectral and textural patterns. Learning exhibits rapid convergence to reasonable levels of performance and, when thoroughly trained, Fo appears to be competitive in discrimination accuracy with other classification techniques.

Delanoy, Richard L.↗

NETL RDE Image Classification Dataset 2020 - 10 Classes

Dataset including high-speed down-axis RDE images used for updated image classification study. This dataset includes 100,000 images with 10 classifications: 1CW, 1CCW, 2CW, 2CCW, 3CW, 3CCW, and Deflagration. Images are cropped to center annulus, and resized to 301x301 pixels. Images are filtered using the AFRL Beta correction factor.

AS↗

Ongoing Work: A Prototype Dataset for Low-flying Autonomous Medical UAS Operations

This paper presents ongoing work to create a dataset for low-flying autonomous medical UAS operations, focused on human stance recognition. This is an exploration of the viability of airborne classification for the Drone as a First Responder (DFR) concept in which a UAS arrives at the scene of an incident before emergency response personnel can get there and provides some level of situational awareness for the personnel arriving to the scene. Future incarnations could also see the UAS administer some level of care to injured parties at the scene. The data set, focused on detecting human stance, being developed here is the result of 30 test flights at NASA Langley Research Center in early 2024. In addition to flights where the participant (an anthropomorphic testing device or human) is alone in the viewing area holding a particular stance, two emergency scenes have been fabricated and collected through video - ``bike crash'' and ``difficult camping''. These test flights include four human participants. The contribution of this work upon completion will be a publicly available data set for the development of classification engines focused on human stance, and in the future, even triage.

Uncrewed Aerial Systems↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Cloud-Computing and Machine Learning in Support of Country-Level Land Cover and Ecosystem Extent Mapping in Liberia and Gabon

Liberia and Gabon joined the Gaborone Declaration for Sustainability in Africa (GDSA), established in 2012, with the goal of incorporating the value of nature intonational decision making by estimating the multiple services obtained from ecosystems using the natural capital accounting framework. In this study, we produced 30-m resolution 10 classes land cover maps for the 2015 epoch for Liberia and Gabon using the Google Earth Engine (GEE) cloud platform to support the ongoing natural capital accounting efforts in these nations. We pro-pose an integrated method of pixel-based classification using Landsat 8 data, the Random Forest(RF) classifier and ancillary data to produce high quality land cover products to fit abroad range of applications, including natural capital accounting. Our approach focuses on a pre-classification filtering (Masking Phase) based on spectral signature and ancillary data to reduce the number of pixels prone to be misclassified; therefore, increasing the quality of the final product. The proposed approach yields an overall accuracy of 83% and 81% for Liberia and Gabon, respectively, out performing prior land cover products for these countries in both thematic content and accuracy. Our approach, while relatively simple and highly replicable, was able to produce high quality land cover products to fill an observational gap in up to date land cover data at national scale for Liberia and Gabon.

Celio de Sousa↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

NANO.PTML model for read-across prediction of nanosystems in neurosciences. computational model and experimental case of study

Abstract Neurodegenerative diseases involve progressive neuronal death. Traditional treatments often struggle due to solubility, bioavailability, and crossing the Blood-Brain Barrier (BBB). Nanoparticles (NPs) in biomedical field are garnering growing attention as neurodegenerative disease drugs (NDDs) carrier to the central nervous system. Here, we introduced computational and experimental analysis. In the computational study, a specific IFPTML technique was used, which combined Information Fusion (IF) + Perturbation Theory (PT) + Machine Learning (ML) to select the most promising Nanoparticle Neuronal Disease Drug Delivery (N2D3) systems. For the application of IFPTML model in the nanoscience, NANO.PTML is used. IF-process was carried out between 4403 NDDs assays and 260 cytotoxicity NP assays conducting a dataset of 500,000 cases. The optimal IFPTML was the Decision Tree (DT) algorithm which shown satisfactory performance with specificity values of 96.4% and 96.2%, and sensitivity values of 79.3% and 75.7% in the training (375k/75%) and validation (125k/25%) set. Moreover, the DT model obtained Area Under Receiver Operating Characteristic (AUROC) scores of 0.97 and 0.96 in the training and validation series, highlighting its effectiveness in classification tasks. In the experimental part, two samples of NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) were synthesized by thermal decomposition of an iron(III) oleate (FeOl) precursor and structurally characterized by different methods. Additionally, in order to make the as-synthesized hydrophobic NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) soluble in water the amphiphilic CTAB (Cetyl Trimethyl Ammonium Bromide) molecule was employed. Therefore, to conduct a study with a wider range of NP system variants, an experimental illustrative simulation experiment was performed using the IFPTML-DT model. For this, a set of 500,000 prediction dataset was created. The outcome of this experiment highlighted certain NANO.PTML systems as promising candidates for further investigation. The NANO.PTML approach holds potential to accelerate experimental investigations and offer initial insights into various NP and NDDs compounds, serving as an efficient alternative to time-consuming trial-and-error procedures.

60 APPLIED LIFE SCIENCES↗

Differentiable vertex fitting for jet flavor tagging

This work explores the use of differentiable programming to integrate domain knowledge, in the form of domain specific software, into neural networks to develop scientific machine learning systems. We propose a differentiable vertex fitting algorithm that estimates the crossing point of multiple curves. In the high energy physics setting, these curves are defined by particle equations of motion and the crossing point represents the origin of particle production. This differentiable vertex fitting algorithm can be seamlessly integrated into neural networks, and we show its utility and efficacy in the high energy physics application of the classification of jets, i.e., collimated streams of particles in particle detectors whose originating parent particle we aim to classify. We demonstrate how differentiable vertex fitting can be integrated into larger transformer-based models for jet flavor tagging and show improvements in heavy flavor jet classification when compared to baseline models. Published by the American Physical Society 2024

Smith, Rachel E. C. (ORCID:0000000335851262)↗

Harmonized Sentinel-1 SAR Global River Geometry and Inundation Database

Satellite-based observations on river geometries are sporadic in time, space, or both. Most satellite-based surface water maps, river widths, water surface elevations (WSE), slopes, and bathymetry are asynchronized in time and space. The current configuration of satellites such as Sentinel-6 measured the WSE but is missing the river width, slopes, and depths. To advance hydrological sciences research, there is a need to produce a harmonized time series of river geometry data of non-SWOT satellites in partnership with the upcoming SWOT mission. The SWOT satellite will measure river width, height, and slope but missing river depth measurements in space and time. Further, none of these current satellites measure the WSE, river width, and slopes synchronously. In this work, we use the Sentinel-1 SAR satellite data archive from 2015 to the present to create a global river width and surface water database at the reach scale. A modified version of the Sentinel SAR surface water classification algorithm from ASF is used to quantify the surface water extent on the stream approximately every six days (at the equator) at 10m spatial resolution globally. This 10m water mask is fed into a workflow to quantify the river widths, surface water inundations, slopes, and synthetic bathymetry in SWORD (SWOT River Database) stream networks. A Satellite HAND is used to address the cloud obscured surface water observations using a trained machine learning algorithm. We use WSE derived from the Global Water Monitor from NASA GSFC, Hydroweb from LEGOS, and ICESat-2 to harmonize the WSE observation. And Landsat-8/9 and Sentinel-2 water observations to fill the gaps in the Sentinel-1 SAR database. We use Congo River Basin as a test case where we have more than 500 radar altimetry-based WSE, continuous series of Sentinel-1, ICESat-2, Landsat-8/9, and Sentinel-2 observations. A Congo River hydrologic model is used to generate the streamflow discharge. The satellite observed river reaches are assimilated with the stream flows computed by the routing models. And the downstream reaches in the river network without satellite observations get optimized for discharge/river geometry at each observation cycle. Our final product is a harmonized river geometry dataset (reach's water extent, WSE, slope, synthetic bathymetry) for Congo Basin's SWORD reaches.

Chandana Gangodagamage↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Universal Fourier Attack for Time Series

A wide variety of adversarial attacks have been proposed and explored using image and audio data. These attacks are notoriously easy to generate digitally when the attacker can directly manipulate the input to a model, but are much more difficult to implement in the real world. In this paper we present a universal, time invariant attack for general time series data such that the attack has a frequency spectrum primarily composed of the frequencies present in the original data. The universality of the attack makes it fast and easy to implement as no computation is required to add it to an input, while time invariance is useful for real world deployment. Additionally, the frequency constraint ensures the attack can withstand filtering defenses. We demonstrate the effectiveness of the attack on two different classification tasks through both digital and real world experiments, and show that the attack is robust against common transform-and-compare defense pipelines.

97 MATHEMATICS AND COMPUTING↗

Learning to classify quantum phases of matter with a few measurements

We study the identification of quantum phases of matter, at zero temperature, when only part of the phase diagram is known in advance. Following a supervised learning approach, we show how to use our previous knowledge to construct an observable capable of classifying the phase even in the unknown region. By using a combination of classical and quantum techniques, such as tensor networks, kernel methods, generalization bounds, quantum algorithms, and shadow estimators, we show that, in some cases, the certification of new ground states can be obtained with a polynomial number of measurements. An important application of our findings is the classification of the phases of matter obtained in quantum simulators, e.g. cold atom experiments, capable of efficiently preparing ground states of complex many-particle systems and applying simple measurements, e.g. single qubit measurements, but unable to perform a universal set of gates.

quantum machine learning↗

Next-Generation Optical Sensing Technologies for Exploring Ocean Worlds - NASA FluidCam, MiDAR, and NeMO-Net

We highlight three emerging NASA optical technologies that enhance our ability to remotely sense, analyze, and explore ocean worlds–FluidCam and fluid lensing, MiDAR, and NeMO-Net. Fluid lensing is the first remote sensing technology capable of imaging through ocean waves without distortions in 3D at sub-cm resolutions. Fluid lensing and the purpose-built FluidCam CubeSat instruments have been used to provide refraction-corrected 3D multispectral imagery of shallow marine systems from unmanned aerial vehicles (UAVs). Results from repeat 2013 and 2016 airborne fluid lensing campaigns over coral reefs in American Samoa present a promising new tool for monitoring fine-scale ecological dynamics in shallow aquatic systems tens of square kilometers in area. MiDAR is a recently-patented active multispectral remote sensing and optical communications instrument which evolved from FluidCam. MiDAR is being tested on UAVs and autonomous underwater vehicles (AUVs) to remotely sense living and non-living structures in light-limited and analog planetary science environments. MiDAR illuminates targets with high-intensity narrowband structured optical radiation to measure an object’s spectral reflectance while simultaneously transmitting data. MiDAR is capable of remotely sensing reflectance at fine spatial and temporal scales, with a signal-to-noise ratio 10-10(exp 3) times higher than passive airborne and spaceborne remote sensing systems, enabling high-framerate multispectral sensing across the ultraviolet, visible, and near-infrared spectrum. Preliminary results from a 2018 mission to Guam show encouraging applications of MiDAR to imaging coral from airborne and underwater platforms whilst transmitting data across the air-water interface. Finally, we share NeMO-Net, the Neural Multi-Modal Observation & Training Network for Global Coral Reef Assessment. NeMO-Net is a machine learning technology under development that exploits high-resolution data from FluidCam and MiDAR for augmentation of low-resolution airborne and satellite remote sensing. NeMO-Net is intended to harmonize the growing diversity of 2D and 3D remote sensing with in situ data into a single open-source platform for assessing shallow marine ecosystems globally using active learning for citizen-science based training. Preliminary results from four-class Q17 coral classification have an accuracy of 94.4%. Together, these maturing technologies present promising scalable, practical, and cost-efficient innovations that address current observational and technological challenges in optical sensing of marine systems.

Ved Chirayath↗