Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Chemical signature characterization with hyperspectral imagery: novel deep learning model architectures and physically-motivated data augmentation techniques

The high spectral resolution afforded by Hyperspectral Imaging (HSI) sensors is poised to bring unprecedented advancements to signature characterization applications. Thus far, much of the research in the machine learning field devoted to HSI applications has focused on a few specific tasks like land-use land-cover classification. In land classification tasks, spatial information is very important, and model architectures are often designed to leverage spatial contexts. However, it is unclear how well these spatially-tuned models will translate to tasks where spectral information is critical, like the detection and characterization of chemicals. In this work, we compare spectral models (inputs are 1D spectra) and spatial-spectral models (inputs are 3D cubes) in the context of predicting chemical concentration maps. We find that spatial-spectral models perform the best, though we find a wide range in performance across the different architectures tested. Additionally, we find that model performance is impacted by the availability of training data, particularly in scenarios where the training data doesn't fully capture the true variance of real-world conditions. We find that data augmentation can help mitigate sparse coverage of observed parameter space (e.g., seasonal or geographic variability in ground cover), and present augmentation strategies that are tailored to hyperspectral data.

• Artificial intelligence (AI) / machine learning ↗

Fast Solution in Sparse LDA for Binary Classification

An algorithm that performs sparse linear discriminant analysis (Sparse-LDA) finds near-optimal solutions in far less time than the prior art when specialized to binary classification (of 2 classes). Sparse-LDA is a type of feature- or variable- selection problem with numerous applications in statistics, machine learning, computer vision, computational finance, operations research, and bio-informatics. Because of its combinatorial nature, feature- or variable-selection problems are NP-hard or computationally intractable in cases involving more than 30 variables or features. Therefore, one typically seeks approximate solutions by means of greedy search algorithms. The prior Sparse-LDA algorithm was a greedy algorithm that considered the best variable or feature to add/ delete to/ from its subsets in order to maximally discriminate between multiple classes of data. The present algorithm is designed for the special but prevalent case of 2-class or binary classification (e.g. 1 vs. 0, functioning vs. malfunctioning, or change versus no change). The present algorithm provides near-optimal solutions on large real-world datasets having hundreds or even thousands of variables or features (e.g. selecting the fewest wavelength bands in a hyperspectral sensor to do terrain classification) and does so in typical computation times of minutes as compared to days or weeks as taken by the prior art. Sparse LDA requires solving generalized eigenvalue problems for a large number of variable subsets (represented by the submatrices of the input within-class and between-class covariance matrices). In the general (fullrank) case, the amount of computation scales at least cubically with the number of variables and thus the size of the problems that can be solved is limited accordingly. However, in binary classification, the principal eigenvalues can be found using a special analytic formula, without resorting to costly iterative techniques. The present algorithm exploits this analytic form along with the inherent sequential nature of greedy search itself. Together this enables the use of highly-efficient partitioned-matrix-inverse techniques that result in large speedups of computation in both the forward-selection and backward-elimination stages of greedy algorithms in general.

Moghaddam, Baback↗

Safe and Robust Binary Classification and Fault Detection Using Reinforcement Learning

In this paper, we propose a learning-based method utilizing the Soft Actor-Critic (SAC) algorithm to train a binary Support Vector Machine (SVM) classifier. This classifier is designed to identify valid input spaces in high-dimensional, highly constrained systems while minimizing the total runtime of offline simulations. The simulations adapt their runtime based on the likelihood that a given training input will be informative to the classifier. Furthermore, we introduce a method for using the trained SAC model to predict whether a desired system input is likely to violate constraints, along with a technique to adjust the input as necessary. Additionally, we explore the potential of this model to detect faults or adversarial attacks within the system. The effectiveness of our approach is demonstrated through various simulations of challenging classification problems and a constrained quadrotor model.

Netter, Josh [Georgia Institute of Technology, Atl↗

Few measurement shots challenge generalization in learning to classify entanglement

The ability to extract general laws from a few known examples depends on the complexity of the problem and on the amount of training data. In the quantum setting, the learner's generalization performance is further challenged by the destructive nature of quantum measurements that, together with the no-cloning theorem, limits the amount of information that can be extracted from each training sample. In this paper we focus on hybrid quantum learning techniques where classical machine-learning methods are paired with quantum algorithms and show that, in some settings, the uncertainty coming from a few measurement shots can be the dominant source of errors. We identify an instance of this possibly general issue by focusing on the classification of maximally entangled vs. separable states, showing that this toy problem becomes challenging for learners unaware of entanglement theory. Finally, we introduce an estimator based on classical shadows that performs better in the big data, few copy regime. Our results show that the naive application of classical machine-learning methods to the quantum setting is problematic, and that a better theoretical foundation of quantum learning is required.

97 MATHEMATICS AND COMPUTING↗

Assessment of fine-tuned large language models for real-world chemistry and material science applications

The current generation of large language models (LLMs) has limited chemical knowledge. Recently, it has been shown that these LLMs can learn and predict chemical properties through fine-tuning. Using natural language to train machine learning models opens doors to a wider chemical audience, as field-specific featurization techniques can be omitted. In this work, we explore the potential and limitations of this approach. We studied the performance of fine-tuning three open-source LLMs (GPT-J-6B, Llama-3.1-8B, and Mistral-7B) for a range of different chemical questions. We benchmark their performances against “traditional” machine learning models and find that, in most cases, the fine-tuning approach is superior for a simple classification problem. Depending on the size of the dataset and the type of questions, we also successfully address more sophisticated problems. The most important conclusions of this work are that, for all datasets considered, their conversion into an LLM fine-tuning training set is straightforward and that fine-tuning with even relatively small datasets leads to predictive models. These results suggest that the systematic use of LLMs to guide experiments and simulations will be a powerful technique in any research study, significantly reducing unnecessary experiments or computations.

Van Herck, Joren↗

Citizen science for IceCube: Name that Neutrino

Name that Neutrino is a citizen science project where volunteers aid in classification of events for the IceCube Neutrino Observatory, an immense particle detector at the geographic South Pole. From March 2023 to September 2023, volunteers did classifications of videos produced from simulated data of both neutrino signal and background interactions. Name that Neutrino obtained more than 128,000 classifications by over 1800 registered volunteers that were compared to results obtained by a deep neural network machine-learning algorithm. Possible improvements for both Name that Neutrino and the deep neural network are discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Automated Analysis of a Large-Scale Sky Survey: The SKICAT System

We describe the application of decision tree based classification techniques to the development of an automated tool for the reduction of a large scientific data set. The primary benefits of the SKICAT approach are increased data reduction throughput, repeatability, and consistency of classification.

data analysis image databases↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

A 3D Citizen Science Video Game for NeMO-Net, the NASA Neural Multi-Modal Observation and Training Network for Global Coral Reef Assessment

NeMO-Net, the NASA neural multi-modal observation and training network for global coral reef assessment, is an open-source deep convolutional neural network aimed at accurately assessing the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. We present here the active learning component of the project, which consists of an interactive video game prototype for tablet and mobile devices where players are able to intuitively label morphology classifications over mm-scale 3D coral reef imagery. Active learning applications present a novel methodology for engaging the public while efficiently providing large-scale training and test data for increasingly complex and data-intensive machine learning algorithms. NeMO-Net trains players on domain-specific knowledge through interactive tutorials and periodically checks players' input against pre-classified coral imagery to gauge their accuracy and utilize in-game mechanics to provide personalized classification training. Players can rate the classifications of other players, unlock rewards and join a global community as they explore and classify coral reefs and other shallow marine environments.

Citizen Science↗

Application of Data Cubes for Improving Detection of Water Cycle Extreme Events

As part of an ongoing NASA-funded project to remove a longstanding barrier to accessing NASA data (i.e., accessing archived time-step array data as point-time series), for the hydrology and other point-time series-oriented communities, "data cubes" are created from which time series files (aka "data rods") are generated on-the-fly and made available as Web services from the Goddard Earth Sciences Data and Information Services Center (GES DISC). Data cubes are data as archived rearranged into spatio-temporal matrices, which allow for easy access to the data, both spatially and temporally. A data cube is a specific case of the general optimal strategy of reorganizing data to match the desired means of access. The gain from such reorganization is greater the larger the data set. As a use case of our project, we are leveraging existing software to explore the application of the data cubes concept to machine learning, for the purpose of detecting water cycle extreme events, a specific case of anomaly detection, requiring time series data. We investigate the use of support vector machines (SVM) for anomaly classification. We show an example of detection of water cycle extreme events, using data from the Tropical Rainfall Measuring Mission (TRMM).

water cycle extreme events↗

Invasion in the Niger Delta: Remote Sensing of Mangrove Conversion to Invasive Nypa fruticans from 2015-2020

Invasive species are a leading threat to biodiversity worldwide. Nypa palm ( Nypa fruticans ) has emerged as the predominant invasive species in the Niger Delta region of Nigeria. While endemic mangroves have high rates of carbon sequestration, stabilize coastlines, and protect biodiversity, Nypa does not provide these services outside its native region of Southeast Asia. Oil exploration and urbanization in this region also exacerbates mangrove loss and Nypa spread. As Nypa is difficult to distinguish from endemic mangrove species in remotely sensed data, estimates of mangrove and ecosystem services losses in Nigeria are highly uncertain. Here, we analyze multisensor satellite data with machine learning to quantify the rapid expansion of Nypa from 2015-2020 in Nigeria. Using Landsat imagery and random forest classification, we quantify total potential Nypa extent in Nigeria in 2019. We then produced a Nypa extent map using iterative combinations of Sentinel-1 SAR, Sentinel-2 MSI, and ALOS PALSAR. Random forest classifications using SAR data from ALOS and Sentinel-1 were best suited for mapping Nypa extent with similar accuracies (78% and 75% respectively). Based on data availability and accuracy, we focused our change analysis on Sentinel-1 SAR. Our results show ~28,000 ha of mangroves were converted to Nypa in Nigeria by 2020 and covered a larger extent than endemic mangroves, compounding the effect of the existing degradation and deforestation in the region. We also compared forest height and complexity estimates from GEDI (Global Ecosystem Dynamics Investigation) LiDAR to further distinguish between endemic mangroves and Nypa in three dimensions. Nypa structural variability, measured by top-of-canopy height, vegetation cover, plant area index, and foliage height diversity, was lower than that of mangroves. At current rates of Nypa expansion, the entire area of study would be invaded by Nypa by 2028, with potentially detrimental consequences to the ecosystem services provided by mangroves.

GEE↗

Persistent Classification: Understanding Adversarial Attacks by Studying Decision Boundary Dynamics

ABSTRACT There are a number of hypotheses underlying the existence of adversarial examples for classification problems. These include the high‐dimensionality of the data, the high codimension in the ambient space of the data manifolds of interest, and that the structure of machine learning models may encourage classifiers to develop decision boundaries close to data points. This article proposes a new framework for studying adversarial examples that does not depend directly on the distance to the decision boundary. Similarly to the smoothed classifier literature, we define a (natural or adversarial) data point to be ( γ , σ)‐stable if the probability of the same classification is at least for points sampled in a Gaussian neighborhood of the point with a given standard deviation . We focus on studying the differences between persistence metrics along interpolants of natural and adversarial points. We show that adversarial examples have significantly lower persistence than natural examples for large neural networks in the context of the MNIST and ImageNet datasets. We connect this lack of persistence with decision boundary geometry by measuring angles of interpolants with respect to decision boundaries. Finally, we connect this approach with robustness by developing a manifold alignment gradient metric and demonstrating the increase in robustness that can be achieved when training with the addition of this metric.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accelerating Thermochemical Equilibrium Calculations for Nuclear Reactor Applications

Thermochemical properties play a key role in modeling and simulation of several key phenomena in nuclear reactors. There has been an increasing interest in incorporating CALPHAD-based formulations in multiphysics simulations including for Molten Salt Reactors where knowledge of phase evolution of the salt and the chemical potentials of various elements are of utmost importance in source term analyses and redox control. However, the size of such simulations is often limited by the high computational cost of full thermodynamic equilibrium calculations. This work discusses the current efforts aimed at accelerating thermochemical equilibrium calculations for multiphysics simulations performed using the open-source finite element / finite volume code Multiphysics Object Oriented Simulation Environment (MOOSE) [1]. While several methods have been proposed for accelerating phase equilibrium calculations [2], most focus on relatively small systems and often rely on a- priori knowledge of the state-space of the system. Nuclear materials, however, are often multi-component systems owing to the evolution of composition under irradiation and an approach based on a-priori mapping of phase diagram is often not enough. This work is aimed at demonstrating an on-the-fly surrogate modeling framework that uses active learning to reduce the number of full equilibrium calculations that must be performed. By combining with efficient coupling approaches, the surrogate framework helps in reducing the computational cost of thermodynamic equilibrium informed multiphysics simulations of nuclear materials. The performance is benchmarked against full coupling with the thermochemistry library Thermochimica [3]. This work uses a machine learning based approach for constructing surrogate models to predict the stable phases in a multicomponent system. The surrogates were constructed using neural networks and Gaussian process classification. In this work, we compare the relative performance of the two methods. We also demonstrate the use of caching previous calculations by interpolating the values from nearest neighbors. References [1] Lindsay, A.D., et al. "2.0 – MOOSE: Enabling massively parallel multiphysics simulation", SoftwareX, 20 (2022): 101202. [2] Roos, W.A. and Zietsman J.H. "Accelerating complex chemical equilibrium calculations – A Review", Calphad, 77 (2022): 102380. [3] Piro, M.H.A., et al. "The thermochemistry library Thermochimica", Computational Materials Science, 67 (2013): 266-272.

36 MATERIALS SCIENCE↗

SIGHT: Stacked Integration of Geospatial Hierarchical Typologies for Inferring Building Characteristics

Building characteristics are often absent in building stock datasets, particularly in regions most vulnerable to climate change and requiring effective disaster management strategies. Traditional machine learning approaches, while widely used to predict building attributes, typically neglect the spatial context of the data, leading to less accurate and reliable outcomes. To address these challenges, this paper introduces a novel algorithm, the Stacked Integration of Geospatial Hierarchical Typologies. This algorithm adapts a meta-learning framework to incorporate geospatial context into the predictive modeling process. We demonstrate the utility of the algorithm through two primary use cases: building use type classification and building height prediction. The algorithm consistently achieved or exceeded a 0.94 macro average F1 score across five geographically distinct countries for building use type classification. For building height prediction, it accurately predicted heights with a root mean square error of 3.01 in a comprehensive study using roughly 3.6 million buildings in Japan. These results underscore the benefits of integrating spatial hierarchies into machine learning models, enhancing both predictive accuracy and reliability in geospatial modeling. This work introduces a new algorithm to address the pervasive data sparsity issue in existing building stock datasets.

Adams, Daniel [ORNL] (ORCID:0000000196950577)↗

Graph theory inspired anomaly detection at the LHC

Designing model-independent anomaly detection algorithms for analyzing LHC data remains a central challenge in the search for new physics, due to the high dimensionality of collider events. In this work, we develop a graph autoencoder as an unsupervised, model-agnostic tool for anomaly detection, using the LHC Olympics dataset as a benchmark. By representing jet constituents as a graph, we introduce a method to systematically control the information available to the model through sparse graph constructions that serve as physically motivated inductive biases. Specifically, (1) we construct graph autoencoders based on locally rigid Laman graphs and globally rigid unique graphs, and (2) we explore the clustering of jet constituents into subjets to interpolate between high- and low-level input representations. We obtain the best performance, measured in terms of the Significance Improvement Characteristic curve for an intermediate level of subjet clustering and certain sparse unique graph constructions. We further investigate the role of graph connectivity in jet classification tasks. Our results demonstrate the potential of leveraging graph-theoretic insights to refine and increase the interpretability of machine learning tools for collider experiments.

Automation↗

Decision-tree structures utilizing a phase-transition material

The rich internal physics due to competing electronic phases present in phase-transition materials such as VO2 offer the potential for compact building block design for emerging non-von Neumann computing technologies. Here, based on the relaxation dynamics of an insulator-metal phase transition, we demonstrate experimentally a decision-tree classifier embedded within a single volatile resistive switching device. The tree is constructed by the combination of the voltage pulse and relaxation time and can adapt to different tasks. We use machine learning to analyze the relaxation process, enabling a predictive voltage-relaxation time phase diagram for the electrical resistance state. Classification of the etiology of the chronic cough is presented as a proof-of-principle use case. Further, our approach can be generalized to broader classes of solid-state and solid-liquid interfacial systems that demonstrate a variety of phase relaxations.

36 MATERIALS SCIENCE↗

ICAT: The Interactive Corpus Analysis Tool

The Interactive Corpus Analysis Tool (ICAT) is a Python library for creating dashboards to explore textual datasets and build simple binary classification models to help filter through them and focus on entries of interest. This tool uses a form of interactive machine learning (IML), a paradigm of “machine teaching” (Simard et al., 2017) that sits at the intersection of the fields of human computer interaction (HCI), visual analytics, and machine learning. The intent of ICAT is to allow subject matter experts (SME) with limited to no experience in machine learning to benefit from an iterative human-in-the-loop (HITL) approach to building their own model without needing to understand the details of the underlying algorithm. This interactivity is achieved by allowing the user to create features, label data points, and visually manipulate a representation of the features to manually cluster and investigate data, while a model is trained on the fly based on these actions. ICAT is built on top of the Panel (Holoviz, 2018) library, using a combination of Vega, a custom IPyWidget using D3, and ipyvuetify, and is intended to be used inside of a Jupyter environment.

Martindale, Nathan [Oak Ridge National Laboratory ↗

The Radiation Biology Ontology: A New Tool Supporting FAIR Principles Across Radiation Biology Facilitating Data Discovery and Integration

Development of the Radiation Biology Ontology (RBO) was motivated by the need for a comprehensive, well-structured ontology for encoding radiation biology metadata. The primary use-cases were archiving data in the STORE database (https://www.storedb.org/), the repository for the RadoNorm Project, and in GeneLab (https://genelab.nasa.gov), NASA’s ‘omics database. The scope of radiobiology research ranges from physics to radiation oncology to socio-legal studies; no existing ontology has the necessary breadth or depth. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR radiation biology data.

ontology↗