Search NASASearch

SEARCH · Search NASA

Results for “machine learning classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Classifying Unidentified X-Ray Sources in the Chandra Source Catalog Using A Multiwavelength Machine-Learning Approach

The rapid increase in serendipitous X-ray source detections requires the development of novel approaches to efficiently explore the nature of X-ray sources. If even a fraction of these sources could be reliably classified, it would enable population studies for various astrophysical source types on a much larger scale than currently possible. Classification of large numbers of sources from multiple classes characterized by multiple properties (features) must be done automatically and supervised machine learning (ML) seems to provide the only feasible approach. We perform classification of Chandra Source Catalog version 2.0 (CSCv2) sources to explore the potential of the ML approach and identify various biases, limitations, and bottlenecks that present themselves in these kinds of studies. We establish the framework and present a flexible and expandable Python pipeline, which can be used and improved by others. We also release the training data set of 2941 X-ray sources with confidently established classes. In addition to providing probabilistic classifications of 66,369 CSCv2 sources (21% of the entire CSCv2 catalog), we perform several narrower-focused case studies (high-mass X-ray binary candidates and X-ray sources within the extent of the H.E.S.S. TeV sources) to demonstrate some possible applications of our ML approach. We also discuss future possible modifications of the presented pipeline, which are expected to lead to substantial improvements in classification confidences.

Hui Yang

Neural network-based classification and regression of magnetohydrodynamic modes in tokamaks

We present a machine learning-based magnetohydrodynamic (MHD) classifier and regressor that utilizes real or complex-valued 3D magnetic sensor array data to determine neoclassical tearing mode (NTM) onset times in tokamaks with millisecond accuracy. The input dataset consists of poloidal profiles of complex Fourier amplitudes with an n = 1 toroidal mode number from 144 human-labeled ITER Baseline Scenario discharges in the DIII-D tokamak, spanning both tearing-dominated and sawtooth-dominated regimes. Since m, n = 2,1 NTMs frequently emerge alongside sawteeth at the same frequency in this scenario, the focus is on isolating the m = 1 and m = 2 components of the n = 1 MHD mode near the tearing onset. To improve model regularization and prediction stability, singular value decomposition was applied to balance the sawtooth and tearing datasets. The enriched datasets facilitated training neural networks that learn the key distinguishing features of sawtooth and tearing modes in the poloidal profiles of their magnetic amplitude and phase. When the modes occur independently, the networks achieve perfect classification due to the modes’ distinct characteristics and low measurement noise. In the more experimentally relevant case where both modes coexist, the networks maintain exceptional performance across key metrics. Tests on synthetic data with known ground truth demonstrate the superior accuracy of the neural network trained on complex-valued input compared to models using real amplitude, phase, or pseudo-complex data, achieving both a mean time delay and standard deviation below 1 ms. Notably, standard linear regression methods fitting the dominant singular modes to the data closely match the neural network’s performance. Applying these methods across a broad range of H-mode scenarios will enable future studies to systematically identify dominant NTM triggers as scenario-specific variables, paving the way for more effective tearing mode avoidance strategies in future fusion reactor designs.

machine learning

Automating sky object classification in astronomical survey images

We describe the application of machine classification techniques to the development of an automated tool for the reduction of a large scientific data set. The 2nd Palomer Observatory Sky Survey is nearly completed. This survey provides comprehensive coverage of the northern celestial hemisphere in the form of photographic plates. The plates are being transformed into digitized images whose quality will probably not be surpassed in the next ten to twenty years. The images are expected to contain on the order of 10(exp 7) galaxies and 10(exp 8) stars. Astronomers wish to determine which of these sky objects belong to various classes of galaxies and stars. The size of this data set precludes manual analysis. Our approach is to develop a software system which integrates the functions of independently developed techniques for image processing and data classification. Digitized sky images are passed through image processing routines to identify sky objects and to extract a set of features for each object. These routines are used to help select a useful set of attributes for classifying sky objects. Then GID3* and O-BTree, two inductive learning techniques, learn classification decision trees from examples. These classifiers will be used to process the rest of the data. This paper gives an overview of the machine learning techniques used, describes the details of our specific application, and reports the initial encouraging results. The results indicate that our approach is well-suited to the problem. The primary benefits of the approach are increased data reduction throughput and consistency of classification. The classification rules which are the product of the inductive learning techniques will form an object, examinable basis for classifying sky objects. A final, not to be underestimated benefit is that astronomers will be freed from the tedium of an intensely visual task to pursue more challenging analysis and interpretation problems based on automatically cataloged data.

Fayyad, Usama M.

State Predictor of Classification Cognitive Engine Applied to Channel Fading

This study presents the application of machine learning (ML) to a space-to-ground communication link, showing how ML can be used to detect the presence of detrimental channel fading. Using this channel state information, the communication link can be used more efficiently by reducing the amount of lost data during fading. The motivation for this work is based on channel fading observed during on-orbit operations with NASA's Space Communication and Navigation (SCaN) testbed on the International Space Station (ISS). This paper presents the process to extract a target concept (fading and not-fading) from the raw data. The pre-processing and data exploration effort is explained in detail, with a list of assumptions made for parsing and labelling the dataset. The model selection process is explained, specifically emphasizing the benefits of using an ensemble of algorithms with majority voting for binary classification of the channel state. Experimental results are shown, highlighting how an end-to-end communication system can utilize knowledge of the channel fading status to identity fading and take appropriate action. With a laboratory testbed to emulate channel fading, the overall performance is compared to standard adaptive methods without fading knowledge, such as adaptive coding and modulation.

Fading

Support Vector Machines for Hyperspectral Remote Sensing Classification

The Support Vector Machine provides a new way to design classification algorithms which learn from examples (supervised learning) and generalize when applied to new data. We demonstrate its success on a difficult classification problem from hyperspectral remote sensing, where we obtain performances of 96%, and 87% correct for a 4 class problem, and a 16 class problem respectively. These results are somewhat better than other recent results on the same data. A key feature of this classifier is its ability to use high-dimensional data without the usual recourse to a feature selection step to reduce the dimensionality of the data. For this application, this is important, as hyperspectral data consists of several hundred contiguous spectral channels for each exemplar. We provide an introduction to this new approach, and demonstrate its application to classification of an agriculture scene.

Gualtieri, J. Anthony

Multi-Class Anomaly Detection in Flight Data using Semi-Supervised Explainable Deep Learning Model

Identifying precursor for safety incidents in aviation data is a crucial task, yet extremely challenging. The main approach, in practice, leverages domain expertise to define expected tolerances in system’s behavior and alarm exceedance from such safety margins. However, this approach is incapable of identifying unknown risk and vulnerabilities. Machine learning has been long studied and deployed to identify precursors for such anomalies, with the great challenge of the need for sufficient labelled set of data to achieve a reliable and accurate performance. In this article, we develop an explainable deep semi-supervised model for anomaly detection in aviation, building upon recent advancements in the machine learning literature. The proposed model combines feature engineering and classification in the feature space, while leveraging all available data (labelled and unlabeled). Validating on two case studies of anomaly detection in take-off and landing phases of commercial aircraft, we show that our model is able to outperform state-of-the-art supervised anomaly detection model and reach significantly high accuracy and low false alarm with minimum amount of available labelled data.

Anomaly Detection

Astronaut Photography of the Earth: A Long-Term Dataset for Earth Systems Research, Applications, and Education

The NASA Earth observations dataset obtained by humans in orbit using handheld film and digital cameras is freely accessible to the global community through the online searchable database at https://eol.jsc.nasa.gov, and offers a useful compliment to traditional ground-commanded sensor data. The dataset includes imagery from the NASA Mercury (1961) through present-day International Space Station (ISS) programs, and currently totals over 2.6 million individual frames. Geographic coverage of the dataset includes land and oceans areas between approximately 52 degrees North and South latitudes, but is spatially and temporally discontinuous. The photographic dataset includes some significant impediments for immediate research, applied, and educational use: commercial RGB films and camera systems with overlapping bandpasses; use of different focal length lenses, unconstrained look angles, and variable spacecraft altitudes; and no native geolocation information. Such factors led to this dataset being underutilized by the community but recent advances in automated and semi-automated image geolocation, image feature classification, and web-based services are adding new value to the astronaut-acquired imagery. A coupled ground software and on-orbit hardware system for the ISS is in development for planned deployment in mid-2017; this system will capture camera pose information for each astronaut photograph to allow automated, full georegistration of the data. The ground system component of the system is currently in use to fully georeference imagery collected in response to International Disaster Charter activations, and the auto-registration procedures are being applied to the extensive historical database of imagery to add value for research and educational purposes. In parallel, machine learning techniques are being applied to automate feature identification and classification throughout the dataset, in order to build descriptive metadata that will improve search capabilities. It is expected that these value additions will increase interest and use of the dataset by the global community.

Stefanov, William L.

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING

Using machine learning techniques to automate sky survey catalog generation

We describe the application of machine classification techniques to the development of an automated tool for the reduction of a large scientific data set. The 2nd Palomar Observatory Sky Survey provides comprehensive photographic coverage of the northern celestial hemisphere. The photographic plates are being digitized into images containing on the order of 10(exp 7) galaxies and 10(exp 8) stars. Since the size of this data set precludes manual analysis and classification of objects, our approach is to develop a software system which integrates independently developed techniques for image processing and data classification. Image processing routines are applied to identify and measure features of sky objects. Selected features are used to determine the classification of each object. GID3* and O-BTree, two inductive learning techniques, are used to automatically learn classification decision trees from examples. We describe the techniques used, the details of our specific application, and the initial encouraging results which indicate that our approach is well-suited to the problem. The benefits of the approach are increased data reduction throughput, consistency of classification, and the automated derivation of classification rules that will form an objective, examinable basis for classifying sky objects. Furthermore, astronomers will be freed from the tedium of an intensely visual task to pursue more challenging analysis and interpretation problems given automatically cataloged data.

Fayyad, Usama M.

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat

A Morphological Model to Separate Resolved–Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream

Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy Surveys (LS). We develop a supervised machine learning model based on the Gradient Boosting algorithm XGBoost. The features input to the model are purely morphological and are derived from the tabulated LS data products. We train the model using ∼2 × 10 5 LS sources in the COSMOS field with HST morphological labels and evaluate the model performance on LS sources with spectroscopic classification from the DESI Data Release 1 (∼2 × 10 7 objects) and the Sloan Digital Sky Survey Data Release 17 (∼3 × 10 6 objects), as well as on ∼2 × 10 8 Gaia stars. A significant fraction of LS sources are not observed in every LS filter, and we therefore build a “Hybrid” model as a linear combination of two XGBoost models, each containing features combining aperture flux measurements from the “blue” (gr) and “red” (iz) filters. The Hybrid model shows a reasonable balance between sensitivity and robustness, and achieves higher accuracy and flexibility compared to the LS morphological typing. With the Hybrid model, we provide classification scores for ∼3 × 10 9 LS sources, making this the largest ever machine learning catalog separating resolved and unresolved sources. The catalog has been incorporated into the real-time pipeline of the La Silla Schmidt Southern Survey (LS4), enabling the identification of extragalactic transients within the LS4 alert stream.

astrostatistics

Hybrid Quantum–Classical Graph Transformers for Efficient Sentiment Analysis

Quantum Machine Learning (QML) offers a promising paradigm that leverages quantum computing principles to develop efficient and expressive models for learning from complex and structured data. Recent advances in natural language processing (NLP) and artificial intelligence (AI) have demonstrated capabilities in understanding, generating, and reasoning over linguistic and multimodal information. In this work, we present the Quantum Graph Transformer (QGT), a hybrid quantum–classical architecture that extends graph transformer capabilities through quantum self-attention. The QGT models variable-length sentences as token graphs, where both the embedding encoding and the self-attention mechanisms are implemented using parameterized quantum circuits (PQCs), enabling efficient contextual learning with significantly fewer trainable parameters. We train QGT using both fully connected and 𝑘 -nearest-neighbor graph structures and evaluate it on five benchmark sentiment-classification datasets. Experimental results show that QGT consistently achieves higher or comparable accuracy to existing quantum NLP models and outperforms a Classical Graph Transformer (CGT) baseline with identical architecture, achieving 29.4 × fewer parameters while requiring 3–5 × fewer samples to reach comparable performance. These findings highlight the potential of graph-based quantum models as scalable and data-efficient architectures for natural language understanding.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Evaluating Material Design Principles for Calcium-Ion Mobility in Intercalation Cathodes

Multivalent-ion batteries offer an alternative to Li-based technologies, with the potential for greater sustainability, improved safety, and higher energy density, primarily due to their rechargeable system featuring a passivating metal anode. Although a system based on the Ca 2+ /Ca couple is particularly attractive given the low electrochemical plating potential of Ca 2+ , the remaining challenge for a viable rechargeable Ca battery is to identify Ca cathodes with fast ion transport. In this work, a high-throughput computational pipeline is adapted to (1) discover novel Ca cathodes in a largely unexplored space of empty intercalation hosts and (2) develop material design rules for Ca-ion mobility. One candidate from the screening, W 2 O 3 (PO 4 ) 2 , is confirmed to have a low Nudged Elastic Band (NEB) barrier of 168 meV within a one-dimensional (1D) ion percolation topology. This candidate is subsequently synthesized and electrochemically tested, achieving reversible Ca cycling with a capacity of 25 mA h/g. To further accelerate the screening for promising Ca intercalation electrodes, machine learning (ML) Random Forest (RF) and Extreme Gradient Boosting (XGB) classification models are created with local environment descriptors based on a large, structurally and chemically diverse dataset of minimum energy pathways, spanning over 5,000 density functional theory (DFT) site energy calculations. Accuracies of 92% are achieved, material design metrics are quantified, ML force-fields are leveraged in an accelerated iteration of the screening, and a total of 27 novel Ca cathode materials are highlighted for further investigation.

25 ENERGY STORAGE

Simultaneous Probe of the Charm and Bottom Quark Yukawa Couplings Using $t\bar{t}$𝐻 Events

A search for the standard model Higgs boson decaying to a charm quark-antiquark pair, 𝐻→$c\bar{c}$, produced in association with a top quark-antiquark pair ($t\bar{t}$𝐻) is presented. The search is performed with data from proton-proton collisions at √𝑠 =13 TeV, corresponding to an integrated luminosity of 138 fb−1. Advanced machine learning techniques are employed for jet flavor identification and event classification. The Higgs boson decay to a bottom quark-antiquark pair is measured simultaneously and the observed $t\bar{t}$𝐻(𝐻→$b\bar{b}$) event rate relative to the standard model expectation is 0.91$^{+0.26}_{−0.22}$. The observed (expected) upper limit on the product of production cross section and branching fraction 𝜎⁡($t\bar{t}$𝐻)⁢ℬ⁡(𝐻→$c\bar{c}$) is 0.11 (0.13) pb at 95% confidence level, corresponding to 7.8 (8.7) times the standard model prediction. When combined with the previous search for 𝐻 →$c\bar{c}$ via associated production with a 𝑊 or 𝑍 boson, the observed (expected) 95% confidence interval on the Higgs-charm Yukawa coupling modifier, 𝜅 𝑐 , is |𝜅 𝑐 | < 3.5 (2.7), the most stringent constraint to date.

Bottom quark

Medics: Medical Decision Support System for Long-Duration Space Exploration

The Autonomous Medical Operations (AMO) group at NASA Ames is developing a “medical decision support system” to enable astronauts on long-duration exploration missions to operate autonomously. The system will support clinical actions by providing medical interpretation advice and procedural recommendations during emergent care and clinical work performed by crew. The current state of development of the system, called MedICS (Medical Interpretation Classification and Segmentation) includes two separate aspects: a set of machine learning diagnostic models trained to analyze organ images and patient health records, and an interface to ultrasound diagnostic hardware and to medical repositories. Three sets of images of different organs and medical records were utilized for training machine learning models for various analyses, as follows: 1. Pneumothorax condition (collapsed lung). The trained model provides a positive or negative diagnosis of the condition. 2. Carotid artery occlusion. The trained model produces a diagnosis of 5 different occlusion levels (including “normal”). 3. Ocular retinal images. The model extracts optic disc pixels (image segmentation). This is a precursor step for advanced autonomous fundus clinical evaluation algorithms to be implemented in FY20. 4. Medical health records. The model produces a differential diagnosis for any particular individual, based on symptoms and other health and demographic information. A probability is calculated for each of 25 most common conditions. The same model provides the likelihood of survival. All results are provided with a confidence level. Item 1 images were provided by the US Army and were part of a data set for the clinical treatment of injured battlefield soldiers. This condition is relevant to possible space mishaps, due to pressure management issues. Item 2 images were provided by Houston Methodist Hospital, and item 3 health records were acquired from the MIT laboratory of computational physiology. The machine learning technology utilized is deep multilayer networks (Deep Learning), and new models will continue to be produced, as relevant data is made available and specific health needs of astronaut crews are identified. The interfacing aspects of the system include a GUI for running the different models, and retrieving and storing data, as well as support for integration with an augmented reality (AR) system deployed at JSC by Tietronix Software Inc. (HoloLens). The AR system provides guidance for the placement of an ultrasound transducer that captures images to be sent to the MedICS system for diagnosis. The image captured and the associated diagnosis appear in the technician’s AR visual display.

Colombano, Silvano

PhotonIDs: ML-Powered Photon Identification System for Dark Count Elimination

Reliable single photon detection is the foundation for practical quantum communication and networking. However, today's superconducting nanowire single photon detector(SNSPD) inherently fails to distinguish between genuine photon events and dark counts, leading to degraded fidelity in long-distance quantum communication. In this work, we introduce PhotonIDs, a machine learning-powered photon identification system that is the first end-to-end solution for real-time discrimination between photons and dark count based on full SNSPD readout signal waveform analysis. PhotonIDs ~demonstrates: 1) an FPGA-based high-speed data acquisition platform that selectively captures the full waveform of signal only while filtering out the background data in real time; 2) an efficient signal preprocessing pipeline, and a novel pseudo-position metric that is derived from the physical temporal-spatial features of each detected event; 3) a hybrid machine learning model with near 98% accuracy achieved on photon/dark count classification. Additionally, proposed PhotonIDs ~ is evaluated on the dark count elimination performance with two real-world case studies: (1) 20 km quantum link, and (2) Erbium ion-based photon emission system. Our result demonstrates that PhotonIDs ~could improve more than 31.2 times of signal-noise-ratio~(SNR) on dark count elimination. PhotonIDs ~ marks a step forward in noise-resilient quantum communication infrastructure.

Linne, Karl C. [Chicago U.] (ORCID:000900091870358

Rapid, antibiotic incubation-free determination of tuberculosis drug resistance using machine learning and Raman spectroscopy

Tuberculosis (TB) is the world’s deadliest infectious disease, with over 1.5 million deaths and 10 million new cases reported anually. The causative organism Mycobacterium tuberculosis (Mtb) can take nearly 40 d to culture, a required step to determine the pathogen’s antibiotic susceptibility. Both rapid identification and rapid antibiotic susceptibility testing of Mtb are essential for effective patient treatment and combating antimicrobial resistance. Here, we demonstrate a rapid, culture-free, and antibiotic incubation-free drug susceptibility test for TB using Raman spectroscopy and machine learning. We collect few-to-single-cell Raman spectra from over 25,000 cells of the Mtb complex strain Bacillus Calmette-Guérin (BCG) resistant to one of the four mainstay anti-TB drugs, isoniazid, rifampicin, moxifloxacin, and amikacin, as well as a pan-susceptible wildtype strain. By training a neural network on this data, we classify the antibiotic resistance profile of each strain, both on dried samples and on patient sputum samples. On dried samples, we achieve >98% resistant versus susceptible classification accuracy across all five BCG strains. In patient sputum samples, we achieve ~79% average classification accuracy. We develop a feature recognition algorithm in order to verify that our machine learning model is using biologically relevant spectral features to assess the resistance profiles of our mycobacterial strains. Finally, we demonstrate how this approach can be deployed in resource-limited settings by developing a low-cost, portable Raman microscope that costs <$5,000. We show how this instrument and our machine learning model enable combined microscopy and spectroscopy for accurate few-to-single-cell drug susceptibility testing of BCG.

60 APPLIED LIFE SCIENCES

Optimizing Neutrino Flavor Conversion Measurements through Machine Learning

The phenomenon of neutrino flavor conversion whereby the flavor of a neutrino particle can change between its time of production and later detection was the first definitive evidence of physics beyond the Standard Model. Some of the oscillation parameters used to describe this conversion are not yet well measured, leaving important questions still open regarding flavor conversion both in vacuum and as neutrinos travel through matter. NOvA is a long-baseline neutrino oscillation experiment that uses Fermilab's predominantly $\nu_\mu$ NuMI beam. A 14 kton oil-based liquid scintillator far detector 810 km away is used to measure neutrino oscillation through the $\nu_\mu$ disappearance and $\nu_e$ appearance channels. Super-K is a 50 kton water Cherenkov detector, which measures the disappearance of $\nu_e$ produced during solar fusion. The high density environment of the sun decreases the $\nu_e$ survival probability at higher energies observable in Super-K compared to the vacuum-dominated oscillations at lower energies. However, the transition region is overshadowed by radioactive background in the detector. In both of these experiments, the separation of neutrino detection events from background and classification of neutrino flavor are crucial tasks that benefit from the introduction of machine learning. Chapter 1 gives an overview of neutrinos and the context under which flavor conversion is measured in this dissertation. Chapters 2-7 present the results of a Bayesian sampling approach for the latest NOvA 3-flavor oscillation analysis with $26.6 \times 10^{20}$ protons on target in neutrino mode and $12.5 \times 10^{20}$ in antineutrino mode collected over 10 years. Chapters 8-13 present the results of extending the Super-K solar analysis to lower energies during its fourth phase with 2970 days of livetime.

Yankelevich, Alejandro Jaime [UC, Irvine]