Search NASA⌕ Search

SEARCH · Search NASA

Results for “machine learning classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

GeoAI advances in specific landform mapping

Landform mapping (also referred to as geomorphology or geomorphometry) can be divided into two domains: general and specific (Evans 2012). Whereas general landform mapping categorizes all elements of the study area into landform classes, such as ridges, valleys, peaks, and depressions, the mapping of specific landforms requires the delineation (even if fuzzy) of individual landforms. The former is mainly driven by physical properties such as elevation, slope, and curvature. The latter, however, must consider the cognitive (human) reasoning that discriminates individual landforms in addition to these physical properties (Arundel and Sinha 2018). Both mapping forms are important. General geomorphometry is needed to understand geological and ecological processes and as boundary layer input to climate and environmental models. Specific geomorphometry supports such activities as disaster management and recovery, emergency response, transportation, and navigation. In the United States, individual landforms of interest are named in the U.S. Geological Survey (USGS) Geographic Names Information System, a point dataset captured specifically to digitize geographic names from the USGS Historical Topographic Map Collection (HTMC). Named landform extent is represented only by the name placement in the HTMC. Recent work has investigated CNN-based deep learning methods to capture these extents in machine-readable form. These studies first relied on physical properties (Arundel et al. 2020) and then included the HTMC as a band in RGB images in limited testing (Arundel et al. 2023). Results from the HTMC dataset surpassed those using just physical properties and using the HTMC alone performed best due to the hillshading and elevation (contour) data incorporated into the topographic maps. However, results fell short of an operational capacity to map all named landforms in the United States. Thus, our current work expands upon past research by focusing on the HTMC and physical information as inputs and the named landform label extents. Specifically, we propose to leverage pre-trained foundation models for segmentation and optical character recognition (OCR) models to jointly map landforms in the United States. Our approach aims to bridge the disparities among the independent information sources to facilitate informed decision-making. The modeling pipeline performs (1) segmentation using the physical information and (2) information extraction using OCR, in parallel. Then a computer vision approach merges the two branches into a labeled segmentation. References: Arundel, Samantha T., Wenwen Li, and Sizhe Wang. 2020. “GeoNat v1.0: A Dataset for Natural Feature Mapping with Artificial Intelligence and Supervised Learning.” Transactions in GIS 24 (3): 556–72. https://doi.org/10.1111/tgis.12633. Arundel, Samantha T, and Gaurav Sinha. 2018. “Validating GEOBIA Based Terrain Segmentation and Classification for Automated Delineation of Cognitively Salient Landforms BT - Proceedings of Workshops and Posters at the 13th International Conference on Spatial Information Theory (COSIT 2017).” In Proceedings of Workshops and Posters at the 13th International Conference on Spatial Information Theory (COSIT 2017), Lecture Notes in Geoinformation and Cartography, edited by Paolo Fogliaroni, Andrea Ballatore, and Eliseo Clementini, 9–14. Cham: Springer International Publishing. Arundel, Samantha T., Gaurav Sinha, Wenwen Li, David P. Martin, Kevin G. McKeehan, and Philip T. Thiem. 2023. “Historical Maps Inform Landform Cognition in Machine Learning.” Abstracts of the ICA 6 (August): 1–2. https://doi.org/10.5194/ica-abs-6-10-2023. Evans, Ian S. 2012. “Geomorphometry and Landform Mapping: What Is a Landform?” Geomorphology 137 (1): 94–106. https://doi.org/10.1016/j.geomorph.2010.09.029.

machine learning↗

Towards a Marine Stratus Climatology on Drizzle Occurrence from CALIPSO

Marine stratus are a predominant feature of our planet with the annual mean coverage exceeding 20%. They strongly reflect sunlight, yet exert only a modest effect on outgoing infrared radiation, providing a significant net cooling to the Earth’s radiative balance. Their formation is coupled to boundary layer circulations that are driven, in part, by cloud top radiative cooling and evaporative cooling from precipitation in downdrafts. Understanding how these cloud systems evolve as the climate changes is a key question that requires additional information on their lifecycle and microphysical properties to accurately represent their behavior in global circulation models. From a large-scale perspective, insight into the microphysical properties of marine stratus at cloud top can be realized through estimates of the effective radius (Re) of the droplet size distributions derived from MODIS observations. Estimates on the occurrence of rain/drizzle are available from CloudSat. Together these observations indicate that precipitation frequently occurs in clouds with higher cloud top Re. This relationship is consistent with the well documented shift in cloud top droplet size distributions towards fewer, yet larger droplets prior the onset of precipitation. Here we report on a new and complementary set observations from the CALIPSO mission. The approach derives an extinction-to-backscatter ratio (Sc, also known as the cloud lidar ratio) using an established relationship that depends on observations of the lidar attenuated backscatter and volume depolarization ratio within the cloud. Because Sc is strongly and inversely related to Re, a change in the derived Sc from higher to lower values corresponds to a change in the droplet size distribution as seen by MODIS. This change in the lidar signals at cloud top clearly identifies clouds that are capable of precipitation. The presentation provides a brief overview of the approach for deriving Sc and compares CALIOP-derived Sc with observations from other techniques. CALIOP classifications of drizzling clouds, based on the retrieved values Sc, are compared to independent, collocated assessments of drizzle occurrence reported in the standard CloudSat data products. Regional and seasonal comparisons highlight the strengths and weaknesses of the two sensors. A machine learning approach that combines information from both CALIOP and CloudSat showcases possible improvements in the global identification of scenes likely to contain rain-bearing clouds.

CALIPSO↗

Anomalous electroweak physics unraveled via evidential deep learning

The ever-growing ecosystem of beyond standard model (BSM) calculations and parametrizations has motivated the development of systematic methods for making quantitative cross-comparisons over the wide range of possible models, especially with controllable uncertainties. In this setting, the language of uncertainty quantification (UQ) furnishes useful metrics for assessing statistical overlaps and discrepancies among BSM and related models. In this study, we leverage recent machine learning (ML) developments in evidential deep learning (EDL) for UQ to separate data (aleatoric) and knowledge (epistemic) uncertainties in a model-discrimination setting. We construct several potentially BSM-motivated scenarios for the anomalous electroweak interaction (AEWI) of neutrinos with nucleons in deep inelastic scattering ( v DIS). These scenarios are then quantitatively mapped, as a demonstration, alongside Monte Carlo replicas of the CT18 PDFs used to calculate the $\varDelta \chi ^{2}$ statistic for a typical multi-GeV v DIS experiment, CDHSW. Our framework effectively highlights areas of model agreement and provides a classification of out-of-distribution (OOD) samples. By offering the opportunity to quantitatively understand model overlaps, the approach presented in this work can help facilitate efficient BSM model exploration and exclusion for future New Physics searches.

AI↗

The phase space distance between collider events

How can one fully harness the power of physics encoded in relativistic N-body phase space? Topologically, phase space is isomorphic to the product space of a simplex and a hypersphere and can be equipped with explicit coordinates and a Riemannian metric. This natural structure that scaffolds the space on which all collider physics events live opens up new directions for machine learning applications and implementation. Here we present a detailed construction of the phase space manifold and its differential line element, identifying particle ordering prescriptions that ensure that the metric satisfies necessary properties. We apply the phase space metric to several binary classification tasks, including discrimination of high-multiplicity resonance decays or boosted hadronic decays of electroweak bosons from QCD processes, and demonstrate powerful performance on simulated data. Our work demonstrates the many benefits of promoting phase space from merely a background on which calculations take place to being geometrically entwined with a theory’s dynamics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Quantum Transfer Learning to Boost Dementia Detection

Dementia is a devastating condition with profound implications for individuals, families, and healthcare systems. Early and accurate detection of dementia is critical for timely intervention and improved patient outcomes. While classical machine learning and deep learning approaches have been explored extensively for dementia prediction, these solutions often struggle with high-dimensional biomedical data and large-scale datasets, quickly reaching computational and performance limitations. To address this challenge, quantum machine learning (QML) has emerged as a promising paradigm, offering faster training and advanced pattern recognition capabilities. This work aims to demonstrate the potential of quantum transfer learning (QTL) to enhance the performance of a weak classical deep learning model applied to a binary classification task for dementia detection. Besides, we show the effect of noise on the QTL-based approach, investigating the reliability and robustness of this method. Using the OASIS 2 dataset, we show how quantum techniques can transform a suboptimal classical model into a more effective solution for biomedical image classification, highlighting their potential impact on advancing healthcare technology.

Bhowmik, Sounak [University of Tennessee, Knoxvill↗

Multivariate Time Series Intermittent Fault Detectionin Controller Area Network CAN

Fault detection in Controller Area Network (CAN) systems is crucial for ensuring the reliability and safety of automotive and industrial applications. This study investigates and compares the effectiveness of time series classification models for supervised fault detection in CAN data. This repository contains the code and data for our benchmarking experiment aimed at detecting intermittent faults in automotive Controller Area Network (CAN) data. The goal of this project is to compare various machine learning (ML) and deep learning (DL) models using different Time Series Cross-Validation (TSCV) techniques to evaluate their effectiveness in a streaming environment for fault detection.

Hespeler, Steven [Oak Ridge National Laboratory (O↗

Confusion-Driven Machine Learning of Structural Phases of a Flexible, Magnetic Stockmayer Polymer

We use a semisupervised, neural-network-based machine learning technique, the confusion method, to investigate structural transitions in magnetic polymers, which we model as chains of magnetic colloidal nanoparticles characterized by dipole–dipole and Lennard-Jones interactions. As input for the neural network, we use the particle positions and magnetic dipole moments of equilibrium polymer configurations, which we generate via replica-exchange Wang–Landau simulations. We demonstrate that by measuring the classification accuracy of neural networks, we can effectively identify transition points between multiple structural phases without any prior knowledge of their existence or location. We corroborate our findings by investigating relevant conventional order parameters. Our study furthermore examines previously unexplored low-temperature regions of the phase diagram, where we find new structural transitions between highly ordered helicoidal polymer configurations.

36 MATERIALS SCIENCE↗

Quality of Candidate Flights and Submission Prediction in Collaborative Digital Departure Reroute

Collaborative Digital Departure Reroute (CDDR) enables the reroute of flights using a flight operator proposed set of alternative route options, referred to as Trajectory Option Set (TOS), in order to reduce delay on the airport's surface and in the Metroplex environment. The reroute functionality is enabled through NASA's Digital Information Platform (DIP). TOS candidate flights are defined as flights with an alternative route with delay savings greater than the flight operator defined relative trajectory cost. This paper analyzes the TOS candidate flights at Dallas/Fort Worth International Airport (KDFW) in the North Texas Metroplex to gain insight into which candidate flights are higher quality through a scoring method. This insight will inform refinements to help CDDR focus on high quality reroute opportunities. Binary classification models for predicting the flight operator's submission of candidate flights are also explored in this paper.

Machine Learning↗

Quality of Candidate Flights and Submission Prediction in Collaborative Digital Departure Reroute

Collaborative Digital Departure Reroute (CDDR) enables the reroute of flights using a flight operator proposed set of alternative route options, referred to as Trajectory Option Set (TOS), in order to reduce delay on the airport's surface and in the Metroplex environment. The reroute functionality is enabled through NASA's Digital Information Platform (DIP). TOS candidate flights are defined as flights with an alternative route with delay savings greater than the flight operator defined relative trajectory cost. This paper analyzes the TOS candidate flights at Dallas/Fort Worth International Airport (KDFW) in the North Texas Metroplex to gain insight into which candidate flights are higher quality through a scoring method. This insight will inform refinements to help CDDR focus on high quality reroute opportunities. Binary classification models for predicting the flight operator's submission of candidate flights are also explored in this paper.

Machine Learning↗

SafeAeroBERT: Towards a Safety-Informed Aerospace-Specific Language Model

As aviation systems continue to operate with high traffic, large amounts of documents containing safety-relevant data continue to be generated via reporting systems such as the ASRS. Advanced natural language processing techniques, specifically pre-trained language models, have shown great success in domain-specific applications; however, the text in aviation safety reports is inundated with jargon and thus not fully utilized by general pre-trained models. In this research, we work towards developing a safety-informed aerospace-specific language model by pre-training a Bidirectional Encoder Representations from Transformer (BERT) model on reports from the Aviation Safety Reporting System and the National Transportation Safety Board. The resulting model, called SafeAeroBERT, is fine-tuned for the specific task of document classification, and can be further tuned for named-entity recognition, relation detection, information retrieval, and summarization. Results from the classification task are compared between SafeAeroBERT, the base BERT, and SciBERT models and show SafeAeroBERT outperforms the general BERT and SciBERT on classifying reports about human factors, aircraft, and procedure. SafeAeroBERT can be used on custom tasks, not limited to document classification, and is intended to aid an intelligent knowledge manager for safety report repositories.

Aviation↗

JGI-Trichoderma v1.0

There is a series of Python and bash scripts to parse genomics datasets used to evaluate the coevolution of gene families and the feature importance of gene families using an SVM classifier. - Cover analysis: takes a list of single-copy genes in a set of genomes, aligns and builds the gene trees to determine if two gene families have a signature of covariation with one another. It parses the files to run phykit cover script described here: https://jlsteenwyk.com/PhyKIT/usage/index.html - SVM-classifier: This Python script is an SVM-based genomic classifier designed for biological data analysis. It combines machine learning with feature selection to identify important genomic markers and classify biological samples. Core Functionality: The script uses Support Vector Machines from scikit-learn to classify genomic data, incorporating SelectKBest for automated feature selection and leave-one-out cross-validation for performance assessment. It operates in multiple modes: feature ranking, optimal combination discovery, and sample prediction. Primary Applications: Genomic sample classification and biomarker discovery Feature importance analysis in high-dimensional biological datasets Prediction of sample categories based on genomic profiles Research applications requiring robust classification of biological data Key Advantages: High-dimensional handling: SVMs excel with genomic data's typical high feature-to-sample ratios Integrated feature selection: Reduces noise and computational overhead while identifying key markers Probability estimation: Provides confidence scores essential for biological interpretation Validation robustness: Leave-one-out cross-validation ensures reliable performance metrics Operational flexibility: Multiple analysis modes support different research phases from exploration to prediction

Stecca Steindorff, Andrei [Lawrence Berkeley Natio↗

Simulated meteorological impacts of offshore wind turbines and sensitivity to the amount of added turbulence kinetic energy

Offshore wind energy projects are currently in development off the east coast of the United States and may influence the local meteorology of the region. Wind power production and other commercial uses in this area are related to atmospheric conditions, and so it is important to understand how future wind plants may change the local meteorology. In the absence of measurements of potential wind plant impacts on meteorology, simulations offer the next-best possible insight into wake effects on boundary layer height, temperature, fluxes, and wind speeds. However, simulation tools that capture these effects offer multiple options for representing the amount of turbine-added turbulence that may impact assessments of micrometeorological effects. To explore this sensitivity, we compare 1 year of simulations from the Weather Research and Forecasting (WRF) model with and without wind plants incorporated, focusing on the lease area south of Massachusetts and Rhode Island. The simulations with wind plants are repeated to include both the maximum and minimum amounts of added turbulence to provide bounds on the potential impacts. We assess changes in wind speeds, 2 m temperature, surface heat flux, turbulence kinetic energy (TKE), and boundary layer height during different stability classifications and ambient wind speeds over the entire year and compare results for the degree of added turbulence in the wind plant simulations. Because the wake behavior may be a function of boundary layer stability, in this paper, we also present a machine learning algorithm to quantify the area and distance of the wake generated by the wind plant. This analysis enables us to identify the relationship between wake extent and boundary layer height. We find that hub-height wind speed is reduced within and downwind of the wind plant, with the strongest impacts occurring during stable conditions and faster wind speeds in region 3 of the turbine power curve, although impacts lessen as wind speeds increase past 15 m s−1. In contrast, wind speeds near the surface decrease when no turbine-added turbulence is included but can increase for stably stratified conditions when 100 % of possible TKE is included in the simulations. TKE increases at hub height in the simulations with added TKE for all stability classes, suggesting that atmospheric stability does not immediately modify the TKE generated by turbines. Negligible changes in hub-height TKE manifest in the simulations without the added TKE. At the surface, TKE increases in the simulations with maximum added turbulence only for unstable conditions. In the no-added-turbulence simulations, surface TKE decreases slightly in neutral and unstable simulations. Differences in 2 m temperatures and surface heat fluxes are small but vary considerably with atmospheric stability and the amount of added TKE. Boundary layer heights increase within the wind plant when turbine-added turbulence is included and decrease slightly downwind during stable conditions. In contrast, with no added turbulence, the boundary layer height is in general reduced in stable conditions with wind speeds less than 15 m s −1 and slightly increased in neutral conditions. Finally, shallower upwind boundary layer heights tend to correlate with larger wake areas and distances, though other factors likely also play a role in determining the extent of the wind plant wake. These simulation-based results provide a bound for micrometeorological impacts of wind plant wakes: simulations that couple the atmosphere to the ocean may reduce these impacts, and we await observational verification.

17 WIND ENERGY↗

Benchmark Models for Classification of Radiation Type Induced in Immune Cells

NASA Biological and Physical Sciences and the Science Mission Directorate have published a benchmark dataset of mouse immune cells subjected to radiation-induced DNA damage. The dataset comprises ML-ready microscopic imagery of said cells, including labels indicating radiation type and dose. The machine learning team at NASA Interagency Implementation and Advanced Concept Team (IMPACT) created multiple benchmark models. Initially, we conducted a preliminary analysis using thresholding. The algorithm used thresholds on average brightness of the available images to classify them into their respective radiation type. We also tested machine learning approaches. Convolutional Neural Networks (CNN) emerged as the best-performing model. This poster presents the benchmark scores obtained by the models.

Vishal Perekadan↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Machine learning reveals genes impacting oxidative stress resistance across yeasts

Reactive oxygen species (ROS) are highly reactive molecules encountered by yeasts during routine metabolism and during interactions with other organisms, including host infection. Here, we characterize the variation in resistance to the ROS-inducing compound tert -butyl hydroperoxide across the ancient yeast subphylum Saccharomycotina and use machine learning (ML) to identify gene families whose sizes are predictive of ROS resistance. The most predictive features are enriched in gene families related to cell wall organization and include two reductase gene families. We estimate the quantitative contributions of features to each species’ classification to guide experimental validation and show that overexpression of the old yellow enzyme (OYE) reductase increases ROS resistance in Kluyveromyces lactis , while Saccharomyces cerevisiae mutants lacking multiple mannosyltransferase-encoding genes are hypersensitive to ROS. Altogether, this work provides a framework for how ML can uncover genetic mechanisms underlying trait variation across diverse species and inform trait manipulation for clinical and biotechnological applications.

59 BASIC BIOLOGICAL SCIENCES↗

One-shot gas detection with transformer paired neural networks in Mako collected longwave infrared hyperspectral imagery

To date, careful data treatment workflows and statistical detectors are used to perform hyperspectral image (HSI) detection of any gas contained in a spectral library, which is often expanded with physics models to incorporate different spectral characteristics. In general, surrounding evidence or known gas-release parameters are used to provide confidence in or confirm detection capability, respectively. This makes quantifying detection performance difficult as it is nearly impossible to develop an absolute ground truth for gas target pixel presence in collected HSI. Consequently, developing and comparing new detection methods, especially machine learning (ML)-based methods, is susceptible to subjectivity in derived detection map quality. Here, in this work, we demonstrate the first use of transformer-based paired neural networks (PNNs) for one-shot gas target detection for multiple gases while providing quantitative classification and detection metrics for their use on labeled data. Terabytes of training data are generated from a database of long-wave infrared HSI obtained from historical Mako sensor campaigns over Los Angeles. By incorporating labels, singular signature representations, and a model development pipeline, we can tune and select PNNs to detect multiple gas targets that are not seen in training on a quantitative basis. We additionally assess our test set detections using interpretability techniques widely employed with ML-based predictors, but less common with detection methods relying on learned latent spaces.

Hyperspectral imaging↗

A Machine Learning Concept for DTN Routing

This paper discusses the concept and architecture of a machine learning based router for delay tolerant space networks. The techniques of reinforcement learning and Bayesian learning are used to supplement the routing decisions of the popular Contact Graph Routing algorithm. An introduction to the concepts of Contact Graph Routing, Q-routing and Naive Bayes classification are given. The development of an architecture for a cross-layer feedback framework for DTN (Delay-Tolerant Networking) protocols is discussed. Finally, initial simulation setup and results are given.

Delay Tolerant Networks↗