Search NASASearch

SEARCH · Search NASA

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif

Real-Time Anomaly Detection for Beyond Standard Model Searches in ProtoDUNE Horizontal Drift

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events---making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving $31.9 \pm 0.2$\% ($26.6 \pm 0.2$\%) $\nu$ efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, $17.5 \pm 0.3$\% ($18.3 \pm 0.3$\%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, Cameron C. [Cincinnati U., RWC]

Real-Time Anomaly Detection for Beyond Standard Model Searches in ProtoDUNE Horizontal Drift

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events---making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving $31.9 \pm 0.2$\% ($26.6 \pm 0.2$\%) $\nu$ efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, $17.5 \pm 0.3$\% ($18.3 \pm 0.3$\%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, Cameron C. [Cincinnati U., RWC]

Real-Time Anomaly Detection for Searches Beyond the Standard Model in the ProtoDUNE Horizontal Drift Detector

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events—making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving 31.9 ± 0.2% (26.6 ± 0.2%) ν efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, 17.5 ± 0.3% (18.3 ± 0.3%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, C. [Cincinnati U., RWC]

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics

Exploring urban typologies using comprehensive analysis of transportation dynamics

Abstract As urban areas continue to expand and develop, categorizing cities into typologies offers a valuable framework for understanding metropolitan dynamics and fostering inter-city collaboration. However, existing typologies related to urban mobility have limitations, failing to consider cities within a single large urban region and often overlooking crucial dimensions such as trip demand and traffic flow. In this paper, we introduce a transportation-focused characterization for cities within a large urban region, specifically the San Francisco Bay Area, California. We incorporate over 40 metrics across five transportation dimensions: trip demand, road network, multi-modal network, traffic flow, and land use. Specifically, for the trip demand dimension, we include metrics capturing residents’ trip characteristics, such as mode share, intra-city trips, and inter-city trips. Additionally, we analyze the purpose of trips entering the city to gain a deeper understanding of incoming trip patterns. In the traffic flow dimension, we examine metrics like vehicle miles traveled, delay, and congestion to assess the traffic conditions on the street network. These, combined with other dimensions, provide a comprehensive view of a city’s transportation dynamics. Using unsupervised machine learning clustering methods, we identified eight distinct typologies for the Bay Area: Live Work Cities; Job and Activity Magnet Cities; Anchor Cities; Multi-modal Cities; Hyper-connected Cities; Low-density Residential Cities; Medium-density Residential Cities; and Mixed-use Residential Cities. Our findings show that many clusters are strongly influenced by trip demand and traffic flow metrics. Finally, we examine the practicality of this typology and its potential to guide collaborative transportation management strategies. The typologies provide a foundation for dialogue among Bay Area cities, focusing on evaluating shared characteristics and leveraging successes or challenges to develop unified strategies for transportation management.

Kuncheria, Anu

Time-resolved spray characterization via unified optical flow and binarization technique

This work leverages an unsupervised machine learning and advanced image processing techniques to characterize the breakup of fuel sprays in a small-scale combustor under reacting conditions, providing valuable insights into near-nozzle flow phenomenology. The proposed methodology integrates an improved optical flow model on a convolutional neural network to extract flow vectors with a binarization technique to assess droplets’ size and shape across the region of interest. The velocimetry approach demonstrates superior performance compared to a state-of-the-art optical flow model when applied to high-speed X-ray phase contrast spray images, achieving more accurate and reliable flow predictions. Moreover, breakup processes are quantified by breakup length and sphericity in accordance with velocity estimations, allowing a more complete characterization of the flow. This study establishes a robust methodology for analyzing spray morphology and primary breakup in compact combustors, contributing valuable means of understanding and optimizing fuel spray behavior in advanced combustion systems.

42 ENGINEERING

Real-space visualization of a defect-mediated charge density wave transition

Here, we study the coupled charge density wave (CDW) and insulator-to-metal transitions in the 2D quantum material 1T-TaS 2 . By applying in situ cryogenic 4D scanning transmission electron microscopy with in situ electrical resistance measurements, we directly visualize the CDW transition and establish that the transition is mediated by basal dislocations (stacking solitons). We find that dislocations can both nucleate and pin the transition and locally alter the transition temperature T c by nearly ~75 K. This finding was enabled by the application of unsupervised machine learning to cluster five-dimensional, terabyte scale datasets, which demonstrate a one-to-one correlation between resistance—a global property—and local CDW domain-dislocation dynamics, thereby linking the material microstructure to device properties. This work represents a major step toward defect-engineering of quantum materials, which will become increasingly important as we aim to utilize such materials in real devices.

4D-STEM

The mass profiles of dwarf galaxies from Dark Energy Survey lensing

We present a novel approach to extracting dwarf galaxies from photometric data to measure their average halo mass profile with weak lensing. We characterize their stellar mass and redshift distributions with a spectroscopic calibration sample. By combining the ${\sim} 5000\,\mathrm{deg}^2$ multiband photometry from the Dark Energy Survey and redshifts from the Satellites Around Galactic Analogs Survey with an unsupervised machine learning method, we select a low-mass galaxy sample spanning redshifts $z\lt 0.3$ and divide it into three mass bins. From low to high median mass, the bins contain [146 420, 330 146, 275 028] galaxies and have median stellar masses of $\log _{10}(M_*/\text{M}_\odot)=\left[8.52\substack{+0.57 -0.76},\, 9.02\substack{+0.50 -0.64},\, 9.49\substack{+0.50 -0.58}\right]$ . We measure the stacked excess surface mass density profiles, $\Delta \Sigma (R)$, of these galaxies using galaxy–galaxy lensing with a signal-to-noise ratio of [14, 23, 28]. Through a simulation-based forward-modelling approach, we fit the measurements to constrain the stellar-to-halo mass relation and find the median halo mass of these samples to be $\log _{10}(M_{\rm halo}/\text{M}_\odot)$ = [$10.67\substack{+0.2 -0.4}$, $11.01\substack{+0.14 -0.27}$, $11.40\substack{+0.08 -0.15}$]. The cold dark matter profiles are consistent with NFW (Navarro, Frenk, and White) profiles over scales ${\lesssim} 0.15 \, {h}^{-1}$ Mpc. We find that ${\sim} 20$ per cent of the dwarf galaxy sample are satellites. This is the first measurement of the halo profiles and masses of such a comprehensive, low-mass galaxy sample. The techniques presented here pave the way for extracting and analysing even lower mass dwarf galaxies and for more finely splitting galaxies by their properties with future photometric and spectroscopic survey data.

dark matter

Identifying Outliers in AI-based Image Compression

Image compression using artificial intelligence (AI) is becoming increasingly prevalent across various fields, including scientific research. Scientific instruments can generate hundreds of images per second, and effectively compressing these images with high compression ratios is crucial for facilitating scientific discoveries. However, automatically detecting outlier cases, where compression may not have succeeded or where interesting scientific phenomena are present, poses a significant challenge. To address this, we have developed a methodology based on unsupervised machine learning techniques for detecting outlier compressed images. This methodology utilizes metrics such as peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), structural texture similarity index measure (STSIM), and deep image and structural texture similarity index (DISTS). We have evaluated our methodology on several unlabeled datasets, including microscopy and x-ray images, and have successfully identified multiple outlier images using our proposed approach. Furthermore, our approach has enabled us to identify image semantics that are valuable for post-experiment analysis by scientists.

Data Analysis

Search for Beyond the Standard Model physics with anomaly detection in multilepton final states in pp collisions at s=13TeV with the ATLAS detector

A model-agnostic search for Beyond the Standard Model physics is presented, targeting final states with at least four light leptons (electrons or muons). The search regions are separated by event topology and unsupervised machine learning is used to identify anomalous events in the full 140 fb-1$$^{-1}$$ of proton–proton collision data collected with the ATLAS detector during Run 2. No significant excess above the Standard Model background expectation is observed. Model-agnostic limits are presented in each topology, along with limits on several benchmark models including vector-like leptons, wino-like charginos and neutralinos, or smuons. Limits are set on the flavourful vector-like lepton model for the first time.

Aad, G

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]

Synoptic Weather Regime Classifications for June, July, August, and September, 2022

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A). This dataset includes the data in June, July, August, and September; the last year of the data is 2022.

54 ENVIRONMENTAL SCIENCES

Synoptic Weather Regime Classifications for the whole year, from 2014 to 2015

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A).

54 ENVIRONMENTAL SCIENCES

Synoptic Weather Regime Classifications for June, July, August, from 2000 to 2024

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A).

node_som_ml

Synoptic Weather Regime Classifications for March, April and May, from 2000 to 2025

The synoptic weather regime classification has become a highly demanded product for the ARM site in recent years. This type of regime classification has shown applications in various studies and topics, including aerosol-cloud interactions, land-atmosphere interactions, and cloud radiative effects. The VAP employs an unsupervised machine learning method, Self-organizing map (SOM), to classify weather regimes for each day of the AMF campaigns and fixed sites, using ERA5 data. This idea is mainly based on our published study for TRACER in Wang et al. (2022, JGR-A).

node_som_ml