Search NASA⌕ Search

SEARCH · Search NASA

Results for “Unsupervised Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Can Selforganizing Maps Accurately Predict Photometric Redshifts?

We present an unsupervised machine-learning approach that can be employed for estimating photometric redshifts. The proposed method is based on a vector quantization called the self-organizing-map (SOM) approach. A variety of photometrically derived input values were utilized from the Sloan Digital Sky Survey's main galaxy sample, luminous red galaxy, and quasar samples, along with the PHAT0 data set from the Photo-z Accuracy Testing project. Regression results obtained with this new approach were evaluated in terms of root-mean-square error (RMSE) to estimate the accuracy of the photometric redshift estimates. The results demonstrate competitive RMSE and outlier percentages when compared with several other popular approaches, such as artificial neural networks and Gaussian process regression. SOM RMSE results (using delta(z) = z(sub phot) - z(sub spec)) are 0.023 for the main galaxy sample, 0.027 for the luminous red galaxy sample, 0.418 for quasars, and 0.022 for PHAT0 synthetic data. The results demonstrate that there are nonunique solutions for estimating SOM RMSEs. Further research is needed in order to find more robust estimation techniques using SOMs, but the results herein are a positive indication of their capabilities when compared with other well-known methods

Way, Michael J.↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology↗

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory ↗

Optimal Electrification Using Renewable Energies: Microgrid Installation Model with Combined Mixture k-Means Clustering Algorithm, Mixed Integer Linear Programming, and Onsset Method

Optimal planning and design of microgrids are priorities in the electrification of off-grid areas. Indeed, in one of the Sustainable Development Goals (SDG 7), the UN recommends universal access to electricity for all at the lowest cost. Several optimization methods with different strategies have been proposed in the literature as ways to achieve this goal. This paper proposes a microgrid installation and planning model based on a combination of several techniques. The programming language Python 3.10 was used in conjunction with machine learning techniques such as unsupervised learning based on K-means clustering and deterministic optimization methods based on mixed linear programming. These methods were complemented by the open-source spatial method for optimal electrification planning: onsset. Four levels of study were carried out. The first level consisted of simulating the model obtained with a cluster, which is considered based on the elbow and k-means clustering method as a case study. The second level involved sizing the microgrid with a capacity of 40 kW and optimizing all the resources available on site. The example of the different resources in the Togo case was considered. At the third level, the work consisted of proposing an optimal connection model for the microgrid based on voltage stability constraints and considering, above all, the capacity limit of the source substation. Finally, the fourth level involved a planning study of electrification strategies based mainly on microgrids according to the study scenario. The results of the first level of study enabled us to obtain an optimal location for the centroid of the cluster under consideration, according to the different load positions of this cluster. Then, the results of the second level of study were used to highlight the optimal resources obtained and proposed by the optimization model formulated based on the various technology costs, such as investment, maintenance, and operating costs, which were based on the technical limits of the various technologies. In these results, solar systems account for 80% of the maximum load considered, compared to 7.5% for wind systems and 12.5% for battery systems. Next, an optimal microgrid connection model was proposed based on the constraints of a voltage stability limit estimated to be 10% of the maximum voltage drop. The results obtained for the third level of study enabled us to present selective results for load nodes in relation to the source station node. Finally, the last results made it possible to plan electrification using different network technologies and systems in the short and long term. The case study of Togo was taken into account. The various results obtained from the different techniques provide the necessary leads for a feasibility study for optimal electrification of off-grid areas using microgrid systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Unsupervised Learning for Improved Gamma-Ray Spectrometry in Pixelated Cadmium Zinc Telluride (CZT) Detectors

Machine learning has been found to be ubiquitously useful across many industries, presenting an opportunity to improve radiation detection performance using data-driven algorithms. Improved detector resolution can aid in the detection, identification, and quantification of radionuclides. Here, in this work, a novel, data-driven, unsupervised learning approach is developed to improve detector spectral characteristics by learning, and subsequently rejecting, poorly performing regions of the pixelated detector. Feature engineering is used to fit individual characteristic photo peaks to a Doniach lineshape with a linear background model. Then, principal component analysis is used to learn a lower-dimension latent space representation of each photo peak where the pixels are clustered, and subsequently ranked, based on the cluster mean distance to an optimal point. Pixels within the worst cluster(s) are rejected to improve the full-width at half-maximum (FWHM) by 10% to 15% (relative to the bulk detector) at 50% net efficiency when applied to training data obtained from measurements of a 100 μCi 154 Eu source using a H3D M400i pixelated cadmium zinc telluride detector. These results compare well with, but do not outperform, a greedy algorithm that accumulates pixels in order of FWHM from lowest to highest used as a benchmark. In the future, this approach can be extended to include the detector energy and angular response. Finally, the model is applied to newly seen natural and enriched uranium spectra relevant for nuclear safeguards applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Combined Machine Learning and Molecular Dynamics Reveal Two States of Hydration of a Single Functional Group of Cationic Polymeric Brushes

The state of hydration of a macromolecular system regulates a plethora of different properties of such a system. In this article, we develop a novel machine learning (ML) approach, based on the unsupervised clustering algorithm, for probing the hydration behavior of the {N(CH 3 ) 3 } + functional group of the PMETAC [Poly(2-(methacryloyloxy)ethyl trimethylammonium chloride] polyelectrolyte (PE) brush system. The PE brushes and the brush-supported water molecules and counterions (chloride ions) are first described using all-atom molecular dynamics (MD) simulations. The simulation data is subsequently used in our ML framework to identify that (1) the {N(CH 3 ) 3 } + functional groups of the PMETAC brushes have two distinct hydration states with one state (state 1) being characterized by less structured water molecules and the other state (state 2) being characterized by more structured water molecules and (2) an enhancement in the brush grafting density leads to the progressive dissapparenace of state 2. An increase in the grafting density increases the number of chloride counterions in a given volume around the {N(CH 3 ) 3 } + functional group and increases the number of shared water molecules between the {N(CH 3 ) 3 } + and Cl - . The chloride counterions are associated with a hydration layer with much less structured water molecules. Therefore, with an increase in the grafting density, an increase in the percentage of shared water molecules leads to the prevalence of the hydration state [of the {N(CH 3 ) 3 } + moiety] with less structured water molecules. Finally, we explain how the present findings are commensurate with two key previous related results, namely a significantly large chloride ion mobility inside the PMETAC brush layer and the {N(CH 3 ) 3 } + -Cl - average distance remaining independent of the PMETAC brush grafting density. Furthermore, we anticipate that the combined ML-MD-simulation approach proposed in this study can be adapted to probe other soft matter systems to reveal new insights of the underlying mechanisms of emergent phenomenon.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unsupervised Clustering and Supervised Regression Learning to Select High Temperature Oxidation-Resistant Materials

High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.

Romedenne, Marie [ORNL] (ORCID:0000000317936561)↗

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

ArcjetCV: A New Machine Learning Application for Extracting Time-Resolved Recession Measurements From Arc Jet Test Videos

Arc jet Computer Vision (ArcjetCV) is a software application built to automate analysis of arc jet ground test video footage. This includes tracking material recession, sting arm motion, and the shock-material standoff distance. This provides a new capability to resolve and validate new physics associated with non-linear processes. This is an essential step to reduce testing, modeling, and validation uncertainties for heatshield material performance. ArcjetCV uses several types of machine learning (convolutional neural net: CNN, decision tree: DT, k-means unsupervised clustering: KM) to automate the video processing pipeline. These include inferring the start/stop of time segments of interest (1D CNN), measuring the time-dependent 2D recession of the material samples (2D CNN, DT), measuring the time-dependent shock standoff distance (2D CNN, DT), and post-processing cleaning of the recession data (KM). The software also provides a graphical user interface for ease of use. The results of using this tool on arc jet videos show non-linear time-dependent effects can be important for certain materials.

machine learning↗

ArcjetCV: a new machine learning application for extracting time-resolved recession measurements from arc jet test videos

Arc jet Computer Vision (ArcjetCV) is a software application built to automate analysis of arc jet ground test video footage. This includes tracking material recession, sting arm motion, and the shock-material standoff distance. This provides a new capability to resolve and validate new physics associated with non-linear processes. This is an essential step to reduce testing, modeling, and validation uncertainties for heatshield material performance. ArcjetCV uses several types of machine learning (convolutional neural net: CNN, decision tree: DT, k-means unsupervised clustering: KM) to automate the video processing pipeline. These include inferring the start/stop of time segments of interest (1D CNN), measuring the time-dependent 2D recession of the material samples (2D CNN, DT), measuring the time-dependent shock standoff distance (2D CNN, DT), and post-processing cleaning of the recession data (KM). The software also provides a graphical user interface for ease of use. The results of using this tool on arc jet videos show non-linear time-dependent effects can be important for certain materials.

machine learning↗

ArcjetCV: A New Machine Learning Application for Extracting Time-Resolved Recession Measurements From Arc Jet Test Videos

Arc jet Computer Vision (ArcjetCV) is a software application built to automate analysis of arc jet ground test video footage. This includes tracking material recession and the shock-material standoff distance. This provides a new capability to resolve and validate new physics associated with non-linear processes. This is an essential step to reduce testing, modeling, and validation uncertainties for heatshield material performance. ArcjetCV uses several types of machine learning (convolutional neural net: CNN, decision tree: DT, k-means unsupervised clustering: KM) to automate the video processing pipeline. These include inferring the start/stop of time segments of interest (1D CNN), measuring the time-dependent 2D recession of the material samples (2D CNN, DT), measuring the time-dependent shock standoff distance (2D CNN, DT), and post-processing cleaning of the recession data (KM). The software also provides a graphical user interface for ease of use. The results of using this tool on arc jet videos show non-linear time-dependent effects can be important for certain materials and characterizing certain failure modes.

machine learning↗