Search NASASearch

SEARCH · Search NASA

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Aurora Detection From Nighttime Lights for Earth and Space Science Applications

This research leverages data from the Day/Night Band (DNB) of the Visible Infrared Imaging Radiometer (VIIRS) instrument onboard the Suomi National Polar-orbiting Partnership (S-NPP) satellite. We demonstrate the value of mining the VIIRS DNB for aurora and describe our use of unsupervised machine learning to create a binary mask for aurora occurrence. This mask can be used to flag aurora-contaminated observations for NASA's nighttime lights products for Earth science applications. The identification of auroral regions can also be used for Space Weather applications, for example, for comparison with aurora forecast model and with other satellite- or ground-based aurora observations. The DNB is a broadband channel that is sensitive to wavelengths from 500 to 900 nm, which covers most of the visible light spectrum, and as the name implies, captures light even at night with a sensitivity at the nanowatt level. This band is suitable for aurora observations since the light emitted by the aurora tends to be dominated by emissions from atomic oxygen, resulting in a greenish glow at a wavelength of 557.7 nm, especially at an altitude of 110 km. This study compares the global nighttime derived aurora regions for 17 and 18 March with the NOAA Space Weather Prediction Center's (SWPC) probability product for the St. Patrick's Day geomagnetic storm in 2015. VIIRS sensors are slated to be added to the next generation of polar-orbiting operational satellites. Our novel automated approach to aurora identification opens up an efficient way to leverage this unique data source.

Aurora

Climatology of Global Precipitation Measurement Mission Precipitation Regimes and Implications for Global Estimates of Vertical Winds

The Global Precipitation Measurement (GPM) mission Validation Network (VN) framework leverages over 118 ground-based polarimetric Doppler radars to validate a large subset of precipitation measurements and retrievals from the GPM Dual-frequency Precipitation Radar (DPR). Recently, GPM DPR reflectivity profiles within the VN have been classified according to their convective regime using unsupervised machine learning techniques. The archetypal regimes are stratiform, convective, mixed stratiform-convective (e.g., transition regions), and “other” (e.g., peripheral regions of light precipitation). Subcategories within these four primary regimes vary according to the characteristic depth of included reflectivity profiles, resulting in 12 main GPM DPR precipitation profile categories. Polarimetry of ground-based Doppler radars in the VN offers additional insights into the types of precipitation, while pairs of radars positioned near each other enable retrieval of vertical winds via dual-Doppler analysis. Geometrically matched to the DPR reflectivity profiles in the GPM VN, these ground-based data and retrievals contribute more detailed characterization of the distinct kinematic and microphysical structures associated with each of the 12 DPR precipitation regimes. DPR reflectivity profiles linked with wind in the VN are restricted to GPM overpasses of proximal radar pairs that allow dual-Doppler analysis. Although a limited subset of DPR profiles in the VN are matched with vertical motion, agreement between the reflectivity structures paired with wind data and those of the greater DPR dataset in the VN suggest that estimates of vertical motion may be inferred in regions without ground-based measurements. We present a climatology of the 12 convective regimes identified within the DPR VN dataset as well as early efforts to estimate the kinematic and microphysical structures of precipitation profiles within the greater GPM DPR dataset by applying machine learning techniques. Precipitation data paired with global estimates of vertical winds from these efforts offer early insight to and support upcoming missions to retrieve convective mass flux, including the Investigation of Convective Updrafts (INCUS) in the Tropics and the global Atmosphere Observing System (AOS).

Precipitation

Machine Learning in the Context of Laser-Induced Breakdown Spectroscopy

The integration of machine learning (ML) with Laser-Induced Breakdown Spectroscopy (LIBS) has revolutionized the analytical capabilities of LIBS. The combi-nation of both methods enables more accurate and efficient data analysis. While LIBS itself is a powerful technique for elemental analysis, the vast amount of spectral data it generates can be hard to interpret. Machine learning addresses these challenges by leveraging algorithms that can learn from data, identify patterns, and make predictions without explicit programming for the interpretation of each specific task. In LIBS application, ML techniques are used to enhance various analytical processes. For example, ML algorithms can classify materials based on their spectral fingerprints, predict the concentration of elements in a sample, and identify underlying patterns within complex datasets. Here, this application improves the precision of LIBS analyses while significantly reducing the time required for data processing and interpretation. In this chapter, the fundamental concepts of ML will be discussed first. Following this, the process of data splitting and the importance of feature selection will be examined. Several machine learning methods will then be closely examined, exploring how each can benefit LIBS analysis and highlighting their respective advantages and shortcomings. This structured approach will provide a comprehensive understanding of the integration of ML in the context of LIBS analysis.

47 OTHER INSTRUMENTATION

Using Machine Learning to Identify Novel Hydroclimate States

Anthropogenic climate change is expected to alter drought risk in the future. However, droughts are not uncommon or unprecedented, as documented in tree-ring-based reconstructions of the summer average Palmer drought severity index (PDSI). Using an unsupervised machine-learning method trained on these reconstructions of pre-industrial climate, we identify outliers: years in which the spatial pattern of PDSI is unusual relative to ‘normal' variability. We show that in many regions, outliers are more frequently identified in the twentieth and twenty-first centuries. This trend is more pronounced when the regional drought atlases are combined into a single global dataset. By definition, outlier patterns at the 10% level are expected to occur once per decade, but from 1950 to 2000 more than 6 years per decade are identified as outliers in the global drought atlas (GDA). Extending the GDA through 2020 using an observational dataset suggests that anomalous global drought conditions are present in 80% of years in the twenty-first century. Our results indicate, without recourse to climate models, that the world is more frequently experiencing drought conditions that are highly unusual in the context of past natural climate variability.

Drought risk

PixelLearn

PixelLearn is an integrated user-interface computer program for classifying pixels in scientific images. Heretofore, training a machine-learning algorithm to classify pixels in images has been tedious and difficult. PixelLearn provides a graphical user interface that makes it faster and more intuitive, leading to more interactive exploration of image data sets. PixelLearn also provides image-enhancement controls to make it easier to see subtle details in images. PixelLearn opens images or sets of images in a variety of common scientific file formats and enables the user to interact with several supervised or unsupervised machine-learning pixel-classifying algorithms while the user continues to browse through the images. The machinelearning algorithms in PixelLearn use advanced clustering and classification methods that enable accuracy much higher than is achievable by most other software previously available for this purpose. PixelLearn is written in portable C++ and runs natively on computers running Linux, Windows, or Mac OS X.

Mazzoni, Dominic

Anomaly detection in collider physics via factorized observables

To maximize the discovery potential of high-energy colliders, experimental searches should be sensitive to unforeseen new physics scenarios. This goal has motivated the use of machine learning for unsupervised anomaly detection. In this paper, we introduce a new anomaly detection strategy called : factorized observables for regressing conditional expectations. Our approach is based on the inductive bias of factorization, which is the idea that the physics governing different energy scales can be treated as approximately independent. Assuming factorization holds separately for signal and background processes, the appearance of nontrivial correlations between low- and high-energy observables is a robust indicator of new physics. Under the most restrictive form of factorization, a machine-learned model trained to identify such correlations will in fact converge to the optimal new physics classifier. We test on a benchmark anomaly detection task for the Large Hadron Collider involving collimated sprays of particles called jets. By teasing out correlations between the kinematics and substructure of jets, our method can reliably extract percent-level signal fractions. This strategy for uncovering new physics adds to the growing toolbox of anomaly detection methods for collider physics with a complementary set of assumptions. Published by the American Physical Society 2024

Astronomy & Astrophysics

Real-time tracking and analysis of gas bubble dynamics in laser powder bed fusion using in-situ X-ray characterization and machine learning

Porosity defects remain a significant challenge in the laser powder bed fusion (LPBF) process, adversely affecting the mechanical properties and reliability of additively manufactured components. Here, this study investigates the real-time formation and trajectory of gas bubbles during LPBF of Al6061 alloy using advanced in-situ X-ray characterization and machine learning. The unsupervised Gaussian mixture model and particle tracking algorithm developed are able to precisely track and quantify the properties of gas bubbles and keyhole pores. Our analysis identified five distinct types of gas bubble formation and movement patterns, emphasizing the diverse origins and behaviors of these defects. It enables precise quantification of trajectories, velocities, and morphological changes of gas bubbles, offering a granular view of the subsurface dynamics within the melt pool. Additionally, we explored keyhole-induced pore dynamics, revealing the critical role of keyhole oscillation and collapse for the formation of both large and small gas pores. It defines four different regions of gas bubble movement within the melt pool, providing a clearer understanding of how local fluid dynamics affect pore behavior. The results underscore the importance of integrating in-situ experimental observation and automated machine learning to develop a more robust predictive model for defect formation in LPBF.

In-situ X-ray imaging

NMF-Based Anomaly Detection in CMS 2D Tracking Occupancy Histograms

The CMS experiment relies on Data Quality Monitoring (DQM) to ensure that recorded collision data are suitable for physics analysis. During LHC Run 3, each run contains many lumisections and tracking monitoring elements, making offline inspection challenging, especially for localized detector effects that may appear only for short periods of time. This poster presents an unsupervised machine-learning approach to identify anomalous lumisections in CMS tracking occupancy histograms using Non-Negative Matrix Factorization (NMF). The workflow uses offline CMS DQMIO tracking histograms retrieved with the CMS DIALS API and organized as two-dimensional occupancy maps for each lumisection. After selecting stable lumisections, the occupancy maps are normalized and arranged into a non-negative data matrix. The NMF model learns a compact set of basis patterns describing normal tracking occupancy. Each lumisection is then reconstructed from these learned components, and the reconstruction error is used as an anomaly score. Large residuals indicate occupancy patterns that deviate from normal detector behavior and are flagged for further inspection. This NMF-based approach provides a fast and interpretable way to flag lumisections whose tracking occupancy patterns differ from normal detector behavior. Preliminary studies show sensitivity to known tracking anomalies, and ongoing work is focused on validating the method across additional Run 3 Pixel and Strip detector issues.

Rodríguez Ramos, Iliomar [Puerto Rico U., Mayaguez

Can Selforganizing Maps Accurately Predict Photometric Redshifts?

We present an unsupervised machine-learning approach that can be employed for estimating photometric redshifts. The proposed method is based on a vector quantization called the self-organizing-map (SOM) approach. A variety of photometrically derived input values were utilized from the Sloan Digital Sky Survey's main galaxy sample, luminous red galaxy, and quasar samples, along with the PHAT0 data set from the Photo-z Accuracy Testing project. Regression results obtained with this new approach were evaluated in terms of root-mean-square error (RMSE) to estimate the accuracy of the photometric redshift estimates. The results demonstrate competitive RMSE and outlier percentages when compared with several other popular approaches, such as artificial neural networks and Gaussian process regression. SOM RMSE results (using delta(z) = z(sub phot) - z(sub spec)) are 0.023 for the main galaxy sample, 0.027 for the luminous red galaxy sample, 0.418 for quasars, and 0.022 for PHAT0 synthetic data. The results demonstrate that there are nonunique solutions for estimating SOM RMSEs. Further research is needed in order to find more robust estimation techniques using SOMs, but the results herein are a positive indication of their capabilities when compared with other well-known methods

Way, Michael J.

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma

Novel CHI3L1 ‐Associated Angiogenic Phenotypes Define Glioma Microenvironments: Insights From Multi‐Omics Integration

ABSTRACT The CHI3L1 signaling pathway significantly influences glioma angiogenesis, but its role in the tumor microenvironment (TME) remains elusive. We propose a novelCHI3L1‐associated vascular phenotype classification for glioma through integrative analyses of multiple datasets with bulk and single‐cell transcriptome, genomics, digital pathology, and clinical data. We investigated the biological characteristics, genomic alterations, therapeutic vulnerabilities, and immune profiles within these phenotypes through a comprehensive multi‐omics approach. We constructed the vascular‐related risk (VR) score based onCHI3L1‐associated vascular signatures (CAVS) identified by machine learning algorithms. Utilizing unsupervised consensus clustering, gliomas were stratified into three distinct vascular phenotypes: Cluster A, marked by high vascularization and stromal activation with a relatively low levels of tumor‐infiltrating lymphocytes (TILs); Cluster B, characterized by moderate vascularization and stromal activity, coupled with a high density of TILs; and Cluster C, defined by low vascularization and sparse immune cell infiltration. We observed that the CAVS effectively indicated glioma‐associated angiogenesis and immune suppression by single‐cell RNA‐seq analysis. Moreover, the high‐VR‐score group exhibited enhanced angiogenic activity, reduced immune response, resistance to immunotherapy, and poorer clinical outcomes. The VR score independently predicted glioma prognosis and, combined with a nomogram, provided a robust clinical decision‐making tool. Potential drug prediction based on transcription factors for high‐risk patients was also performed. Our study reveals thatCHI3L1‐associated vascular phenotypes shape distinct immune landscapes in gliomas, offering insights for optimizing therapeutic strategies to improve patient outcomes.

Oncology

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science

Unsupervised Clustering and Supervised Regression Learning to Select High Temperature Oxidation-Resistant Materials

High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.

Romedenne, Marie [ORNL] (ORCID:0000000317936561)

Unsupervised domain adaptation for radioisotope identification in gamma spectroscopy

Training machine learning models for radioisotope identification using gamma spectroscopy remains an elusive challenge for many practical applications, largely stemming from the difficulty of acquiring and labeling large, diverse experimental datasets. Simulations can mitigate this challenge, but the accuracy of models trained on simulated data can deteriorate substantially when deployed to an out-of-distribution operational environment. In this study, we demonstrate that unsupervised domain adaptation (UDA) can improve the ability of a model trained on synthetic data to generalize to a new testing domain, provided unlabeled data from the target domain are available. Conventional supervised techniques are unable to utilize this data because the absence of isotope labels precludes defining a supervised classification loss. Instead, we first pretrain a spectral classifier using labeled synthetic data and subsequently leverage unlabeled target data to align the learned feature representations between the source and target domains. We compare a range of different UDA techniques, finding that minimizing the maximum mean discrepancy (MMD) between source and target feature vectors yields the most consistent improvement to testing scores. For instance, using a custom transformer-based neural network, we achieved a testing accuracy of $0.904 \pm 0.022$ on an experimental LaBr test set after performing unsupervised feature alignment via MMD minimization, compared to $0.754 \pm 0.014$ before alignment. Overall, our results highlight the potential of using UDA to adapt a radioisotope classifier trained on synthetic data for real-world deployment.

Lalor, Peter W.