Search NASASearch

SEARCH · Search NASA

Results for “Gaussian mixture model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Galaxy cluster profiles: a Gaussian mixture model approach to halo miscentering

Measurements of the galaxy density and weak-lensing profiles of galaxy clusters typically rely on an assumed cluster center, which is taken to be the brightest cluster galaxy or other proxies for the true halo center defined as the minimum in the potential well. Departure of the assumed cluster center from the true halo center bias the resultant profile measurements, an effect known as miscentering bias. Currently, miscentering is typically modeled in stacked profiles of clusters with a two parameter model. We use an alternate approach in which the profiles of individual clusters are used with the corresponding likelihood computed using a Gaussian mixture model. We test the approach using halos and the corresponding subhalo profiles from the IllustrisTNG hydrodynamic simulations. We obtain significantly improved estimates of the miscentering parameters for both 3D and projected 2D profiles relevant for imaging surveys. We discuss applications to upcoming cosmological surveys. Our Python package for the Gaussian mixture model is publicly available at https://github.com/KyleMiller1/Halo-Miscentering-Mixture-Model.

Bayesian reasoning

Building molecular model series from heterogeneous CryoEM structures using Gaussian mixture models and deep neural networks

Cryogenic electron microscopy (CryoEM) produces structures of macromolecules at near-atomic resolution. However, building molecular models with good stereochemical geometry from those structures can be challenging and time-consuming, especially when many structures are obtained from datasets with conformational heterogeneity. Here we present a model refinement protocol that automatically generates series of molecular models from CryoEM datasets, which describe the dynamics of the macromolecular system and have near-perfect geometry scores. This method makes it easier to interpret the movement of the protein complex from heterogeneity analysis and to compare the structural dynamics observed from CryoEM data with results from other experimental and simulation techniques.

59 BASIC BIOLOGICAL SCIENCES

Leveraging Gaussian Mixture Models for Detecting Anomalies in Time-Series Data

Test systems must be capable of classifying measured data as expected or anomalous in real time. Anomalous results may portend system failure, and, if undetected, may result in damage to the unit, test equipment, or potential harm to personnel. This report investigates the use of Gaussian Mixture Models (GMMs) as a clustering tool in classifying time-series data.

Wilke, Rudeger H.T. [Sandia National Laboratories

G-Mapper: Learning a Cover in the Mapper Construction

The Mapper algorithm is a visualization technique in topological data analysis (TDA) that outputs a graph reflecting the structure of a given dataset. However, the Mapper algorithm requires tuning several parameters in order to generate a “nice” Mapper graph. This paper focuses on selecting the cover parameter. We present an algorithm that optimizes the cover of a Mapper graph by splitting a cover repeatedly according to a statistical test for normality. Our algorithm is based on G-means clustering, which searches for the optimal number of clusters in 𝑘-means by iteratively applying the Anderson–Darling test. Our splitting procedure employs a Gaussian mixture model to carefully choose the cover according to the distribution of the given data. In conclusion, experiments for synthetic and real-world datasets demonstrate that our algorithm generates covers so that the Mapper graphs retain the essence of the datasets, while also running significantly faster than a previous iterative method.

G-means clustering

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE

A comparison of probabilistic generative frameworks for molecular simulations

Generative artificial intelligence is now a widely used tool in molecular science. Despite the popularity of probabilistic generative models, numerical experiments benchmarking their performance on molecular data are lacking. Here, in this work, we introduce and explain several classes of generative models, broadly sorted into two categories: flow-based models and diffusion models. We select three representative models: neural spline flows, conditional flow matching, and denoising diffusion probabilistic models, and examine their accuracy, computational cost, and generation speed across datasets with tunable dimensionality, complexity, and modal asymmetry. Our findings are varied, with no one framework being the best for all purposes. In a nutshell, (i) neural spline flows do best at capturing mode asymmetry present in low-dimensional data, (ii) conditional flow matching outperforms other models for high-dimensional data with low complexity, and (iii) denoising diffusion probabilistic models appear the best for low-dimensional data with high complexity. Our datasets include a Gaussian mixture model and the dihedral torsion angle distribution of the Aib9 peptide, generated via a molecular dynamics simulation. We hope our taxonomy of probabilistic generative frameworks and numerical results may guide model selection for a wide range of molecular tasks.

Artificial intelligence

Constraints on White Dwarf Hydrogen Layer Masses Using Gravitational Redshifts

The hydrogen envelope is the outermost layer of a DA white dwarf; it makes up the entirety of the stellar photosphere, and yet its typical extent is difficult to model theoretically and remains poorly observationally constrained. As a result, hydrogen envelope mass is a substantial source of systematic uncertainty in the physical properties of white dwarfs, including overall masses and cooling ages. In this work, we fit a Gaussian mixture model to gravitational redshifts from high-resolution spectroscopy, paired with radius measurements from Gaia BP/RP spectra, to measure the mass–radius relation for a sample of 468 white dwarfs. Our results are in excellent agreement with the predicted mass–radius relations of state-of-the-art evolutionary models, including those from the MESA Isochrones and Stellar Tracks (MIST) library. We find that mass–radius relations such as those from MIST that assume a thick and mass-dependent hydrogen envelope are preferred by the observed probability density function over models that assume a hydrogen envelope of constant mass. Proper treatment of the evolution of white dwarf progenitors is thus important for accurately modeling the mass–radius relation. Our results indicate that gravitational redshift measurements of large samples of white dwarfs in wide binaries are promising probes of the hydrogen envelope masses of DA white dwarfs.

Astronomy and AstroPhysics

Real-time tracking and analysis of gas bubble dynamics in laser powder bed fusion using in-situ X-ray characterization and machine learning

Porosity defects remain a significant challenge in the laser powder bed fusion (LPBF) process, adversely affecting the mechanical properties and reliability of additively manufactured components. Here, this study investigates the real-time formation and trajectory of gas bubbles during LPBF of Al6061 alloy using advanced in-situ X-ray characterization and machine learning. The unsupervised Gaussian mixture model and particle tracking algorithm developed are able to precisely track and quantify the properties of gas bubbles and keyhole pores. Our analysis identified five distinct types of gas bubble formation and movement patterns, emphasizing the diverse origins and behaviors of these defects. It enables precise quantification of trajectories, velocities, and morphological changes of gas bubbles, offering a granular view of the subsurface dynamics within the melt pool. Additionally, we explored keyhole-induced pore dynamics, revealing the critical role of keyhole oscillation and collapse for the formation of both large and small gas pores. It defines four different regions of gas bubble movement within the melt pool, providing a clearer understanding of how local fluid dynamics affect pore behavior. The results underscore the importance of integrating in-situ experimental observation and automated machine learning to develop a more robust predictive model for defect formation in LPBF.

In-situ X-ray imaging

Dominant balance-based adaptive mesh refinement for incompressible fluid flows

This work introduces a novel adaptive mesh refinement (AMR) method that utilizes dominant balance analysis (DBA) for efficient and accurate grid adaptation in computational fluid dynamics (CFD) simulations. The proposed method leverages a Gaussian mixture model (GMM) to classify grid cells into active and passive regions based on the dominant physical interactions within the equation space. By modeling truncation error probabilistically from discretized terms, the method identifies regions of high interaction where numerical accuracy is most sensitive to resolution. Unlike traditional AMR strategies, this approach does not rely on heuristic-based sensors or user-defined thresholds, providing a fully automated and problem-independent framework for AMR. Applied to the incompressible Navier-Stokes equations for steady and unsteady flow past a cylinder, the DBA-based AMR method achieves comparable accuracy to high-resolution grids while reducing computational costs by up to 70 %. The validation highlights the method’s effectiveness in capturing complex flow features while minimizing grid cells, directing computational resources toward regions with the most critical dynamics. This modular and scalable strategy is adaptable to a wide range of applications, presenting a promising tool for efficient high-fidelity simulations in CFD and other multiphysics domains.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

End-to-End Automated Segmentation Framework for Four-Dimensional Scanning Transmission Electron Microscopy Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) is powerful for rapidly characterizing arrays of nanoparticles produced via high-throughput synthesis. However, such 4D-STEM datasets typically contain thousands of nanoparticles, each characterized by thousands of diffraction patterns spatially distributed across the nanoparticle, necessitating efficient and comprehensive analysis. We propose an end-to-end segmentation framework to automatically segment each nanoparticle into regions with distinct composition/orientation of crystal grains, using only the 4D-STEM data. Bragg disk information is extracted in a physics-informed manner from the diffraction patterns at each spatial location and combined with the real space coordinates to form feature vectors. These feature vectors are then used as inputs to a Gaussian mixture model (GMM) to segment the nanoparticle into distinct regions. We also develop two visualization tools based on the GMM outputs to infer the interface transition and the degree of superposition. Our framework comprehensively integrates machine learning tools and physics knowledge, and provides a basis for substantially compressing enormous 4D-STEM datasets, e.g., by replacing the full 4D-STEM dataset for each nanoparticle with only a single set of Bragg disk features for each distinct crystal grain identified in the nanoparticle. In this article, we demonstrate the power of our framework by presenting results for real, complex datasets.

47 OTHER INSTRUMENTATION

Artificial-intelligence-assisted analysis of 28 Si * → 7⁢𝛼 breakup data

Mid-weight 𝛼-conjugate nuclei are predicted to possess exotic toroid like resonances with high angular momenta. The search for these states in 28 Si* is the main point of two published experimental investigations of the peripheral 28 Si + 12 C reaction by Cao and collaborators and by Hannaman and collaborators. In this work, we develop a novel artificial intelligence (AI) based machine learning method utilizing the Gaussian Mixture Model (GMM) to analyze available experimental and theoretical data. Here, we additionally study the reaction with the Hybrid 𝛼-Cluster (H⁡𝛼⁢C) model. In all the examined data, our results suggest the presence of underlying structure which is close to that predicted for toroidal states.

Breakup reactions

A Comparison of Machine Learning Methods of Association Tested on Dense Nodal Arrays

The association of phase picks to form events is one of the fundamental components of seismology. Large and dense sensor networks, such as >1000 geophone arrays (and distributed acoustic sensing), offer unique challenges in association due to the vast numbers of observations and high likelihood of errant picks. In addition, the large number of stations can greatly increase the time it takes to perform the association. For this reason, machine learning (ML) methods might provide a more optimal method of association for such networks. In this work, we examine how well ML methods (e.g., Gaussian mixture model association, PhaseLink, and Graph Earthquake Neural Interpretation Engine) can incorporate dense seismic arrays into regional networks and how well they handle the increasing numbers of stations. Here, we test their capabilities on two dense seismic deployments, one within Rock Valley Nevada (52 nodes and a 9-station sparse local network), and the LArge-n Seismic Survey in Oklahoma dense nodal array (>1800 vertical-component geophones). Processing data from these two different styles of dense seismic deployments allows testing of how the ML algorithms can merge array data with a broader regional network, how they deal with poorly picked phases, and how they handle anthropogenic noise. We compare the ML-associated bulletins to those obtained using the Rapid Earthquake Association and Location algorithm, a more traditional method of association. We find that there are very small differences in results between the methods for small networks (<100 stations) with low pick rates. For large networks (>1000), there are enough errant picks that some of the ML methods start to create false events out of noise. We also find that the ML methods vary in computation time significantly but are all faster than the traditional method tested here.

58 GEOSCIENCES

FNCL Enhancements Implementation (FY25 Annual Report)

The FNCL investigations team at Lawrence Livermore National Laboratory (LLNL) has completed research and development of hardware, signal processing, and analysis tools to enhance the measurement capabilities of both the current CAEN SyS VeryFuel Fast Neutron Collar (FNCL) instrument and a next-generation FNCL prototype. The team successfully built and commissioned the LLNL Demonstrator System: a fully integrated, three-panel detector system featuring higher segmentation, plastic scintillators (EJ-276D), Silicon Photomultipliers (SiPMs), no high-voltage requirement, a reduced electronic footprint, and the LLNL-developed Gaussian Mixture Model Pulse Shape Discrimination (GMM-PSD) signal processing. An extensive experimental campaign was conducted at LLNL’s Inherently Safe Subcritical Assembly (ISSA) facility using both the baseline FNCL and the LLNL Demonstrator. The campaign results validated system performance, calibration stability, and the effectiveness of advanced signal processing and analysis algorithms in a relevant environment.

and physical protection

Entropy-Assisted Quality Pattern Identification in Finance

Short-term patterns in financial time series form the cornerstone of many algorithmic trading strategies, yet extracting these patterns reliably from noisy market data remains a formidable challenge. In this paper, we propose an entropy-assisted framework for identifying high-quality, non-overlapping patterns that exhibit consistent behavior over time. We ground our approach in the premise that historical patterns, when accurately clustered and pruned, can yield substantial predictive power for short-term price movements. To achieve this, we incorporate an entropy-based measure as a proxy for information gain: patterns that lead to high one-sided movements in historical data yet retain low local entropy are more “informative” in signaling future market direction. Compared to conventional clustering techniques such as K-means and Gaussian Mixture Models (GMMs), which often yield biased or unbalanced groupings, our approach emphasizes balance over a forced visual boundary, ensuring that quality patterns are not lost due to over-segmentation. By emphasizing both predictive purity (low local entropy) and historical profitability, our method achieves a balanced representation of Buy and Sell patterns, making it better suited for short-term algorithmic trading strategies. This paper offers an in-depth illustration of our entropy-assisted framework through two case studies on Gold vs. USD and GBPUSD. While these examples demonstrate the method’s potential for extracting high-quality patterns, they do not constitute an exhaustive survey of all possible asset classes.

Physics

SpecDis: Value Added Distance Catalog for 4 Million Stars from DESI Year-1 Data

We present the SpecDis value-added stellar distance catalog accompanying DESI Data Release 1. SpecDis trains a feed-forward neural network (NN) with Gaia parallaxes and gets the distance estimates. To build up an unbiased training sample, we do not apply selections on parallax error or signal-to-noise (S/N) of the stellar spectra, and instead, we incorporate parallax error into the loss function. Moreover, we employ principal component analysis to reduce the noise and dimensionality of stellar spectra. Validated by independent external samples of member stars with precise distances from globular clusters, dwarf galaxies, stellar streams, combined with blue horizontal branch stars, we demonstrate that our distance measurements show no significant bias up to 100 kpc, and are much more precise than Gaia parallax beyond 7 kpc. The median distance uncertainties are 23%, 19%, 11%, and 7% for S/N < 20, 20 ≤ S/N < 60, 60 ≤ S/N < 100, and S/N ≥ 100. Selecting stars with ${\mathrm{log}}\,g\lt 3.8$ and distance uncertainties smaller than 25%, we have more than 74,000 giant candidates within 50 kpc of the Galactic center and 1500 candidates beyond this distance. Additionally, we develop a Gaussian mixture model to identify unresolvable equal-mass binaries by modeling the discrepancy between the NN-predicted and the geometric absolute magnitudes from Gaia parallaxes and identify 120,000 equal-mass binary candidates. Our final catalog provides distances and distance uncertainties for >4 million stars, offering a valuable resource for Galactic astronomy.

astronomy data analysis

New Measurements of the Lyα Forest Continuum and Effective Optical Depth with LyCAN and DESI Y1 Data

Abstract We present the Ly α Continuum Analysis Network (LyCAN), a convolutional neural network that predicts the unabsorbed quasar continuum within the rest-frame wavelength range of 1040–1600 Å based on the red side of the Ly α emission line (1216–1600 Å). We developed synthetic spectra based on a Gaussian mixture model representation of nonnegative matrix factorization (NMF) coefficients. These coefficients were derived from high-resolution, low-redshift ( z < 0.2) Hubble Space Telescope/Cosmic Origins Spectrograph (COS) quasar spectra. We supplemented this COS-based synthetic sample with an equal number of DESI Year 5 mock spectra. LyCAN performs extremely well on testing sets, achieving a median error in the forest region of 1.5% on the DESI mock sample, 2.0% on the COS-based synthetic sample, and 4.1% on the original COS spectra. LyCAN outperforms principal component analysis (PCA) and NMF-based prediction methods using the same training set by 40% or more. We predict the intrinsic continua of 83,635 DESI Year 1 spectra in the redshift range of 2.1 ≤ z ≤ 4.2 and perform an absolute measurement of the evolution of the effective optical depth. This is the largest sample employed to measure the optical depth evolution to date. We fit a power law of the form τ ( z ) = τ 0 ( 1 + z ) γ to our measurements and find τ 0 = (2.46 ± 0.14) × 10 −3 and γ = 3.62 ± 0.04. Our results show particular agreement with high-resolution, ground-based observations around z = 2, indicating that LyCAN is able to predict the quasar continuum in the forest region with only spectral information outside the forest.

79 ASTRONOMY AND ASTROPHYSICS

The Velocity Field of Our Milky Way Outer Stellar Halo Based on DESI DR2

Using 64,000 halo K giants from the Dark Energy Spectroscopic Instrument second Data Release (DR2), we decompose the Milky Way (MW) stellar halo between 3 and 160 kpc into metal-rich (MR) and metal-poor (MP) components via a Gaussian mixture model. The two populations are nearly equal in number but chemically and kinematically distinct: MR stars occupy highly radial orbits with velocity anisotropy of β ≈ 0.94 and metallicity dispersion σ [Fe/H] ≈ 0.17 dex, without obvious dependence on distance, and are mainly contributed by Gaia-Sausage/Enceladus (GSE) debris. The MR component dominates the inner 30 kpc and reemerges beyond 50 kpc, implying GSE debris can extend to ∼70–80 kpc. MP stars exhibit a weaker radial bias of β ≈ 0.46, decreasing to −0.5 beyond 80 kpc, and with a larger metallicity dispersion of σ [Fe/H] ≈ 0.46 dex, showing signatures of multiple minor mergers. Both components exhibit net prograde rotation at ∼10–30 kpc with a stronger azimuthal signal in the MP population. The nonequilibrium motions of the outer halo (>50 kpc) are quantified with a dipole-plus-contraction velocity field. We find that the outer halo is simultaneously contracting (μ compr = −19 km s −1 , distance-independent) and subject to reflex motions (μ dipole increases from −19 to −44 km s −1 with radius), reflecting the perturbation from the Large Magellanic Cloud (LMC). We also confirm a linear dependence of mean polar velocity for the outer stellar halo on μ dipole , a direct consequence of the LMC and MW interaction. Our results provide a quantitative distance-resolved description of the MW’s last major accretion event and its ongoing response to the first infall of the LMC.

Li, Songting [Shanghai Jiao Tong University; Shang