Search NASA⌕ Search

SEARCH · Search NASA

Results for “Mixture modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

InClass nets: independent classifier networks for nonparametric estimation of conditional independence mixture models and unsupervised classification

Abstract Conditional independence mixture models (CIMMs) are an important class of statistical models used in many fields of science. We introduce a novel unsupervised machine learning technique called the independent classifier networks (InClass nets) technique for the nonparameteric estimation of CIMMs. InClass nets consist of multiple independent classifier neural networks (NNs), which are trained simultaneously using suitable cost functions. Leveraging the ability of NNs to handle high-dimensional data, the conditionally independent variates of the model are allowed to be individually high-dimensional, which is the main advantage of the proposed technique over existing non-machine-learning-based approaches. Two new theorems on the nonparametric identifiability of bivariate CIMMs are derived in the form of a necessary and a (different) sufficient condition for a bivariate CIMM to be identifiable. We use the InClass nets technique to perform CIMM estimation successfully for several examples. We provide a public implementation as a Python package called RainDancesVI.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Likelihood Maximization and Moment Matching in Low SNR Gaussian Mixture Models

We derive an asymptotic expansion for the log-likelihood of Gaussian mixture models (GMMs) with equal covariance matrices in the low signal-to-noise regime. The expansion reveals an intimate connection between two types of algorithms for parameter estimation: the method of moments and likelihood optimizing algorithms such as Expectation-Maximization (EM). We show that likelihood optimization in the low SNR regime reduces to a sequence of least squares optimization problems that match the moments of the estimate to the ground truth moments one by one. This connection is a stepping stone towards the analysis of EM and maximum likelihood estimation in a wide range of models. A motivating application for the study of low SNR mixture models is cryo-electron microscopy data, which can be modeled as a GMM with algebraic constraints imposed on the mixture centers. We discuss the application of our expansion to algebraically constrained GMMs, among other example models of interest. © 2022 The Authors. Communications on Pure and Applied Mathematics published by Wiley Periodicals LLC.

97 MATHEMATICS AND COMPUTING↗

Predicting Damages to Remainder Parcels in Right-of-Way Acquisitions for Expanding Transportation Infrastructure: Using a Truncated Finite-Mixture Model

Right-of-way acquisition is a critical component of transportation infrastructure development. Transportation infrastructure projects cannot proceed without proper right-of-way acquisition or may face significant delays. State Departments of Transportation frequently acquire parcels of land for roadway expansion projects. A majority of these acquisitions can be partial takings, referring to a portion of a parcel that is acquired. The remainder of the property usually suffers economic changes due to the partial acquisition, which can be calculated as damage percentages. The damage percentage represents the extent to which the remaining land or property value has been diminished due to the acquisition. It reflects the remaining property value percentage that may have been lost or compromised due to the acquisition. Here, this study aims to provide a robust model to estimate damage percentages to the remainder parcels that may help state Departments of Transportation appraisers make early predictions about the damages in cases involving partial takings. The research uses 509 appraisal reports from the Tennessee Department of Transportation to identify the key parcel attributes that influence the percentage of damages. Three regression models are developed: a linear regression model, a finite-mixture model (FMM), and a truncated FMM with two latent classes. The modeling results show that the truncated FMM with two classes outperforms the other models. To validate the models, actual sales data is collected and analyzed for 59 properties, and the results suggest that the model predictions are fairly accurate. A predictive tool is developed based on the models to help appraisers anticipate right-of-way damages under different scenarios and can provide early predictions about the damages.

42 ENGINEERING↗

Enter Gaussian Mixture Modeling Extensions for Improved False Discovery Rate Estimation in GC-MS Metabolomics

Identifying small molecules (e.g., metabolites) is key towards driving scientific advancement in metabolomics, and gas chromatography–mass spectrometry (GC-MS) is an analytic method that may be applied to facilitate this process. The typical GC-MS identification workflow involves quantifying the similarity of an observed sample spectrum and other features (e.g. retention index) to that of several references, noting the compound of the best-matching reference spectrum as the identified metabolite. While a deluge of similarity metrics exists, none characterize the error rate of generated identifications, thereby presenting an unknown risk of false identification or discovery. To quantify this unknown risk, we propose a model-based framework for estimating the false discovery rate (FDR) among a set of identifications. Extending the traditional mixture modeling framework, our method incorporates both similarity score and experimental information in estimating the FDR. We apply these models to identification lists derived from across 548 samples of varying complexity and sample type (e.g., fungal species, standard mixtures, etc.), comparing their performance to that of the traditional Gaussian mixture model (GMM). Through simulation, we additionally assess the impact of reference library size on the accuracy of FDR estimates. In comparing the best performing model extensions to the GMM, our results indicate relative decreases in median absolute estimation error (MAE) ranging from 12% to 70%, based on comparisons of the median MAEs across all hit-lists. Results indicate that these relative performance improvements generally hold despite library size, however FDR estimation error typically worsens as the set of reference compounds diminishes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Continuum shock mixture models for Ni+Al multilayers: Individual layers and bulk equations of state

Continuum shock mixture models are reviewed and applied to determine the equations of state for five different compositions of Ni x Al y ⁠, as well as bulk Ni+Al reactive multilayers, by combining the fundamental property data for elemental nickel and aluminum. From the literature, we down-select and evaluate two analytical models for the mixture Hugoniot, i.e., the well-known method of kinetic energy averaging (KEA) and a recent model proposed by Jordan and Baer [J. Appl. Phys. 111, 083516 (2012)]. Fundamentally, the former method assumes pressure equilibrium, whereas the latter assumes a common particle velocity and mixture sound speed from compressible two-phase cavitating flows. Additionally, we construct thermodynamically complete equations of state by fitting Einstein oscillator series models for the specific heat at constant volume. Finally, the solid solution approximation is invoked for intermetallic compositions, which are not strictly physical mixtures. Overall, the KEA model provides a better fit to the available Ni x Al y and Ni+Al multilayer shock compression data; however, there are combinations of material properties where the performance of these two models is thought to be reversed. Moreover, the results of this work include the first analytical solution of Jordan–Baer that does not require numerical root finding, as well as proposed modifications to the Einstein oscillator series to incorporate some effects of local pressure–temperature equilibrium and reaction–diffusion. Future work is planned that will use these equations of state in mesoscale simulations to study shock-induced reaction in Ni+Al multilayers, and the intended application is illustrated with a brief 2D hydrocode example.

36 MATERIALS SCIENCE↗

Heterogeneous Data Type Mixture Model

This code performs estimation of a mixture model involving heterogeneous data types. The resulting model can be used to calculate likelihoods of underlying model components.

Kaplan, AlanD↗

Earthquake Phase Association Using a Bayesian Gaussian Mixture Model

Earthquake phase association algorithms aggregate picked seismic phases from a network of seismometers into individual seismic events and play an important role in earthquake monitoring and research. Dense seismic networks and improved phase picking methods produce massive seismic phase datasets, particularly for earthquake swarms and aftershocks occurring closely in time and space, making phase association a challenging problem. Here, we present a new association method, the Gaussian Mixture Model Association (GaMMA), that combines the Gaussian mixture model with earthquake location, origin time, and magnitude estimation. We treat earthquake phase association as an unsupervised clustering problem in a probabilistic framework, where each earthquake corresponds to a cluster of P and S phases with a hyperbolic moveout of arrival times and a decay of amplitude with distance. We use the multivariate Gaussian distribution to model the collection of phase picks of an event; and the mean of the multivariate Gaussian distribution is given by the predicted arrival time and amplitude from the causative event. We carry out the pick assignment to each earthquake and determine earthquake source parameters (i.e., earthquake location, origin time, and magnitude) under the maximum likelihood criterion using the Expectation-Maximization algorithm. The GaMMA method does not require typical association steps of other algorithms, such as grid-search or supervised training. The results for both synthetic tests and for the 2019 Ridgecrest earthquake sequence show that GaMMA effectively associates phases from a temporally and spatially dense earthquake sequence while producing useful estimates of earthquake location and magnitude.

58 GEOSCIENCES↗

Anomaly Detection in Connected and Autonomous Vehicle Trajectories Using LSTM Autoencoder and Gaussian Mixture Model

Connected and Autonomous Vehicles (CAVs) technology has the potential to transform the transportation system. Although these new technologies have many advantages, the implementation raises significant concerns regarding safety, security, and privacy. Anomalies in sensor data caused by errors or cyberattacks can cause severe accidents. To address the issue, this study proposed an innovative anomaly detection algorithm, namely the LSTM Autoencoder with Gaussian Mixture Model (LAGMM). This model supports anomalous CAV trajectory detection in the real-time leveraging communication capabilities of CAV sensors. The LSTM Autoencoder is applied to generate low-rank representations and reconstruct errors for each input data point, while the Gaussian Mixture Model (GMM) is employed for its strength in density estimation. The proposed model was jointly optimized for the LSTM Autoencoder and GMM simultaneously. The study utilizes realistic CAV data from a platooning experiment conducted for Cooperative Automated Research Mobility Applications (CARMAs). The experiment findings indicate that the proposed LAGMM approach enhances detection accuracy by 3% and precision by 6.4% compared to the existing state-of-the-art methods, suggesting a significant improvement in the field.

33 ADVANCED PROPULSION SYSTEMS↗

Galaxy cluster profiles: a Gaussian mixture model approach to halo miscentering

Measurements of the galaxy density and weak-lensing profiles of galaxy clusters typically rely on an assumed cluster center, which is taken to be the brightest cluster galaxy or other proxies for the true halo center defined as the minimum in the potential well. Departure of the assumed cluster center from the true halo center bias the resultant profile measurements, an effect known as miscentering bias. Currently, miscentering is typically modeled in stacked profiles of clusters with a two parameter model. We use an alternate approach in which the profiles of individual clusters are used with the corresponding likelihood computed using a Gaussian mixture model. We test the approach using halos and the corresponding subhalo profiles from the IllustrisTNG hydrodynamic simulations. We obtain significantly improved estimates of the miscentering parameters for both 3D and projected 2D profiles relevant for imaging surveys. We discuss applications to upcoming cosmological surveys. Our Python package for the Gaussian mixture model is publicly available at https://github.com/KyleMiller1/Halo-Miscentering-Mixture-Model.

Bayesian reasoning↗

Automated integration gate selection for Gaussian mixture model pulse shape discrimination

Pulse shapes differ between neutron and gamma particles when measured with detector devices employing pulse shape discriminating (PSD) scintillators. Digitized waveforms can be used in detection systems to perform pulse shape discrimination for this application. Prior Gaussian Mixture Model (GMM) methods require access to the pulse full-waveform. Reducing the waveform to a smaller set of combined samples reduces computational cost while affecting PSD performance. In this work, we develop a method for selecting the best performing combination of integration gates, or contiguous summed segments of the digitized pulse for PSD. The method uses a discrimination score based on the GMM PSD approach. Furthermore, the final selection is performed using Bayesian Optimization. PSD detection results are compared with varying numbers of selected gates on time-of-flight (TOF) data. This method can be used to fully automate the selection of gates in an unsupervised (without ground truth labels) setting.

42 ENGINEERING↗

Lockdown impacts on residential electricity demand in India: A data-driven and non-intrusive load monitoring study using Gaussian mixture models

This study evaluates the effect of complete nationwide lockdown in 2020 on residential electricity demand across 13 Indian cities and the role of digitalisation using a public smart meter dataset. We undertake a data-driven approach to explore the energy impacts of work-from-home norms across five dwelling typologies. Our methodology includes climate correction, dimensionality reduction and machine learning-based clustering using Gaussian Mixture Models of daily load curves. Results show that during the lockdown, maximum daily peak demand increased by 150-200% as compared to 2018 and 2019 levels for one room-units (RM1), one bedroom-units (BR1) and two bedroom-units (BR2) which are typical for low- and middle-income families. While the upper-middle- and higher-income dwelling units (i.e., three (3BR) and more-than-three bedroom-units (M3BR)) saw night-time demand rise by almost 44% in 2020, as compared to 2018 and 2019 levels. Our results also showed that new peak demand emerged for the lockdown period for RM1, BR1 and BR2 dwelling typologies. We found that the lack of supporting socioeconomic and climatic data can restrict a comprehensive analysis of demand shocks using similar public datasets, which informed policy implications for India's digitalisation. We further emphasised improving the data quality and reliability for effective data-centric policymaking.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Building molecular model series from heterogeneous CryoEM structures using Gaussian mixture models and deep neural networks

Cryogenic electron microscopy (CryoEM) produces structures of macromolecules at near-atomic resolution. However, building molecular models with good stereochemical geometry from those structures can be challenging and time-consuming, especially when many structures are obtained from datasets with conformational heterogeneity. Here we present a model refinement protocol that automatically generates series of molecular models from CryoEM datasets, which describe the dynamics of the macromolecular system and have near-perfect geometry scores. This method makes it easier to interpret the movement of the protein complex from heterogeneity analysis and to compare the structural dynamics observed from CryoEM data with results from other experimental and simulation techniques.

59 BASIC BIOLOGICAL SCIENCES↗

Improving resolution and resolvability of single-particle cryoEM structures using Gaussian mixture models

Cryogenic electron microscopy is widely used in structural biology, but its resolution is often limited by the dynamics of the macromolecule. Here, we developed a refinement protocol based on Gaussian mixture models that integrates particle orientation and conformation estimation, and improves the alignment for flexible domains of protein structures. We demonstrated this protocol on multiple datasets, resulting in improved resolution and resolvability, locally and globally, by visual and quantitative measures.

59 BASIC BIOLOGICAL SCIENCES↗

Red Dragon: a redshift-evolving Gaussian mixture model for galaxies

ABSTRACT Precision-era optical cluster cosmology calls for a precise definition of the red sequence (RS), consistent across redshift. To this end, we present the Red Dragon algorithm: an error-corrected multivariate Gaussian mixture model (GMM). Simultaneous use of multiple colours and smooth evolution of GMM parameters result in a continuous RS and blue cloud (BC) characterization across redshift, avoiding the discontinuities of red fraction inherent in swapping RS selection colours. Based on a mid-redshift spectroscopic sample of SDSS galaxies, an RS defined by Red Dragon selects quiescent galaxies (low specific star formation rate) with a balanced accuracy of over $90{{\ \rm per\ cent}}$. This approach to galaxy population assignment gives more natural separations between RS and BC galaxies than hard cuts in colour–magnitude or colour–colour spaces. The Red Dragon algorithm is publicly available at bitbucket.org/wkblack/red-dragon-gamma/.

79 ASTRONOMY AND ASTROPHYSICS↗

Bayesian mixture model approach to quantifying the empirical nuclear saturation point

The equation of state (EOS) in the limit of infinite symmetric nuclear matter exhibits an equilibrium density, $n_0 \approx 0.16 \, \mathrm{fm}^{-3}$, at which the pressure vanishes and the energy per particle attains its minimum, $E_0 \approx -16 \, \mathrm{MeV}$. Although not directly measurable, the nuclear saturation point $(n_0,E_0)$ can be extrapolated by density functional theory (DFT), providing tight constraints for microscopic interactions derived from chiral effective field theory (EFT). However, when considering several DFT predictions for $(n_0,E_0)$ from Skyrme and Relativistic Mean Field (RMF) models together, a discrepancy between these model classes emerges at high confidence levels that each model prediction's uncertainty cannot explain. How can we leverage these DFT constraints to rigorously benchmark nuclear saturation properties of chiral interactions? To address this question, we present a Bayesian mixture model that combines multiple DFT predictions for $(n_0,E_0)$ using an efficient conjugate prior approach. The inferred posterior distribution for the saturation point's mean and covariance matrix follows a Normal-inverse-Wishart class, resulting in posterior predictives in the form of correlated, bivariate $t$-distributions. The DFT uncertainty reports are then used to mix these posteriors using an ordinary Monte Carlo approach. At the 95\% credibility level, we estimate $n_0 \approx 0.157 \pm 0.010 \, \mathrm{fm}^{-3}$ and $E_0 \approx -15.97 \pm 0.40 \, \mathrm{MeV}$ for the marginal (univariate) $t$-distributions. Combined with chiral EFT calculations of the pure neutron matter EOS, we obtain bivariate normal distributions for the nuclear symmetry energy and its slope parameter evaluated at $n_0$: $S_v \approx 32.0 \pm 1.1 \, \mathrm{MeV}$ and $L\approx 52.6\pm 8.1 \, \mathrm{MeV}$ (95\%), respectively. Furthermore, our Bayesian framework is publicly available, so practitioners can readily use and extend our results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Gaussian Mixture Model Solvers for the Boltzmann Equation

This report documents our experience constructing a numerical method for the collisional Boltzmann equation that is capable of accurately capturing the collisionless through strongly collisional limits. We explore three different functional representations and present a detailed account of a numerical method based on a spatially dependent Gaussian mixture model (GMM). The Kullback-Leibler divergence is used as a closeness measure and various expectation maximization (EM) solution algorithms are implemented to find a compact representation in velocity space for distribution functions that exhibit significant non-Maxwellian character. We discuss issues that appear with this representation over a range of Knudsen numbers for a prototypical test problem and demonstrate that the strongly collisional limit recovers a solution to Euler's equations. Looking forward, this approach is broadly applicable to the non-relativistic and relativistic collisional Vlasov equations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Leveraging Gaussian Mixture Models for Detecting Anomalies in Time-Series Data

Test systems must be capable of classifying measured data as expected or anomalous in real time. Anomalous results may portend system failure, and, if undetected, may result in damage to the unit, test equipment, or potential harm to personnel. This report investigates the use of Gaussian Mixture Models (GMMs) as a clustering tool in classifying time-series data.

Wilke, Rudeger H.T. [Sandia National Laboratories ↗

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and↗