Search NASA⌕ Search

SEARCH · Search NASA

Results for “Posterior regularized”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Posterior Regularized Bayesian Neural Network incorporating soft and hard knowledge constraints

Neural Networks (NNs) have been widely used in supervised learning due to their ability to model complex nonlinear patterns, often presented in high-dimensional data such as images and text. However, traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs (BNNS) could help measure the uncertainty by considering the distributions of the NN model parameters. Besides, domain knowledge is commonly available and could improve the performance of BNNs if it can be appropriately incorporated. In this work, we propose a novel Posterior-Regularized Bayesian Neural Network (PR-BNN) model by incorporating different types of knowledge constraints, such as the soft and hard constraints, as a posterior regularization term. Furthermore, we propose to combine the augmented Lagrangian method and the existing BNN solvers for efficient inference. Furthermore, the experiments in simulation and two case studies about aviation landing prediction and solar energy output prediction have shown the knowledge constraints and the performance improvement of the proposed model over traditional BNNs without the constraints.

14 SOLAR ENERGY↗

Posterior Regularized Bayesian Neural Network

Traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs(BNNs) could help measure the confidence level by using distributions in NNs modeling. Besides, knowledge is commonly available and could improve the performance of BNNs if it can be properly incorporated. In this work, we propose a novel Posterior-Regularized BNN(PR-BNN) model by incorporating soft and hard constraints as a posterior regularization term. We also propose an augmented Lagrangian method and stochastic optimization algorithm for efficient updating via Monte Carlo sampling. The simulations and case studies for solar PV plants have shown the performance improvement of the proposed model over traditional BNNs.

97 MATHEMATICS AND COMPUTING↗

Posterior Regularized Bayesian Neural Network

Traditional NNs often lack the ability for uncertainty quantification. Bayesian NNs(BNNs) could help measure the confidence level by using distributions in NNs modeling. Besides, knowledge is commonly available and could improve the performance of BNNs if it can be properly incorporated. In this work, we propose a novel Posterior-Regularized BNN(PR-BNN) model by incorporating soft and hard constraints as a posterior regularization term. We also propose an augmented Lagrangian method and stochastic optimization algorithm for efficient updating via Monte Carlo sampling. The simulations and case studies in solar energy prediction have shown the performance improvement of the proposed model over traditional BNNs.

14 SOLAR ENERGY↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

Estimating location parameters in a mixture

The problem of estimating the parameters in a finite mixture is considered. The approach is based on an integral equation formulation of the form h sub t (x) = integral (limits b and a) f(x-y) g sut t (y)dy where h sub t is a smoothed version of h and g sub t is a prior function that tends to be concentrated on the translation values. A solution for g sub t that uses the method of regularization and one based on a posterior operator approach is considered. Numerical simulations are presented to bring out some of the estimation and numerical problems of these approaches.

Heydorn, R. P.↗

Monotonic Gaussian Process for Physics-Constrained Machine Learning With Materials Science Applications

Physics-constrained machine learning is emerging as an important topic in the field of machine learning for physics. One of the most significant advantages of incorporating physics constraints into machine learning methods is that the resulting model requires significantly less data to train. By incorporating physical rules into the machine learning formulation itself, the predictions are expected to be physically plausible. Gaussian process (GP) is perhaps one of the most common methods in machine learning for small datasets. In this paper, we investigate the possibility of constraining a GP formulation with monotonicity on three different material datasets, where one experimental and two computational datasets are used. The monotonic GP is compared against the regular GP, where a significant reduction in the posterior variance is observed. The monotonic GP is strictly monotonic in the interpolation regime, but in the extrapolation regime, the monotonic effect starts fading away as one goes beyond the training dataset. Imposing monotonicity on the GP comes at a small accuracy cost, compared to the regular GP. The monotonic GP is perhaps most useful in applications where data are scarce and noisy, and monotonicity is supported by strong physical evidence.

36 MATERIALS SCIENCE↗

Assessing correlated truncation errors in modern nucleon-nucleon potentials

We test the BUQEYE model of correlated effective field theory (EFT) truncation errors on Reinert, Krebs, and Epelbaum's semilocal momentum-space implementation of the chiral EFT (𝜒⁢EFT ) expansion of the nucleon-nucleon (NN) potential. This Bayesian model hypothesizes that dimensionless coefficient functions extracted from the order-by-order corrections to NN observables can be treated as draws from a Gaussian process (GP). We combine a variety of graphical and statistical diagnostics to assess when predicted observables have a 𝜒⁢EFT convergence pattern consistent with the hypothesized GP statistical model. Our conclusions are that, first, the BUQEYE model is generally applicable to the potential investigated here, which enables statistically principled estimates of the impact of higher EFT orders on observables. Second, parameters defining the extracted coefficients such as the expansion parameter 𝑄 must be well chosen for the coefficients to exhibit a regular convergence pattern—a property we exploit to obtain posterior distributions for such quantities. Third, the assumption of GP stationarity across lab energy and scattering angle is not generally met; this necessitates adjustments in future work. We provide a workflow and interpretive guide for our analysis framework, and show what can be inferred about probability distributions for 𝑄, the EFT breakdown scale Λ 𝑏 , the scale associated with soft physics in the 𝜒⁢EFT potential 𝑚 eff , and the GP hyperparameters. All our results can be reproduced using a publicly available Jupyter notebook, which can be straightforwardly modified to analyze other 𝜒⁢EFT NN potentials.

Bayesian methods↗

Neural posterior unfolding

Differential cross section measurements are the currency of scientific exchange in particle and nuclear physics. A key challenge for these analyses is the correction for detector distortions, known as deconvolution or unfolding. Binned unfolding of cross section measurements traditionally rely on the regularized inversion of the response matrix that represents the detector response, mapping pre-detector (`particle level') observables to post-detector (`detector level') observables. In this paper we introduce Neural Posterior Unfolding, a modern, Bayesian approach that leverages normalizing flows for unfolding. By using normalizing flows for neural posterior estimation, NPU offers several key advantages including implicit regularization through the neural network architecture, fast amortized inference that eliminates the need for repeated retraining, and direct access to the full uncertainty in the unfolded result. In addition to introducing NPU, we implement a classical Bayesian unfolding method called Fully Bayesian Unfolding (FBU) in modern Python so it can also be studied. These tools are validated on simple Gaussian examples and then tested on simulated jet substructure examples from the Large Hadron Collider (LHC). We find that the Bayesian methods are effective and worth additional development to be analysis ready for cross section measurements at the LHC and beyond.

Analysis and statistical methods↗

Ensemble variational Fokker-Planck methods for data assimilation

Particle flow filters solve Bayesian inference problems by smoothly transforming a set of particles into samples from the posterior distribution. Particles move in state space under the flow of an McKean-Vlasov-Itˆo process. This work introduces the Variational Fokker-Planck (VFP) framework for data assimilation, a general approach that includes previously known particle flow filters as special cases. The McKean-Vlasov-Itˆo process that transforms particles is defined via an optimal drift that depends on the selected diffusion term. It is established that the underlying probability density - sampled by the ensemble of particles - converges to the Bayesian posterior probability density. For a finite number of particles the optimal drift contains a regularization term that nudges particles toward becoming independent random variables. Based on this analysis, we derive computationally-feasible approximate regularization approaches that penalize the mutual information between pairs of particles, and avoid particle collapse. Moreover, the diffusion plays a role akin to a particle rejuvenation approach that aims to alleviate particle collapse. The VFP framework is very flexible. Different assumptions on prior and intermediate probability distributions can be used to implement the optimal drift, and localization and covariance shrinkage can be applied to alleviate the curse of dimensionality. A robust implicit-explicit method is discussed for the efficient integration of stiff McKean- Vlasov-Itˆo processes. Here, the effectiveness of the VFP framework is demonstrated on three progressively more challenging test problems, namely the Lorenz ’63, Lorenz ’96 and the quasi-geostrophic equations.

97 MATHEMATICS AND COMPUTING↗

Galaxy cluster matter profiles - I. Self-similarity, mass calibration, and observable-mass relation validation employing cluster mass posteriors

We present a study of the weak lensing inferred matter profiles ΔΣ(R) of 698 South Pole Telescope (SPT) thermal Sunyaev-Zel’dovich effect (tSZE) selected and MCMF optically confirmed galaxy clusters in the redshift range 0.25 < z < 0.94 that have associated weak gravitational lensing shear profiles from the Dark Energy Survey (DES). Rescaling these profiles to account for the mass dependent size and the redshift dependent density produces average rescaled matter profiles ΔΣ(R/R200c)/(ρcritR200c) with a lower dispersion than the unscaled ΔΣ(R) versions, indicating a significant degree of self-similarity. Galaxy clusters from hydrodynamical simulations also exhibit matter profiles that suggest a high degree of self-similarity, with RMS variation among the average rescaled matter profiles with redshift and mass falling by a factor of approximately six and 23, respectively, compared to the unscaled average matter profiles. We employed this regularity in a new Bayesian method for weak lensing mass calibration that employs the so-called cluster mass posterior P(M200|ζ̂, λ̂, z), which describes the individual cluster masses given their tSZE (ζ̂) and optical (λ̂, z) observables. This method enables simultaneous constraints on richness λ-mass and tSZE detection significance ζ-mass relations using average rescaled cluster matter profiles. We validated the method using realistic mock datasets and present observable-mass relation constraints for the SPT×DES sample, where we constrained the amplitude, mass trend, redshift trend, and intrinsic scatter. Our observable-mass relation results are in agreement with the mass calibration derived from the recent cosmological analysis of the SPT×DES data based on a cluster-by-cluster lensing calibration. Our new mass calibration technique offers a higher efficiency when compared to the single cluster calibration technique. We present new validation tests of the observable-mass relation that indicate the underlying power-law form and scatter are adequate to describe the real cluster sample but that also suggest a redshift variation in the intrinsic scatter of the λ-mass relation may offer a better description. In addition, the average rescaled matter profiles offer high signal-to-noise ratio (S/N) constraints on the shape of real cluster matter profiles, which are in good agreement with available hydrodynamical ΛCDM simulations. This high S/N profile contains information about baryon feedback, the collisional nature of dark matter, and potential deviations from general relativity.Key words: gravitational lensing: weak / galaxies: clusters: general / large-scale structure of Universe

79 ASTRONOMY AND ASTROPHYSICS↗

Accelerating Markov Chain Monte Carlo sampling with diffusion models

Global fits of physics models require efficient methods for exploring high-dimensional and/or multimodal posterior functions. We introduce a novel method for accelerating Markov Chain Monte Carlo (MCMC) sampling by pairing a Metropolis-Hastings algorithm with a diffusion model that can draw global samples with the aim of approximating the posterior. We briefly review diffusion models in the context of image synthesis before providing a streamlined diffusion model tailored towards low-dimensional data arrays. We then present our adapted Metropolis-Hastings algorithm which combines local proposals with global proposals taken from a diffusion model that is regularly trained on the samples produced during the MCMC run. Our approach leads to a significant reduction in the number of likelihood evaluations required to obtain an accurate representation of the Bayesian posterior across several analytic functions, as well as for a physical example based on a global fit of parton distribution functions. Our method is extensible to other MCMC techniques, and we briefly compare our method to similar approaches based on normalising flows. A code implementation can be found at https://github.com/NickHunt-Smith/MCMC-diffusion.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian operator inference for data-driven reduced-order modeling

This work proposes a Bayesian inference method for the reduced-order modeling of time-dependent systems. Informed by the structure of the governing equations, the task of learning a reduced-order model from data is posed as a Bayesian inverse problem with Gaussian prior and likelihood. The resulting posterior distribution characterizes the operators defining the reduced-order model, hence the predictions subsequently issued by the reduced-order model are endowed with uncertainty. The statistical moments of these predictions are estimated via a Monte Carlo sampling of the posterior distribution. Since the reduced models are fast to solve, this sampling is computationally efficient. Furthermore, the proposed Bayesian framework provides a statistical interpretation of the regularization term that is present in the deterministic operator inference problem, and the empirical Bayes approach of maximum marginal likelihood suggests a selection algorithm for the regularization hyperparameters. The proposed method is demonstrated on two examples: the compressible Euler equations with noise-corrupted observations, and a single-injector combustion process.

97 MATHEMATICS AND COMPUTING↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Detecting outbreaks using a spatial latent field

In this paper, we present a method for estimating the infection-rate of a disease as a spatial-temporal field. Our data comprises time-series case-counts of symptomatic patients in various areal units of a region. We extend an epidemiological model, originally designed for a single areal unit, to accommodate multiple units. The field estimation is framed within a Bayesian context, utilizing a parameterized Gaussian random field as a spatial prior. We apply an adaptive Markov chain Monte Carlo method to sample the posterior distribution of the model parameters condition on COVID-19 case-count data from three adjacent counties in New Mexico, USA. Our results suggest that the correlation between epidemiological dynamics in neighboring regions helps regularize estimations in areas with high variance (i.e., poor quality) data. Using the calibrated epidemic model, we forecast the infection-rate over each areal unit and develop a simple anomaly detector to signal new epidemic waves. Our findings show that anomaly detector based on estimated infection-rates outperforms a conventional algorithm that relies solely on case-counts.

Safta, Cosmin [Sandia National Laboratories (SNL-C↗

Exercise Effects on the Brain and Sensorimotor Function in Bed Rest

Long duration spaceflight microgravity results in cephalad fluid shifts and deficits in posture control and locomotion. Effects of microgravity on sensorimotor function have been investigated on Earth using head down tilt bed rest (HDBR). HDBR serves as a spaceflight analogue because it mimics microgravity in body unloading and bodily fluid shifts. Preliminary results from our prior 70 days HDBR studies showed that HDBR is associated with focal gray matter (GM) changes and gait and balance deficits, as well as changes in brain functional connectivity. In consideration of the health and performance of crewmembers we investigated whether exercise reduces the effects of HDBR on GM, functional connectivity, and motor performance. Numerous studies have shown beneficial effects of exercise on brain health. We therefore hypothesized that an exercise intervention during HDBR could potentially mitigate the effects of HDBR on the central nervous system. Eighteen subjects were assessed before (12 and 7 days), during (7, 30, and ~70 days) and after (8 and 12 days) 70 days of 6-degrees HDBR at the NASA HDBR facility in UTMB, Galveston, TX, US. Each subject was randomly assigned to a control group or one of two exercise groups. Exercise consisted of daily supine exercise which started 20 days before the start of HDBR. The exercise subjects participated either in regular aerobic and resistance exercise (e.g. squat, heel raise, leg press, cycling and treadmill running), or aerobic and resistance exercise using a flywheel apparatus (rowing). Aerobic and resistance exercise intensity in both groups was similar, which is why we collapsed the two exercise groups for the current experiment. During each time point T1-weighted MRI scans and resting state functional connectivity scans were obtained using a 3T Siemens scanner. Focal changes over time in GM density were assessed using voxel based morphometry (VBM8) under SPM. Changes in resting state functional connectivity was assessed using both a region of interest (ROI, or seed-to-voxel) approach as well as a whole brain intrinsic connectivity (i.e., voxel-to-voxel) analysis. For the ROI analysis we selected 11 ROIs of brain regions that are involved in sensorimotor function (i.e., L. Insular C., L. Putamen, R. Premotor C., L.+R. Primary Motor C., R. Vestibular C., L. Posterior Cingulate G., R. Cerebellum Lobule V + VIIIb + Crus I, and the R. Superior Parietal G.) and correlated their time course of brain activation during rest with all other voxels in the brain. The whole brain connectivity analysis tests changes in the strength of the global connectivity pattern between each voxel and the rest of the brain. Functional mobility was assessed using an obstacle course. Vestibular contribution to balance was measured using Neurocom Sensory Organization Test 5. Behavioral measures were assessed pre-HDBR, and 0, 8 and 12 days post-HDBR. Linear mixed models were used to test for effects of time, group, and group-by-time interactions. Family-wise error corrected VBM revealed significantly larger increases in GM volume in the right primary motor cortex in bed rest control subjects than in bed rest exercise subjects. No other significant group by time interactions in gray matter changes with bed rest were observed. Functional connectivity MRI revealed that the increase in connectivity during bed rest of the left putamen with the bilateral midsagittal precunes and the right cingulate gyrus was larger in bed rest control subjects than in bed rest exercise subjects. Furthermore, the increase in functional connectivity with bed rest of the right premotor cortex with the right inferior frontal gyrus and the right primary motor cortex with the bilateral premotor cortex was smaller in bed rest control subjects than in bed rest exercise subjects. Functional mobility performance was less affected by HDBR in exercise subjects than in control subjects and post HDBR exercise subjects recovered faster than control subjects. The group performance differences and GM changes were not related. Exercise in HDBR partially mitigates the adverse effect of HDBR on functional mobility, particularly during the post-bed rest recovery phase. In addition, exercise appears to result in differential brain structural and functional changes in motor regions such as the primary motor cortex, the premotor cortex and the putamen. Whether these central nervous system changes are related to motor behavioral changes including gait and balance warrants further research.

Koppelmans, V.↗