Search NASA⌕ Search

SEARCH · Search NASA

Results for “Gaussian Processes (GPs)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗

Latent map Gaussian processes for mixed variable metamodeling

Gaussian processes (GPs) are ubiquitously used in sciences and engineering as metamodels. Standard GPs, however, can only handle numerical or quantitative variables. Here we introduce latent map Gaussian processes (LMGPs) that inherit the attractive properties of GPs and are also applicable to mixed data which have both quantitative and qualitative inputs. The core idea behind LMGPs is to learn a continuous, low-dimensional latent space or manifold which encodes all qualitative inputs. To learn this manifold, we first assign a unique prior vector representation to each combination of qualitative inputs. We then use a low-rank linear map to project these priors on a manifold that characterizes the posterior representations. As the posteriors are quantitative, they can be directly used in any standard correlation function such as the Gaussian or Matern. Hence, the optimal map and the corresponding manifold, along with other hyperparameters of the correlation function, can be systematically learned via maximum likelihood estimation. Through a wide range of analytic and real-world examples, we demonstrate the advantages of LMGPs over state-of-the-art methods in terms of accuracy and versatility. In particular, we show that LMGPs can handle variable-length inputs, have an explainable neural network interpretation, and provide insights into how qualitative inputs affect the response or interact with each other. We also employ LMGPs in Bayesian optimization and illustrate that they can discover optimal compound compositions more efficiently than conventional methods that convert compositions to qualitative variables via manual featurization.

42 ENGINEERING↗

Applying Constrained Bayesian Optimization to the Design of Critical Experiments

Often when planning a criticality experiment, many design configurations are iteratively investigated with a Monte Carlo transport code. The goal is that the experiment will be optimal with respect to some variable, like the fraction of fissions occurring at a certain energy range, while simultaneously being critical. Unfortunately, the Monte Carlo transport simulations are expensive, which can ultimately limit the number of configurations that can be explored. In this work, we present how Gaussian processes (GPs) can be used as a reduced-order model in a constrained Bayesian optimization (CBO) algorithm to design a criticality experiment. The GPs replace the Monte Carlo transport simulations that explore the design space. The CBO algorithm efficiently identifies new points in the design space to run the Monte Carlo transport code while respecting the criticality constraint. It does so in a manner that both improves the accuracy of the GP and finds the approximate global optimum. We demonstrate the performance of CBO with the design of a Thermal Epithermal eXperiment (TEX) for the criticality safety validation of nuclear waste models of the Hanford Tank Farm.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reinforcement Learning via Gaussian Processes with Neural Network Dual Kernels

While deep neural networks (DNNs) and Gaussian Processes (GPs) are both popularly utilized to solve problems in reinforcement learning, both approaches feature undesirable drawbacks for challenging problems. DNNs learn complex non-linear embeddings, but do not naturally quantify uncertainty and are often data-inefficient to train. GPs infer posterior distributions over functions, but popular kernels exhibit limited expressivity on complex and high-dimensional data. Fortunately, recently discovered conjugate and neural tangent kernel functions encode the behavior of overparameterized neural networks in the kernel domain. We demonstrate that these kernels can be efficiently applied to regression and reinforcement learning problems by analyzing a baseline case study.We apply GPs with neural network dual kernels to solve reinforcement learning tasks for the first time. We demonstrate, using the well understood mountain-car problem, that GPs empowered with dual kernels perform at least as well as those using the conventional radial basis function kernel. Finally, we conjecture that by inheriting the probabilistic rigor of GPs and the powerful embedding properties of DNNs, GPs using NN dual kernels will empower future reinforcement learning models on difficult domains.

97 MATHEMATICS AND COMPUTING↗

Fast increased fidelity samplers for approximate Bayesian Gaussian process regression

Gaussian processes (GPs) are common components in Bayesian non-parametric models having a rich methodological literature and strong theoretical grounding. The use of exact GPs in Bayesian models is limited to problems containing several thousand observations due to their prohibitive computational demands. We develop a posterior sampling algorithm using H-matrix approximations that scales at O(n log 2 n). We show that this approximation’s Kullback-Leibler divergence to the true posterior can be made arbitrarily small. Though multidimensional GPs could be used with our algorithm, d-dimensional surfaces are modeled as tensor products of univariate GPs to minimize the cost of matrix construction and maximize computational efficiency. We illustrate the performance of this fast increased fidelity approximate GP, FIFA-GP, using both simulated and non-synthetic data sets

97 MATHEMATICS AND COMPUTING↗

Multi-objective Bayesian alloy design using multi-task Gaussian processes

In design applications, correlations among material properties (such as the tendency for stronger materials to be less ductile) are often neglected. This approach is echoed in multi-objective optimization techniques which treat each performance characteristic as an independent objective, aiming to optimize scalar functions and find optimal Pareto fronts. However, this overlooks the statistical relationships between performance characteristics inherent in a material system. To address this, we propose the use of Bayesian optimization, a highly efficient black-box optimization algorithm known for constructing Gaussian processes (GPs) – uncorrelated surrogates - to model objective functions. Rather than evaluating multiple GPs for each objective function separately, we argue for a shift towards jointly modeling these objective functions, considering their statistical correlations. This integrated approach utilizes naturally occurring relationships among material properties, providing additional information to enhance the performance of the design framework. This requires the replacement of multiple independent GPs with a single multi-task GP, employing a correlation matrix to construct a multi-task kernel function, wherein each task corresponds to a single objective function. Here, we anticipate this refined methodology will better leverage material correlations, improving design optimization results.

36 MATERIALS SCIENCE↗

Combining High-Throughput Experiments and Active Learning to Characterize Deep Eutectic Solvents

The high tunability of deep eutectic solvents (DESs) stems from the ease of changing their precursors and relative compositions. However, measuring the physicochemical properties across large composition and temperature ranges, necessary to properly design target-specific DESs, is tedious and error-prone and represents a bottleneck in the advancement and scalability of DES-based applications. As such, active learning (AL) methodologies based on Gaussian processes (GPs) were developed in this work to minimize the experimental effort necessary to characterize DESs. Owing to its importance for large-scale applications, the reduction of DES viscosity through the addition of a low-molecular-weight solvent was explored as a case study. A high-throughput experimental screening was initially performed on nine different ternary DESs. Then, GPs were successfully trained to predict DES viscosity from its composition and temperature, showcasing the ability of these stochastic, nonparametric models to accurately describe the physicochemical properties of complex mixtures. Finally, the ability of GPs to provide estimates of their own uncertainty was leveraged through an AL framework to minimize the number of data points necessary to obtain accurate viscosity modes. This led to a significant reduction in data requirements, with many systems requiring only five independent viscosity data points to be properly described.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining multitask and transfer learning with deep Gaussian processes for autotuning-based performance engineering

We combine deep Gaussian processes (DGPs) with multitask and transfer learning for the performance modeling and optimization of HPC applications. Deep Gaussian processes merge the uncertainty quantification advantage of Gaussian processes (GPs) with the predictive power of deep learning. Multitask and transfer learning allow for improved learning efficiency when several similar tasks are to be learned simultaneously and when previous learned models are sought to help in the learning of new tasks, respectively. A comparison with state-of-the-art autotuners shows the advantage of our approach on two application problems. In this article, we combine DGPs with multitask and transfer learning to allow for both an improved tuning of an application parameters on problems of interest but also the prediction of parameters on any potential problem the application might encounter.

97 MATHEMATICS AND COMPUTING↗

Bayesian batch optimization for molybdenum versus tungsten inertial confinement fusion double shell target design

Access to reliable, clean energy sources is a major concern for national security. Much research is focused on the “grand challenge” of producing energy via controlled fusion reactions in a laboratory setting. For fusion experiments, specifically inertial confinement fusion (ICF), to produce sufficient energy, the fusion reactions in the ICF fuel need to become self-sustaining and burn deuterium-tritium (DT) fuel efficiently. The recent record-breaking NIF ignition shot was able to achieve this goal as well as produce more energy than used to drive the experiment. This achievement brings self-sustaining fusion-based power systems closer than ever before, capable of providing humans with access to secure, renewable energy. In order to further progress toward the actualization of such power systems, more ICF experiments need to be conducted at large laser facilities such as the United States's National Ignition Facility (NIF) or France's Laser Mega-Joule. The high cost per shot and limited number of shots that are possible per year make it prohibitive to perform large numbers of experiments. As such, experimental design relies heavily on complex predictive physics simulations for high-fidelity “preshot” analysis. These multidimensional, multi-physics, high-fidelity simulations have to account for a variety of input parameters as well as modeling the extreme conditions (pressures and densities) present at ignition. Such simulations (especially in 3D) can become computationally prohibitive to turn around for each ICF experiment. In this work, we explore using Bayesian optimization with Gaussian processes (GPs) to find optimal designs for ICF double shell targets, while keeping computational costs to manageable levels. These double shell targets have an inner shell that grades from beryllium on the outer surface to the higher Z material molybdenum, as opposed to the nominally used tungsten, on the inside in order to trade off between the high performance associated with high density inner shells and capsule stability. We describe our results for “capsule-only” xRAGE simulations to study the physics between different capsule designs, inner shell materials, and potential for future experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Activity coefficient acquisition with thermodynamics–informed active learning for phase diagram construction

This work explores the use of thermodynamics-informed Gaussian processes (GPs) and active learning (AL) to model activity coefficients and construct phase diagrams. Relying on synthetic data generated from an excess Gibbs energy model, GPs were found to accurately describe the activity coefficients of several binary mixtures across large composition and temperature ranges. Moreover, GPs could estimate their own uncertainty and identify composition/temperature regions where activity coefficient data provide the most information to the models. This was leveraged to build AL algorithms targeted at modeling phase equilibria. In many cases, a single active-learning-acquired data point was sufficient to describe the phase diagrams studied. Lastly, the ability of AL to greatly reduce the amount of data needed to obtain accurate models was further verified on experimental case studies, namely individual ion activity coefficients, the solid–liquid and vapor–liquid equilibrium of deep eutectic solvents, and phase equilibria in ternary mixtures.

25 ENERGY STORAGE↗

Scaled Vecchia Approximation for Fast Computer-Model Emulation

Many scientific phenomena are studied using computer experiments consisting of multiple runs of a computer model while varying the input settings. Gaussian processes (GPs) are a popular tool for the analysis of computer experiments, enabling interpolation between input settings, but direct GP inference is computationally infeasible for large datasets. We adapt and extend a powerful class of GP methods from spatial statistics to enable the scalable analysis and emulation of large computer experiments. Specifically, we apply Vecchia’s ordered conditional approximation in a transformed input space, with each input scaled according to how strongly it relates to the computer-model response. The scaling is learned from the data by estimating parameters in the GP covariance function using Fisher scoring. Our methods are highly scalable, enabling estimation, joint prediction, and simulation in near-linear time in the number of model runs. In several numerical examples, our approach substantially outperformed existing methods.

97 MATHEMATICS AND COMPUTING↗

A Fast, Two-dimensional Gaussian Process Method Based on Celerite: Applications to Transiting Exoplanet Discovery and Characterization

Gaussian processes (GPs) are commonly used as a model of stochastic variability in astrophysical time series. In particular, GPs are frequently employed to account for correlated stellar variability in planetary transit light curves. The efficient application of GPs to light curves containing thousands to tens of thousands of data points has been made possible by recent advances in GP methods, including the celerite method. Here we present an extension of the celerite method to two input dimensions where, typically, the second dimension is small. This method scales linearly with the total number of data points when the noise in each large dimension is proportional to the same celerite kernel and only the amplitude of the correlated noise varies in the second dimension. We demonstrate the application of this method to the problem of measuring precise transit parameters from multiwavelength light curves and show that it has the potential to improve transit parameters measurements by orders of magnitude. Applications of this method include transit spectroscopy and exomoon detection, as well a broader set of astronomical problems.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Multi-Fidelity Bayesian Optimization with Gaussian Processes for Double Shell Inertial Confinement Fusion Target Design

Reliable, secure access to energy is a major focus for national security efforts. One potential route to such energy is through fusion reactions in inertial confinement fusion (ICF) experiments. Such experiments are carried out at facilities such as the National Ignition Facility (NIF) in Livermore, California, where high powered lasers are used to compress a DT fuel-containing target to the necessary high temperature, high pressure conditions. These experiments are limited in number, which creates a heavy dependence on high fidelity predictive physics simulations and analysis performed “pre shot,” or before the experiment occurs. Many of these simulations in higher dimensions (2D and 3D) are computationally expensive, so finding optimal simulation-based designs presents its own challenges. In this work, we present our multi-fidelity Bayesian optimization with Gaussian processes (GPs) for ICF double shell targets, where a 1D surrogate model is used to help find a 2D surrogate model, enabling us to find optimal targets in the higher fidelity (2D), while saving computational cost.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Low-Earth Orbit Trajectory Optimization in the Presence of Atmospheric Uncertainty

The previous 20 to 25 years have seen a tremendous increase in space exploration, and with that an increase in the level of logistics planning needed to ensure mission success. For spacecraft that are designed to be periodically re-supplied, a key logistics consumable is propellant, as it constitutes the greatest up-mass on re-supply vehicles. A trajectory design strategy is therefore desired that minimizes propellant usage in order to ease the demand for propellant re-supply missions. This thesis develops such a strategy in three stages, and uses the International Space Station (ISS) as its testbed, as no other LEO spacecraft is more challenging from a space logistics standpoint. First, the ISS trajectory planning problem is formulated as a constrained burn optimization problem assuming a deterministic atmosphere. The cost function is total ∆v, with constraints imposed on longitude of ascending viii node (LAN) and semi-major axis (SMA) altitude. Analytic derivatives are constructed for both the cost and constraints, which are necessary given the 6-week to 2-year time frames being considered. A gradient-based optimizer is then utilized to find locally-optimal solutions to real-world ISS trajectory planning problems. Second, atmospheric uncertainty is addressed by constructing a probabilistic model of space weather data using Gaussian Processes (GPs). Bayesian inference is performed using the GP model to generate mean and covariance estimates for space weather predictions, whose pedigree is assessed against test data. The predictions are then mapped into atmospheric density via the analytic Jacchia-Roberts density model, and the effect of space weather uncertainty on orbital lifetime is examined. Third, an ISS burn execution uncertainty model is developed. This model, along with the space weather uncertainty model, are deployed in a linear covariance analysis to ascertain their combined effect on LAN and SMA altitude dispersions. The deterministic constraints from the original problem are re-formulated as stochastic constraints, where now the constraint uncertainty interval is required to fall within specified bounds. An updated optimization framework is constructed using the original ∆v cost function along with the stochastic constraints to solve the trajectory optimization problem under atmospheric uncertainty. Finally, the complete architecture is summarized for deployment in an operational setting.

Trajectory Optimization↗

Augmenting a Simulation Campaign for Hybrid Computer Model and Field Data Experiments

The Kennedy and O’Hagan (KOH) calibration framework uses coupled Gaussian processes (GPs) to meta-model an expensive simulator (first GP), tune its “knobs” (calibration inputs) to best match observations from a real physical/field experiment and correct for any modeling bias (second GP) when predicting under new field conditions (design inputs). There are well-established methods for placement of design inputs for data-efficient planning of a simulation campaign in isolation, that is, without field data: space-filling, or via criterion like minimum integrated mean-squared prediction error (IMSPE). Analogues within the coupled GP KOH framework are mostly absent from the literature. Here, in this study, we derive a closed form IMSPE criterion for sequentially acquiring new simulator data for KOH. We illustrate how acquisitions space-fill in design space, but concentrate in calibration space. Closed form IMSPE precipitates a closed-form gradient for efficient numerical optimization. We demonstrate that our KOH-IMSPE strategy leads to a more efficient simulation campaign on benchmark problems, and conclude with a showcase on an application to equilibrium concentrations of rare earth elements for a liquid–liquid extraction reaction.

97 MATHEMATICS AND COMPUTING↗

Star–Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

Abstract We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam Subaru Strategic Program data. In our experiments, we evaluate the performance of a principal component analysis embedded GP predictive model against other machine-learning algorithms, including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

Stochastic machine learning via sigma profiles to build a digital chemical space

This work establishes a different paradigm on digital molecular spaces and their efficient navigation by exploiting sigma profiles. To do so, the remarkable capability of Gaussian processes (GPs), a type of stochastic machine learning model, to correlate and predict physicochemical properties from sigma profiles is demonstrated, outperforming state-of-the-art neural networks previously published. The amount of chemical information encoded in sigma profiles eases the learning burden of machine learning models, permitting the training of GPs on small datasets which, due to their negligible computational cost and ease of implementation, are ideal models to be combined with optimization tools such as gradient search or Bayesian optimization (BO). Gradient search is used to efficiently navigate the sigma profile digital space, quickly converging to local extrema of target physicochemical properties. While this requires the availability of pretrained GP models on existing datasets, such limitations are eliminated with the implementation of BO, which can find global extrema with a limited number of iterations. A remarkable example of this is that of BO toward boiling temperature optimization. Holding no knowledge of chemistry except for the sigma profile and boiling temperature of carbon monoxide (the worst possible initial guess), BO finds the global maximum of the available boiling temperature dataset (over 1,000 molecules encompassing more than 40 families of organic and inorganic compounds) in just 15 iterations (i.e., 15 property measurements), cementing sigma profiles as a powerful digital chemical space for molecular optimization and discovery, particularly when little to no experimental data is initially available.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗