Search NASA⌕ Search

SEARCH · Search NASA

Results for “nonparametric estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Deep Nonparametric Estimation of Operators between Infinite Dimensional Spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Deep nonparametric estimation of operators between infinite dimensional spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Deep nonparametric estimation of intrinsic data structures by chart autoencoders: Generalization error and robustness

Autoencoders have demonstrated remarkable success in learning low-dimensional latent features of high-dimensional data across various applications. Assuming that data are sampled near a low-dimensional manifold, we employ chart autoencoders, which encode data into low-dimensional latent features on a collection of charts, preserving the topology and geometry of the data manifold. Our paper establishes statistical guarantees on the generalization error of chart autoencoders, and we demonstrate their denoising capabilities by considering n noisy training samples, along with their noise-free counterparts, on a d-dimensional manifold. By training autoencoders, we show that chart autoencoders can effectively denoise the input data with normal noise. We prove that, under proper network architectures, chart autoencoders achieve a squared generalization error in the order of n–$\frac{2}{d+2}$log 4 n, which depends on the intrinsic dimension of the manifold and only weakly depends on the ambient dimension and noise level. We further extend our theory on data with noise containing both normal and tangential components, where chart autoencoders still exhibit a denoising effect for the normal component. As a special case, our theory also applies to classical autoencoders, as long as the data manifold has a global parametrization. Furthermore, our results provide a solid theoretical foundation for the effectiveness of autoencoders, which is further validated through several numerical experiments.

97 MATHEMATICS AND COMPUTING↗

Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation

Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its theoretical foundations. Nev- ertheless, the majority of these studies examine how well deep neural networks can model functions with uniform regularities. In this paper, we explore a different angle: how deep neural networks can adapt to varying degrees of smoothness in functions and nonuni- form data distributions across different locations and scales. More precisely, we focus on a broad class of functions defined by nonlinear tree-based approximation methods. This class encompasses a range of function types, such as functions with uniform regularities and discontinuous functions. We develop nonparametric approximation and estimation theories for this class using deep ReLU networks. Our results show that deep neural networks are adaptive to the nonuniform smoothness of functions and nonuniform data distributions at different locations and scales. We apply our results to several function classes, and derive the corresponding approximation and generalization errors. The validity of our results is demonstrated through numerical experiments.

97 MATHEMATICS AND COMPUTING↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Advances in statistical methods for cancer surveillance research: an age-period-cohort perspective

Background: Analysis of Lexis diagrams (population-based cancer incidence and mortality rates indexed by age group and calendar period) requires specialized statistical methods. However, existing methods have limitations that can now be overcome using new approaches. Methods: We assembled a “toolbox” of novel methods to identify trends and patterns by age group, calendar period, and birth cohort. We evaluated operating characteristics across 152 cancer incidence Lexis diagrams compiled from United States (US) Surveillance, Epidemiology and End Results Program data for 21 leading cancers in men and women in four race and ethnicity groups (the “cancer incidence panel”). Results: Nonparametric singular values adaptive kernel filtration (SIFT) decreased the estimated root mean squared error by 90% across the cancer incidence panel. A novel method for semi-parametric age-period-cohort analysis (SAGE) provided optimally smoothed estimates of age-period-cohort (APC) estimable functions and stabilized estimates of lack-of-fit (LOF). SAGE identified statistically significant birth cohort effects across the entire cancer panel; LOF had little impact. As illustrated for colon cancer, newly developed methods for comparative age-period-cohort analysis can elucidate cancer heterogeneity that would otherwise be difficult or impossible to discern using standard methods. Conclusions: Cancer surveillance researchers can now identify fine-scale temporal signals with unprecedented accuracy and elucidate cancer heterogeneity with unprecedented specificity. Birth cohort effects are ubiquitous modulators of cancer incidence in the US. The novel methods described here can advance cancer surveillance research.

60 APPLIED LIFE SCIENCES↗

Parametric and Nonparametric Models of U.S. Cost Overruns for Nuclear Power Plants

This study presents new data-driven models to estimate the effect of capacity on the percentage of cost overruns in the United States for nuclear power plant construction projects before and after the Three Mile Island accident. Parametric and nonparametric models have been developed that describe the significant shifts in nuclear energy costs during the dynamic environment. Employing a contemporary descriptive methodology and a quantitative analysis, we furnish a comprehensive overview of the alterations in cost overrun distribution and show the changes observed in other pivotal metrics alongside cost overruns. Our emphasis lies in documenting the fluctuations in cost overruns alongside nuclear reactor capacity levels and the increase of the overnight capital costs to build nuclear reactors. Our results show that increasing the size of nuclear reactors is not a factor statistically significant to decrease the percentage of cost overruns, and the probit model results provide evidence that an increase in size increases the probability of having cost overruns larger than 100% (double the estimated cost). We also compare our findings to two other regions: Asia and Europe.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Assessing equation of state-independent relations for neutron stars with nonparametric models

Relations between neutron star properties that do not depend on the nuclear equation of state offer insights on neutron star physics and have practical applications in data analysis. Such relations are obtained by fitting to a range of phenomenological or nuclear physics equation of state models, each of which may have varying degrees of accuracy. In this study we revisit commonly used relations and reassess them with a very flexible set of phenomenological nonparametric equation of state models that are based on Gaussian processes. Our models correspond to two sets: equations of state which mimic hadronic models, and equations of state with rapidly changing behavior that resemble phase transitions. Here we quantify the accuracy of relations under both sets and discuss their applicability with respect to expected upcoming statistical uncertainties of astrophysical observations. We further propose a goodness-of-fit metric which provides an estimate for the systematic error introduced by using the relation to model a certain equation-of-state set. Overall, the nonparametric distribution is more poorly fit with existing relations, with the I–Love–Q relations retaining the highest degree of universality. Fits degrade for relations involving the tidal deformability, such as the binary-Love and compactness-Love relations, and when introducing phase transition phenomenology. For most relations, systematic errors are comparable to current statistical uncertainties under the nonparametric equation of state distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Combining High-Throughput Experiments and Active Learning to Characterize Deep Eutectic Solvents

The high tunability of deep eutectic solvents (DESs) stems from the ease of changing their precursors and relative compositions. However, measuring the physicochemical properties across large composition and temperature ranges, necessary to properly design target-specific DESs, is tedious and error-prone and represents a bottleneck in the advancement and scalability of DES-based applications. As such, active learning (AL) methodologies based on Gaussian processes (GPs) were developed in this work to minimize the experimental effort necessary to characterize DESs. Owing to its importance for large-scale applications, the reduction of DES viscosity through the addition of a low-molecular-weight solvent was explored as a case study. A high-throughput experimental screening was initially performed on nine different ternary DESs. Then, GPs were successfully trained to predict DES viscosity from its composition and temperature, showcasing the ability of these stochastic, nonparametric models to accurately describe the physicochemical properties of complex mixtures. Finally, the ability of GPs to provide estimates of their own uncertainty was leveraged through an AL framework to minimize the number of data points necessary to obtain accurate viscosity modes. This led to a significant reduction in data requirements, with many systems requiring only five independent viscosity data points to be properly described.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Asymptotic inconsistency of the cumulative algorithm for laser-induced damage probability analysis

The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.

Computational methods↗

A Nonparametric Method for the Inference of Halo Occupation Distributions

The galaxy–halo connection traces processes by which galaxies form and evolve. The halo occupation distribution (HOD) describes the relationship between galaxies and their host dark matter haloes. Measurements of the galaxy two-point correlation function (2PCF) allow us to extract information about the HODs of observed galaxy samples. Several parametric HOD models have been proposed in the literature, but the choice of parameterization restricts the space of possible HODs. To resolve this issue, we introduce a nonparametric HOD fitting method in which we train an emulator to learn the mappings among the galaxy 2PCF, physical properties used to select galaxy samples, and the HOD, all obtained from simulated past light cones constructed with the Santa Cruz semianalytic model. Implementing this emulator within a likelihood analysis framework, we derive constraints on the HOD of a galaxy sample when provided with a measurement of its 2PCF. Using the emulator to accelerate likelihood evaluations, we test the nonparametric HOD approach on a set of 2PCFs for mock galaxy samples drawn from the TNG100-1 simulation and selected above threshold values of stellar mass and star formation rate. Our framework is able to recover TNG100-1 HODs within 0.2 dex. We use the TNG100-1 mocks to tune the reported uncertainties to estimate those expected in the analysis of observations. Comparing to parametric HOD modelling routines applied to the same mock galaxy samples, our approach consistently infers the HOD with comparable or greater precision and accuracy.

Kennedy, Jacob [Rutgers Univ., Piscataway, NJ (Uni↗

Nonparametric Multiparticle Set Methods for Interpreting Environmental Samples

Collection and analysis of environmental samples is commonly used by a range of stakeholders in nuclear safeguards and security contexts. While the ubiquity of samples and their transport in the environment allow regular collection, developing and demonstrating methods for analyzing these samples is difficult. In this work, an environmental sample consists of a set of one or more individual particles. Recent advances in reactor simulation have allowed us to generate data that are more representative of real-world environmental samples, enabling statistically defensible method development and testing. The most notable of these advances is a drastic increase in the number of material depletion regions, which allows our simulations to capture the variation in isotopic composition seen at length scales consistent with environmental samples. Traditional approaches for handling multiparticle samples treat each particle in the sample individually, estimating the quantity of interest (e.g., core-average burnup) resulting from measurement and analysis of signatures (e.g., nuclide assays) from each individual particle. Individual estimates are then averaged to generate a single estimate of the quantity of interest over the entire sample. In this presentation, we introduce two novel approaches for interpreting environmental samples that comprise of multiple particles: (1) the Quantile-Quantile Comparator, which uses a multivariate generalization of quantile-quantile plots for comparing unknown statistical distributions, and (2) the Set Transformer, an attention-based neural network module designed to model interactions among elements (particles) in the input set (sample). Statistically representative sampling cannot be guaranteed as samples are passively collected and are beholden to what particles are available in the environment. These new analysis methods for set-input problems are expected to be more robust than traditional approaches to issues of sampling bias where particles are not uniformly distributed throughout regions of interest, as well as generally outperform traditional approaches by jointly considering all elements in the set. We will present results comparing the performance of traditional single particle approaches and the novel Quantile-Quantile Comparator and Set Transformer for interpretation of simulated environmental samples.

Phathanapirom, Birdy↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

Stellar Mass Calibrations for Local Low-mass Galaxies

The stellar masses of galaxies are measured from integrated light via several methods—however, few of these methods were designed for low-mass (M ⋆ ≲ 10 8 M ⊙ ) “dwarf” galaxies, whose properties (e.g., stochastic star formation, low metallicity) pose unique challenges for estimating stellar masses. In this work, we quantify the precision and accuracy at which stellar masses of low-mass galaxies can be recovered using UV/optical/IR photometry. We use mock observations of 469 low-mass galaxies from a variety of models, including both semi-empirical models (GRUMPY and UniverseMachine-SAGA) and cosmological baryonic zoom-in simulations (MARVELous Dwarfs and FIRE-2), to test literature color–M ⋆ /L relations and multiwavelength spectral energy distribution (SED) mass estimators. We identify a list of “best practices” for measuring stellar masses of low-mass galaxies from integrated photometry. We find that literature color–M ⋆ /L relations are often unable to capture the bursty star formation histories (SFHs) of low-mass galaxies, and we develop an updated prescription for stellar mass based on g − r color that is better able to recover stellar masses for the bursty low-mass galaxies in our sample (with ∼0.1 dex precision). SED fitting can also precisely recover stellar masses of low-mass galaxies, but this requires thoughtful choices about the form of the assumed SFH: Parametric SFHs can underestimate stellar mass by as much as ∼0.4 dex, while nonparametric SFHs recover true stellar masses with insignificant offset (−0.03 ± 0.11 dex). Finally, we also caution that noninformative (wide) dust attenuation priors may introduce M ⋆ uncertainties of up to ∼0.6 dex.

de los Reyes, Mithi A. C. [Amherst College, MA (Un↗

Learning Functions Varying along a Central Subspace

Many functions of interest are in a high-dimensional space but exhibit low-dimensional structures. This paper studies regression of an s-Hölder function in $R^D$ which varies along a central subspace of dimension $d$ while $d \ll D$. A direct approximation of $f$ in $R^D$ with an accuracy $\varepsilon$ requires the number of samples in the order of $\varepsilon^{-(2s+D)/s}$. In this paper, we analyze the generalized contour regression (GCR) algorithm for the estimation of the central subspace and use piecewise polynomials for function approximation. GCR is among the best estimators for the central subspace, but its sample complexity is an open question. In this paper, we partially answer this questions by proving that if a variance quantity is exactly known, GCR leads to a mean squared estimation error of $O(n^{-1})$ for the central subspace. The estimation error of this variance quantity is also given in this paper. The mean squared regression error of $f$ is proved to be in the order of $(n/\log n)^{-\frac{2s}{2s+d}}$, where the exponent depends on the dimension of the central subspace instead of the ambient space . This result demonstrates that GCR is effective in learning the low-dimensional central subspace. We also propose a modified GCR with improved efficiency. Here, the convergence rate is validated through several numerical experiments.

97 MATHEMATICS AND COMPUTING↗