Search NASA⌕ Search

SEARCH · Search NASA

Results for “nonparametric regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Unveiling the effect of composition on nuclear waste immobilization glasses’ durability by nonparametric machine learning

Abstract Ensuring the long-term chemical durability of glasses is critical for nuclear waste immobilization operations. Durable glasses usually undergo qualification for disposal based on their response to standardized tests such as the product consistency test or the vapor hydration test (VHT). The VHT uses elevated temperature and water vapor to accelerate glass alteration and the formation of secondary phases. Understanding the relationship between glass composition and VHT response is of fundamental and practical interest. However, this relationship is complex, non-linear, and sometimes fairly variable, posing challenges in identifying the distinct effect of individual oxides on VHT response. Here, we leverage a dataset comprising 654 Hanford low-activity waste (LAW) glasses across a wide compositional envelope and employ various machine learning techniques to explore this relationship. We find that Gaussian process regression (GPR), a nonparametric regression method, yields the highest predictive accuracy. By utilizing the trained model, we discern the influence of each oxide on the glasses’ VHT response. Moreover, we discuss the trade-off between underfitting and overfitting for extrapolating the material performance in the context of sparse and heterogeneous datasets.

36 MATERIALS SCIENCE↗

Learning Functions Varying along a Central Subspace

Many functions of interest are in a high-dimensional space but exhibit low-dimensional structures. This paper studies regression of an s-Hölder function in $R^D$ which varies along a central subspace of dimension $d$ while $d \ll D$. A direct approximation of $f$ in $R^D$ with an accuracy $\varepsilon$ requires the number of samples in the order of $\varepsilon^{-(2s+D)/s}$. In this paper, we analyze the generalized contour regression (GCR) algorithm for the estimation of the central subspace and use piecewise polynomials for function approximation. GCR is among the best estimators for the central subspace, but its sample complexity is an open question. In this paper, we partially answer this questions by proving that if a variance quantity is exactly known, GCR leads to a mean squared estimation error of $O(n^{-1})$ for the central subspace. The estimation error of this variance quantity is also given in this paper. The mean squared regression error of $f$ is proved to be in the order of $(n/\log n)^{-\frac{2s}{2s+d}}$, where the exponent depends on the dimension of the central subspace instead of the ambient space . This result demonstrates that GCR is effective in learning the low-dimensional central subspace. We also propose a modified GCR with improved efficiency. Here, the convergence rate is validated through several numerical experiments.

97 MATHEMATICS AND COMPUTING↗

Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation

Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its theoretical foundations. Nev- ertheless, the majority of these studies examine how well deep neural networks can model functions with uniform regularities. In this paper, we explore a different angle: how deep neural networks can adapt to varying degrees of smoothness in functions and nonuni- form data distributions across different locations and scales. More precisely, we focus on a broad class of functions defined by nonlinear tree-based approximation methods. This class encompasses a range of function types, such as functions with uniform regularities and discontinuous functions. We develop nonparametric approximation and estimation theories for this class using deep ReLU networks. Our results show that deep neural networks are adaptive to the nonuniform smoothness of functions and nonuniform data distributions at different locations and scales. We apply our results to several function classes, and derive the corresponding approximation and generalization errors. The validity of our results is demonstrated through numerical experiments.

97 MATHEMATICS AND COMPUTING↗

Asymptotic inconsistency of the cumulative algorithm for laser-induced damage probability analysis

The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.

Computational methods↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

Parametric and Nonparametric Models of U.S. Cost Overruns for Nuclear Power Plants

This study presents new data-driven models to estimate the effect of capacity on the percentage of cost overruns in the United States for nuclear power plant construction projects before and after the Three Mile Island accident. Parametric and nonparametric models have been developed that describe the significant shifts in nuclear energy costs during the dynamic environment. Employing a contemporary descriptive methodology and a quantitative analysis, we furnish a comprehensive overview of the alterations in cost overrun distribution and show the changes observed in other pivotal metrics alongside cost overruns. Our emphasis lies in documenting the fluctuations in cost overruns alongside nuclear reactor capacity levels and the increase of the overnight capital costs to build nuclear reactors. Our results show that increasing the size of nuclear reactors is not a factor statistically significant to decrease the percentage of cost overruns, and the probit model results provide evidence that an increase in size increases the probability of having cost overruns larger than 100% (double the estimated cost). We also compare our findings to two other regions: Asia and Europe.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗