Search NASA⌕ Search

SEARCH · Search NASA

Results for “Non-parametric estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING↗

Out-of-distribution detection with non-parametric density estimation for models predicting processing history of uranium ore concentrates

The rapid advancement in machine learning (ML) and computer vision (CV) coincides with the growth of interest in deploying these ML/CV models in numerous fields from medicine to social science. Similar to those areas, we have witnessed a great number of works in materials science employing ML/CV models – neural networks in particular – in their studies in recent years. These models have proven to obtain accurate performance in various tasks. However, these models struggle to attain a similar performance when encountering test samples coming from a distribution that is different from the training set. More importantly, they fail without providing any warning to the users. Therefore, we propose a framework for detecting out-of-distribution (OOD) samples to alert users when a human intervention might be necessary in this work. Specifically, we explore the use of a non-parametric density estimation method to detect OOD samples. Here, we assess OOD detection capability of the proposed framework on ML models developed for categorizing precipitation routes of U 3 O 8 when encountering OOD datasets that contain samples (1) undergone different imaging acquisition process, (2) undergone different material synthesis process, and (3) different materials than ID set. Through those experiments, we achieve an average area under the receiver operating characteristic (AUROC) of at least 91% on average in detecting OOD samples. With minimal overhead cost and superior performance, the proposed framework enables a reliable and safe system when deploying in real-world scenarios.

Convolutional neural networks↗

Generative adversarial networks for scintillation signal simulation in EXO-200

Generative Adversarial Networks trained on samples of simulated or actual events have been proposed as a way of generating large simulated datasets at a reduced computational cost. In this work, a novel approach to perform the simulation of photodetector signals from the time projection chamber of the EXO-200 experiment is demonstrated. The method is based on a Wasserstein Generative Adversarial Network — a deep learning technique allowing for implicit non-parametric estimation of the population distribution for a given set of objects. Our network is trained on real calibration data using raw scintillation waveforms as input. We find that it is able to produce high-quality simulated waveforms an order of magnitude faster than the traditional simulation approach and, importantly, generalize from the training sample and discern salient high-level features of the data. In particular, the network correctly deduces position dependency of scintillation light response in the detector and correctly recognizes dead photodetector channels. Furthermore, the network output is then integrated into the EXO-200 analysis framework to show that the standard EXO-200 reconstruction routine processes the simulated waveforms to produce energy distributions comparable to that of real waveforms. Finally, the remaining discrepancies and potential ways to improve the approach further are highlighted.

47 OTHER INSTRUMENTATION↗

Local modeling for FRF estimation with noisy input measurements

The frequency response function (FRF) is an essential means by which dynamic systems are qualified. In recent years, local modeling approaches have been extensively researched and shown to significantly outperform traditional FRF estimators. However, the standard local modeling approach assumes a perfectly-known system input, which results in biased FRF estimates in the presence of input noise. This paper derives a simple adjustment that can be used to improve FRF estimation for systems subjected to random excitation with noisy input data. This improvement can be implemented with little modification to standard local modeling algorithms and with little additional computational burden. The adjustment is coupled with a model selection procedure to avoid underfitting and overfitting. In conclusion, the methods presented in this paper are validated on a simulation, and they are shown to reduce bias due to input noise.

47 OTHER INSTRUMENTATION↗

Detecting Anomalies in Time Series Using Kernel Density Approaches

This paper introduces a novel anomaly detection approach tailored for time series data with exclusive reliance on normal events during training. Our key innovation lies in the application of kernel-density estimation (KDE) to scrutinize reconstruction errors, providing an empirically derived probability distribution for normal events post-reconstruction. This non-parametric density estimation technique offers a nuanced understanding of anomaly detection, differentiating it from prevalent threshold-based mechanisms in existing methodologies. In post-training, events are encoded, decoded, and evaluated against the estimated density, providing a comprehensive notion of normality. In addition, we propose a data augmentation strategy involving variational autoencoder-generated events and a smoothing step for enhanced model robustness. The significance of our autoencoder-based approach is evident in its capacity to learn normal representation without prior anomaly knowledge. Through the KDE step on reconstruction errors, our method addresses the versatility of anomalies, departing from assumptions tied to larger reconstruction errors for anomalous events. Our proposed likelihood measure then distinguishes normal from anomalous events, providing a concise yet comprehensive anomaly detection solution. The extensive experimental results support the feasibility of our proposed method, yielding significantly improved classification performance by nearly 10% on the UCR benchmark data.

Frehner, Robin↗

SDSS-IV MaNGA: The incidence of major mergers in type I and II AGN host galaxies in the DR15 sample

We present a study on the incidence of major mergers and their impact on the triggering of nuclear activity in 47 type I and 236 type II optically selected AGN from the MaNGA DR15 sample. From an estimate of non-parametric image predictors (Gini, M_20, concentration (C), asymmetry (A), clumpiness (S), Sérsic index (n), and shape asymmetry ()) using the SDSS images, in combination with a Linear Discriminant Analysis Method, we identified major mergers and merger stages. We reinforced our results by looking for bright tidal features in our post-processed SDSS and DESI legacy images. We find a statistically significant higher incidence of major mergers of 29 per cent ± 3 per cent in our type I+II AGN sample compared to 22 per cent ± 0.8 per cent for a non-AGN sample matched in redshift, stellar mass, colour, and morphological type, finding also a prevalence of post-coalescence (51 per cent ± 5 per cent) over pre-coalescence (23 per cent ± 6 per cent) merger stages. The levels of AGN activity among our massive major mergers are similar to those reported in other works using [O iii] tracers. However, similar levels are produced by our AGN-galaxies hosting stellar bars, suggesting that major mergers are important promoters of nuclear activity but are not the main nor the only mechanism behind the AGN triggering. The tidal strength parameter Q was considered at various scales looking for environmental differences that could affect our results on the merger incidence, finding non-significant differences. Finally, the H-H β diagram could be used as an empirical predictor for the flux coming from an AGN source, useful to correct photometric quantities in large AGN samples emerging from surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗

Iron-Chromium-Aluminum Accident Tolerant Fuel Concept Source Term Accident Sequence Analysis - High Burnup Fuel Source Term Accident Sequence Analysis Supplement

To extend NUREG-1465 and high burnup fuel source term (SAND2023-01313) recommendations, representative radiological releases to containment – patterned after NUREG-1465 – have been evaluated for LWRs utilizing iron-chromium-aluminum (FeCrAl) alloys in place of zirconium-based alloys in major core structures (cladding and fuel canisters) and high burnup fuel with enrichments of 8% and 10% for PWRs and BWRs, respectively. Representative radionuclide releases are generated for this accident tolerant fuel concept by applying non-parametric bootstrap methods to MELCOR simulation results. Accident scenarios considered in this analysis include principle contributors to historical core damage frequency estimates for a range of nuclear reactor technologies representative of the operating U.S.A. fleet of nuclear reactors.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Cr-coated Accident Tolerant Fuel Concept Source Term Accident Sequence Analysis - High Burnup Fuel Source Term Accident Sequence Analysis Supplement

To extend NUREG-1465 and high burnup fuel source term (SAND2023-01313) recommendations, representative radiological releases to containment – patterned after NUREG-1465 – have been evaluated for LWRs utilizing the chromium-coating on major zircaloy structures (cladding and fuel canisters) and high burnup fuel with enrichments of 8% and 10% for PWRs and BWRs, respectively. Representative radionuclide releases are generated for this accident tolerant fuel concept by applying non-parametric bootstrap methods to MELCOR simulation results. Accident scenarios considered in this analysis include principle contributors to historical core damage frequency estimates for a range of nuclear reactor technologies representative of the operating U.S.A. fleet of nuclear reactors.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quantifying uncertainty in analysis of shockless dynamic compression experiments on platinum. II. Bayesian model calibration

Dynamic shockless compression experiments provide the ability to explore material behavior at extreme pressures but relatively low temperatures. Typically, the data from these types of experiments are interpreted through an analytic method called Lagrangian analysis. Here, in this work, alternative analysis methods are explored using modern statistical methods. Specifically, Bayesian model calibration is applied to a new set of platinum data shocklessly compressed to 570 GPa. Several platinum equation-of-state models are evaluated, including traditional parametric forms as well as a novel non-parametric model concept. The results are compared to those in Paper I obtained by inverse Lagrangian analysis. The comparisons suggest that Bayesian calibration is not only a viable framework for precise quantification of the compression path, but also reveals insights pertaining to trade-offs surrounding model form selection, sensitivities of the relevant experimental uncertainties, and assumptions and limitations within Lagrangian analysis. The non-parametric model method, in particular, is found to give precise unbiased results and is expected to be useful over a wide range of applications. The calibration results in estimates of the platinum principal isentrope over the full range of experimental pressures to a standard error of 1.6%, which extends the results from Paper I while maintaining the high precision required for the platinum pressure standard.

Brown, Justin Lee↗

Model-Free Probabilistic Forecasting of Nodal Voltages in Distribution Systems

As the penetration of distributed energy resources (DERs) into distribution systems increases, so does the interest in forecasting relevant system variables to help mitigate the associated challenges. One such challenge is the more frequent occurrence of excessive voltages in distribution systems with higher shares of DERs. Accurate and reliable estimates together with forecasts of system states (i.e., nodal voltages) will therefore play a key role in improving the utilization of these variable and uncertain sources while mitigating potential operational risks. Whilst recent literature has explored machine learning (ML) methods for voltage estimation and their extrapolation for a short-time period into the future, few have taken uncertainty quantification into account, and these methods have not yet been translated into operations. This paper discusses the advantages offered by probabilistic voltage forecasts and proposes a non-parametric Bayesian method suitable for forecasting nodal voltages at short-term time horizons while accounting for uncertainties in load and distributed photovoltaic (PV) generation. We demonstrate the value of the proposed Gaussian process (GP) model for a case study using historical forecasts and observation data.

distribution system↗

An Empirical Quantile Estimation Approach for Chance-Constrained Nonlinear Optimization Problems

We investigate an empirical quantile estimation approach to solve chance-constrained nonlinear optimization problems. Our approach is based on the reformulation of the chance constraint as an equivalent quantile constraint to provide stronger signals on the gradient. In this approach, the value of the quantile function is estimated empirically from samples drawn from the random parameters, and the gradient of the quantile function is estimated via a finite-difference approximation on top of the quantile-function-value estimation. We establish a convergence theory of this approach within the framework of an augmented Lagrangian method for solving general nonlinear constrained optimization problems. The foundation of the convergence analysis is a concentration property of the empirical quantile process, and the analysis is divided based on whether or not the quantile function is differentiable. In contrast to the sampling-and-smoothing approach used in the literature, the method developed in this paper does not involve any smoothing function and hence the quantile-function gradient approximation is easier to implement and there are less accuracy-control parameters to tune. Furthermore, we demonstrate the effectiveness of this approach and compare it with a smoothing method for the quantile-gradient estimation. Numerical investigation shows that the two approaches are competitive for certain problem instances.

Applied Probability↗

Parametric and Nonparametric Models of U.S. Cost Overruns for Nuclear Power Plants

This study presents new data-driven models to estimate the effect of capacity on the percentage of cost overruns in the United States for nuclear power plant construction projects before and after the Three Mile Island accident. Parametric and nonparametric models have been developed that describe the significant shifts in nuclear energy costs during the dynamic environment. Employing a contemporary descriptive methodology and a quantitative analysis, we furnish a comprehensive overview of the alterations in cost overrun distribution and show the changes observed in other pivotal metrics alongside cost overruns. Our emphasis lies in documenting the fluctuations in cost overruns alongside nuclear reactor capacity levels and the increase of the overnight capital costs to build nuclear reactors. Our results show that increasing the size of nuclear reactors is not a factor statistically significant to decrease the percentage of cost overruns, and the probit model results provide evidence that an increase in size increases the probability of having cost overruns larger than 100% (double the estimated cost). We also compare our findings to two other regions: Asia and Europe.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

Robust inference of the Galactic Centre gamma-ray excess spatial properties

ABSTRACT The gamma-ray Fermi-LAT Galactic Centre excess (GCE) has puzzled scientists for over 15 yr. Despite ongoing debates about its properties, and especially its spatial distribution, its nature remains elusive. We scrutinize how the estimated spatial morphology of this excess depends on models for the Galactic diffuse emission, focusing particularly on the extent to which the Galactic plane and point sources are masked. Our main aim is to compare a spherically symmetric morphology – potentially arising from the annihilation of dark matter (DM) particles – with a boxy morphology – expected if faint unresolved sources in the Galactic bulge dominate the excess emission. Recent claims favouring a DM-motivated template for the GCE are shown to rely on a specific Galactic bulge template, which performs worse than other templates for the Galactic bulge. We find that a non-parametric model of the Galactic bulge derived from the VISTA Variables in the Via Lactea survey results in a significantly better fit for the GCE than DM-motivated templates. This result is independent of whether a galprop-based model or a more non-parametric ring-based model is used to describe the diffuse Galactic emission. This conclusion remains true even when additional freedom is added in the background models, allowing for non-parametric modulation of the model components and substantially improving the fit quality. When adopted, optimized background models provide robust results in terms of preference for a boxy bulge morphology for the GCE, regardless of the mask applied to the Galactic plane.

Astronomy & Astrophysics↗

Unravelling the orbits of cluster galaxy populations according to their dominant gas ionization source

ABSTRACT We investigate the kinematical and dynamical properties of cluster galaxy populations classified according to their dominant source of gas ionization, namely: star-forming (SF) galaxies, optical active galactic nuclei (AGNs), mixed SF plus AGN ionization (transition objects, T), and quiescent (Q) galaxies. We stack 8892 member galaxies from 336 relaxed galaxy clusters to build an ensemble cluster and estimate the observed projected profiles of numerical density and velocity dispersion, $\sigma _P(R)$, of each galaxy population. The MAMPOSSt code and the Jeans equations inversion technique are used to constrain the velocity anisotropy profiles of the galaxy populations in both parametric and non-parametric ways. We find that Q (SF) galaxies display the lowest (highest) typical cluster-centric distances and velocity dispersion values. Transition galaxies are more concentrated and tend to exhibit lower velocity dispersion values than SF galaxies. Galaxies that host an optical AGN are as concentrated as Q galaxies but display velocity dispersion values similar to those of the SF population. MAMPOSSt is able to find equilibrium solutions that successfully recover the observed $\sigma _P(R)$ profile only for the Q, T, and AGN populations. We find that the orbits of all populations are consistent with isotropy in the inner regions, becoming increasingly radial with the distance from the cluster centre. These results suggest that Q galaxies are in equilibrium within their clusters, while SF galaxies have more recently arrived in the cluster environment. Finally, the T and AGN populations appear to be in an intermediate dynamical state between those of the SF and Q populations.

Valk, Greique A. (ORCID:0009000827731299)↗