Search NASA⌕ Search

SEARCH · Search NASA

Results for “Non-parametric estimation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Poisson tensor completion non-parametric differential entropy estimator

We introduce the Poisson tensor completion (PTC) estimator, a non-parametric differential entropy estimator. The PTC estimator leverages inter-sample relationships to compute a low-rank Poisson tensor decomposition of the frequency histogram. Our crucial observation is that the histogram bins are an instance of a space partitioning of counts and thus can be identified with a spatial Poisson process. The Poisson tensor decomposition leads to a completion of the intensity measure over all bins—including those containing few to no samples—and leads to our proposed PTC differential entropy estimator. A Poisson tensor decomposition models the underlying distribution of the count data and guarantees non-negative estimated values and so can be safely used directly in entropy estimation. Our estimator is the first tensor-based estimator that exploits the underlying spatial Poisson process related to the histogram explicitly when estimating the probability density with low-rank tensor decompositions for the purpose of tensor completion. Furthermore, we demonstrate that our PTC estimator is a substantial improvement over standard histogram-based estimators for sub-Gaussian probability distributions because of the concentration of norm phenomenon.

42 ENGINEERING↗

Out-of-distribution detection with non-parametric density estimation for models predicting processing history of uranium ore concentrates

The rapid advancement in machine learning (ML) and computer vision (CV) coincides with the growth of interest in deploying these ML/CV models in numerous fields from medicine to social science. Similar to those areas, we have witnessed a great number of works in materials science employing ML/CV models – neural networks in particular – in their studies in recent years. These models have proven to obtain accurate performance in various tasks. However, these models struggle to attain a similar performance when encountering test samples coming from a distribution that is different from the training set. More importantly, they fail without providing any warning to the users. Therefore, we propose a framework for detecting out-of-distribution (OOD) samples to alert users when a human intervention might be necessary in this work. Specifically, we explore the use of a non-parametric density estimation method to detect OOD samples. Here, we assess OOD detection capability of the proposed framework on ML models developed for categorizing precipitation routes of U 3 O 8 when encountering OOD datasets that contain samples (1) undergone different imaging acquisition process, (2) undergone different material synthesis process, and (3) different materials than ID set. Through those experiments, we achieve an average area under the receiver operating characteristic (AUROC) of at least 91% on average in detecting OOD samples. With minimal overhead cost and superior performance, the proposed framework enables a reliable and safe system when deploying in real-world scenarios.

Convolutional neural networks↗

Non-Parametric Collision Probability for Low-Velocity Encounters

An implicit, but not necessarily obvious, assumption in all of the current techniques for assessing satellite collision probability is that the relative position uncertainty is perfectly correlated in time. If there is any mis-modeling of the dynamics in the propagation of the relative position error covariance matrix, time-wise de-correlation of the uncertainty will increase the probability of collision over a given time interval. The paper gives some examples that illustrate this point. This paper argues that, for the present, Monte Carlo analysis is the best available tool for handling low-velocity encounters, and suggests some techniques for addressing the issues just described. One proposal is for the use of a non-parametric technique that is widely used in actuarial and medical studies. The other suggestion is that accurate process noise models be used in the Monte Carlo trials to which the non-parametric estimate is applied. A further contribution of this paper is a description of how the time-wise decorrelation of uncertainty increases the probability of collision.

Carpenter, J. Russell↗

A new approach to evaluate gamma-ray measurements

Misunderstandings about the term random samples its implications may easily arise. Conditions under which the phases, obtained from arrival times, do not form a random sample and the dangers involved are discussed. Watson's U sup 2 test for uniformity is recommended for light curves with duty cycles larger than 10%. Under certain conditions, non-parametric density estimation may be used to determine estimates of the true light curve and its parameters.

Dejager, O. C.↗

Quantitative assessment of the cataractogenic potential of very low doses of neutrons

We report on the prevalence and relative biological effectiveness (RBE) for various stages of lens opacification in rats induced by very low doses (2 to 250 mGy) of medium-energy (440 keV) neutrons, compared to those for X rays. Neutron doses were delivered either in a single fraction or in four separate fractions and the irradiated animals were followed for over 100 weeks. At the highest observed dose (250 mGy) and at early observation times, there was evidence of an inverse dose-rate effect; i.e., a fractionated exposure was more potent than a single exposure. Neutron RBEs relative to X rays were estimated using a non-parametric technique. The results were only weakly dependent on time postirradiation. At 30 weeks, for example, 80% confidence intervals for the RBE of acutely delivered neutrons relative to X rays were 8-16 at 250 mGy, 10-20 at 50 mGy, 50-100 at 10 mGy and 250-500 at 2 mGy. The results are consistent with the estimated neutron RBEs in Japanese A-bomb survivors, though broad confidence bounds are present in the Japanese results. Our findings are also consistent with data reported earlier for cataractogenesis induced by heavy ions in rats, mice, and rabbits. We conclude from these results that, at very low doses (<10 mGy), the RBE for neutron-induced cataractogenesis is considerably larger than the RBE of 20 commonly used, and use of a significantly larger value for calculating equivalent dose would be prudent.

NASA Discipline Radiation Health↗

Quantiles, parametric-select density estimation, and bi-information parameter estimators

A quantile-based approach to statistical analysis and probability modeling of data is presented which formulates statistical inference problems as functional inference problems in which the parameters to be estimated are density functions. Density estimators can be non-parametric (computed independently of model identified) or parametric-select (approximated by finite parametric models that can provide standard models whose fit can be tested). Exponential models and autoregressive models are approximating densities which can be justified as maximum entropy for respectively the entropy of a probability density and the entropy of a quantile density. Applications of these ideas are outlined to the problems of modeling: (1) univariate data; (2) bivariate data and tests for independence; and (3) two samples and likelihood ratios. It is proposed that bi-information estimation of a density function can be developed by analogy to the problem of identification of regression models.

Parzen, E.↗

Comparison of Two Methods for Estimating the Sampling-Related Uncertainty of Satellite Rainfall Averages Based on a Large Radar Data Set

The uncertainty of rainfall estimated from averages of discrete samples collected by a satellite is assessed using a multi-year radar data set covering a large portion of the United States. The sampling-related uncertainty of rainfall estimates is evaluated for all combinations of 100 km, 200 km, and 500 km space domains, 1 day, 5 day, and 30 day rainfall accumulations, and regular sampling time intervals of 1 h, 3 h, 6 h, 8 h, and 12 h. These extensive analyses are combined to characterize the sampling uncertainty as a function of space and time domain, sampling frequency, and rainfall characteristics by means of a simple scaling law. Moreover, it is shown that both parametric and non-parametric statistical techniques of estimating the sampling uncertainty produce comparable results. Sampling uncertainty estimates, however, do depend on the choice of technique for obtaining them. They can also vary considerably from case to case, reflecting the great variability of natural rainfall, and should therefore be expressed in probabilistic terms. Rainfall calibration errors are shown to affect comparison of results obtained by studies based on data from different climate regions and/or observation platforms.

Lau, William K. M.↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗

Decision Boundary Feature Extraction for Nonparametric Classification

Feature extraction has long been an important topic in pattern recognition. Although many authors have studied feature extraction for parametric classifiers, relatively few feature extraction algorithms are available for nonparametric classifiers. A new feature extraction algorithm based on decision boundaries for nonparametric classifiers is proposed. It is noted that feature extraction for pattern recognition is equivalent to retaining 'discriminantly informative features' and a discriminantly informative feature is related to the decision boundary. Since nonparametric classifiers do not define decision boundaries in analytic form, the decision boundary and normal vectors must be estimated numerically. A procedure to extract discriminantly informative features based on a decision boundary for non-parametric classification is proposed. Experiments show that the proposed algorithm finds effective features for the nonparametric classifier with Parzen density estimation.

Lee, Chulhee↗

Accurate Biomass Estimation via Bayesian Adaptive Sampling

The following concepts were introduced: a) Bayesian adaptive sampling for solving biomass estimation; b) Characterization of MISR Rahman model parameters conditioned upon MODIS landcover. c) Rigorous non-parametric Bayesian approach to analytic mixture model determination. d) Unique U.S. asset for science product validation and verification.

Wheeler, Kevin R.↗

Parametric and Nonparametric Models of U.S. Cost Overruns for Nuclear Power Plants

This study presents new data-driven models to estimate the effect of capacity on the percentage of cost overruns in the United States for nuclear power plant construction projects before and after the Three Mile Island accident. Parametric and nonparametric models have been developed that describe the significant shifts in nuclear energy costs during the dynamic environment. Employing a contemporary descriptive methodology and a quantitative analysis, we furnish a comprehensive overview of the alterations in cost overrun distribution and show the changes observed in other pivotal metrics alongside cost overruns. Our emphasis lies in documenting the fluctuations in cost overruns alongside nuclear reactor capacity levels and the increase of the overnight capital costs to build nuclear reactors. Our results show that increasing the size of nuclear reactors is not a factor statistically significant to decrease the percentage of cost overruns, and the probit model results provide evidence that an increase in size increases the probability of having cost overruns larger than 100% (double the estimated cost). We also compare our findings to two other regions: Asia and Europe.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

Modeling Fire Severity in Black Spruce Stands in the Alaskan Boreal Forest Using Spectral and Non-Spectral Geospatial Data

Biomass burning in the Alaskan interior is already a major disturbance and source of carbon emissions, and is likely to increase in response to the warming and drying predicted for the future climate. In addition to quantifying changes to the spatial and temporal patterns of burned areas, observing variations in severity is the key to studying the impact of changes to the fire regime on carbon cycling, energy budgets, and post-fire succession. Remote sensing indices of fire severity have not consistently been well-correlated with in situ observations of important severity characteristics in Alaskan black spruce stands, including depth of burning of the surface organic layer. The incorporation of ancillary data such as in situ observations and GIS layers with spectral data from Landsat TM/ETM+ greatly improved efforts to map the reduction of the organic layer in burned black spruce stands. Using a regression tree approach, the R2 of the organic layer depth reduction models was 0.60 and 0.55 (pb0.01) for relative and absolute depth reduction, respectively. All of the independent variables used by the regression tree to estimate burn depth can be obtained independently of field observations. Implementation of a gradient boosting algorithm improved the R2 to 0.80 and 0.79 (pb0.01) for absolute and relative organic layer depth reduction, respectively. Independent variables used in the regression tree model of burn depth included topographic position, remote sensing indices related to soil and vegetation characteristics, timing of the fire event, and meteorological data. Post-fire organic layer depth characteristics are determined for a large (N200,000 ha) fire to identify areas that are potentially vulnerable to a shift in post-fire succession. This application showed that 12% of this fire event experienced fire severe enough to support a change in post-fire succession. We conclude that non-parametric models and ancillary data are useful in the modeling of the surface organic layer fire depth. Because quantitative differences in post-fire surface characteristics do not directly influence spectral properties, these modeling techniques provide better information than the use of remote sensing data alone.

Barrett, K.↗

Remote Sensing of Rain

The first problem addressed concerns passive-microwave rain retrievals. Most current approaches start by building off-line a cloud-model-derived database. Given data, the retrieval algorithms search the database for the microwave temperatures "closest" to the observed data, then after some fine-tuning (performed in different ways by different implementations) the rain is estimated to be that which corresponds to the selected (and fine-tuned) set of database temperatures. These approaches have three drawbacks: they cannot properly take into account the ambiguities which arise from the fact that several rain scenarios can produce the same observed temperatures; they are quite inefficient since they require manipulating a large database along with often complex "fine-tuning" procedures; and they cannot refine their estimates if additional data is available. This past year we have derived closed formulae relating observed microwave brightness temperatures, T(sub b), and the underlying rain rates, R: average T(sub b) =f (rain) and average rain = g (T(sub b)), along with the corresponding covariance matrices. These results are sufficient to describe the conditional probabilities p(R/T(sub b)) and p(T(sub b)/R) to second order. Progress has also been made towards deriving a robust description of the rain drop size distribution (DSD). The widespread approach consisting in parameterizing the DSD as a gamma-distribution in terms of the drop diameter D suffers from the facts that, in reality, the DSD is not a smooth function of D and that the largely arbitrary Gamma model imposes unintended behavior, which has implications on any quantities derived from the DSD model. We have therefore developed a non-parametric yet practical description of the DSD, which is particularly well-suited for use in remote-sensing applications. The diagram on the left shows a comparison between an actual DSD sample and the truncated non-parametric representation. One figure shows the relation between radar reflectivity and rain rate derived using this representation. Validation of the Tropical Rainfall Measuring Mission (TRMM) radar-radiometer combined R and DSD algorithm is underway. This algorithm was designed to make optimal use of the instantaneous reflectivity profiles measured by the TRMM radar and the microwave brightness temperatures measured by the TRMM passive radiometer. So far, it appears to be the most reliable TRMM rain algorithm.

Haddad, Ziad S.↗

Unravelling the orbits of cluster galaxy populations according to their dominant gas ionization source

ABSTRACT We investigate the kinematical and dynamical properties of cluster galaxy populations classified according to their dominant source of gas ionization, namely: star-forming (SF) galaxies, optical active galactic nuclei (AGNs), mixed SF plus AGN ionization (transition objects, T), and quiescent (Q) galaxies. We stack 8892 member galaxies from 336 relaxed galaxy clusters to build an ensemble cluster and estimate the observed projected profiles of numerical density and velocity dispersion, $\sigma _P(R)$, of each galaxy population. The MAMPOSSt code and the Jeans equations inversion technique are used to constrain the velocity anisotropy profiles of the galaxy populations in both parametric and non-parametric ways. We find that Q (SF) galaxies display the lowest (highest) typical cluster-centric distances and velocity dispersion values. Transition galaxies are more concentrated and tend to exhibit lower velocity dispersion values than SF galaxies. Galaxies that host an optical AGN are as concentrated as Q galaxies but display velocity dispersion values similar to those of the SF population. MAMPOSSt is able to find equilibrium solutions that successfully recover the observed $\sigma _P(R)$ profile only for the Q, T, and AGN populations. We find that the orbits of all populations are consistent with isotropy in the inner regions, becoming increasingly radial with the distance from the cluster centre. These results suggest that Q galaxies are in equilibrium within their clusters, while SF galaxies have more recently arrived in the cluster environment. Finally, the T and AGN populations appear to be in an intermediate dynamical state between those of the SF and Q populations.

Valk, Greique A. (ORCID:0009000827731299)↗

Decision boundary feature selection for non-parametric classifier

Feature selection has been one of the most important topics in pattern recognition. Although many authors have studied feature selection for parametric classifiers, few algorithms are available for feature selection for nonparametric classifiers. In this paper we propose a new feature selection algorithm based on decision boundaries for nonparametric classifiers. We first note that feature selection for pattern recognition is equivalent to retaining 'discriminantly informative features', and a discriminantly informative feature is related to the decision boundary. A procedure to extract discriminantly informative features based on a decision boundary for nonparametric classification is proposed. Experiments show that the proposed algorithm finds effective features for the nonparametric classifier with Parzen density estimation.

Lee, Chulhee↗

Modeling and Visualizing Uncertainty in Continuous Variables Predicted using Remotely Sensed Data

The use of remotely sensed images to map continuous biophysical variables, such as those related to terrestrial vegetation amount, sea surface temperature, and many other targets of NASA s Earth Observing System (EOS), includes variable, parametric, positional, spatial support and structural sources of uncertainty. A complete description of uncertainty will lead to a probability distribution at each location, allowing the exploration of the spatial dimension of uncertainty, that is, where the field is not well quantified. To achieve this purpose, convenient visualization tools are required. We have produced such a tool, called PDFVis, that facilitates the display of probability density functions (pdfs) on a per-grid-cell basis. The density estimate from Monte-Carlo generated realizations is interactively displayed as well as parametric and non-parametric summaries of the pdf field (such as mean, median, quartiles, standard deviation, number of modes, and locations of modes). Shaded surface renderings of pdfs along a transect can also be projected onto a plane. This tool will become more useful as richer descriptions of spatial uncertainty become available.

Dungan, Jennifer L.↗

Software Reliability 2002

In FY01 we learned that hardware reliability models need substantial changes to account for differences in software, thus making software reliability measurements more effective, accurate, and easier to apply. These reliability models are generally based on familiar distributions or parametric methods. An obvious question is 'What new statistical and probability models can be developed using non-parametric and distribution-free methods instead of the traditional parametric method?" Two approaches to software reliability engineering appear somewhat promising. The first study, begin in FY01, is based in hardware reliability, a very well established science that has many aspects that can be applied to software. This research effort has investigated mathematical aspects of hardware reliability and has identified those applicable to software. Currently the research effort is applying and testing these approaches to software reliability measurement, These parametric models require much project data that may be difficult to apply and interpret. Projects at GSFC are often complex in both technology and schedules. Assessing and estimating reliability of the final system is extremely difficult when various subsystems are tested and completed long before others. Parametric and distribution free techniques may offer a new and accurate way of modeling failure time and other project data to provide earlier and more accurate estimates of system reliability.

Wallace, Dolores R.↗