Search NASA⌕ Search

SEARCH · Search NASA

Results for “confidence intervals”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Monte Carlo method for constructing confidence intervals with unconstrained and constrained nuisance parameters in the NOvA experiment

Measuring observables to constrain models using maximum-likelihood estimation is fundamental to many physics experiments. Wilks' theorem provides a simple way to construct confidence intervals on model parameters, but it only applies under certain conditions. These conditions, such as nested hypotheses and unbounded parameters, are often violated in neutrino oscillation measurements and other experimental scenarios. Monte Carlo methods can address these issues, albeit at increased computational cost. In the presence of nuisance parameters, however, the best way to implement a Monte Carlo method is ambiguous. Furthermore, this paper documents the method selected by the NOvA experiment, the profile construction. It presents the toy studies that informed the choice of method, details of its implementation, and tests performed to validate it. It also includes some practical considerations which may be of use to others choosing to use the profile construction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Profiled Feldman-Cousins Method for Confidence Interval Construction for the Nova 3-Flavor Oscillation Analysis

The small interaction cross-section of neutrinos makes experimental neutrino physics particularly responsive to technological advancements. A significant development leveraged by the NOvA experiment is large-scale parallel processing, enabling novel computational approaches to longstanding experimental challenges. Central to managing the resulting high-throughput data is NOvA’s implementation of the Freight Train model, designed for efficient data production and handling.This dissertation details the methodology and execution of the NOvA 2024 3-Flavor Oscillation Analysis, supported by a comprehensive dataset spanning ten years. It emphasizes frequentist results refined through the Feldman-Cousins (FC) technique, specifically addressing confidence interval corrections in parameter estimation. The computational intensity associated with Feldman-Cousins arises from extensive Monte Carlo simulations, which were substantially mitigated through parallel computing on the Perlmutter supercomputer at the National Energy Research Scientific Computing Center (NERSC), employing the MPI framework.To further enhance computational efficiency, an Importance Sampling method is introduced and evaluated, demonstrating significant potential to reduce complexity, particularly in exploring extreme parameter space regions. This thesis presents both the successful application of advanced computational resources and the development of sophisticated statistical techniques, aiming to enhance the precision and scope of neutrino oscillation analyses.

Dye ajdye11190@gmail.com, Andrew Joseph [Mississip↗

EPR-based uncertainty validation of the calculated external doses for population exposed in the urals region

Tooth enamel Electron Paramagnetic Resonance (EPR) spectroscopy was used as a method for external dosimetry in the territories contaminated in the 1950s by PA ‘Mayak’ (Urals region) to validate the mean dose estimates predicted by the Techa River Dosimetry System (TRDS). The purpose of this study is to validate the uncertainties of TRDS doses. Ninety percent confidence intervals (90% confidence interval, CI) of dose estimated with both methods were compared for 220 people. All data were grouped according to the width of 90%CI, viz.: (1) 90%CI of EPR-based dose ≤ 90%CI of TRDS prediction (38 cases); (2) 90%CI of EPR-based dose > 90%CI of TRDS prediction (182 cases). About 91% of 90%CIs overlap. In group 1, 100% cases overlap. In group 2, 80% of the cases were non-contradictive (the calculated 90%CI is completely within the measured one). Interval comparison of doses predicted retrospectively and estimated based on individual measurements are non-contradictory and demonstrate a good agreement.

61 RADIATION PROTECTION AND DOSIMETRY↗

Systematic study of projection biases in the weak lensing analysis of cosmic shear and the combination of galaxy clustering and galaxy-galaxy lensing

This paper presents the results of a systematic study of projection biases in the weak lensing analysis of cosmic shear and the combination of galaxy clustering and galaxy-galaxy lensing using data collected during the first year of running the Dark Energy Survey experiment. The study uses Lambda cold dark matter ( Λ CDM ) as the cosmological model and two-point correlation functions for the weak lensing (WL) analysis. The results in this paper show that, independent of the WL analysis, projection biases of more than 1 σ exist and are a function of the position of the true values of the parameters h , n s , Ω b , and Ω ν h 2 with respect to their prior probabilities. For cosmic shear, and the combination of galaxy clustering and galaxy-galaxy lensing, this study shows that the coverage probability of the 68.27% credible intervals ranges from as high as 93% to as low as 16% and that these credible intervals are inflated, on average, by 29% for cosmic shear and 20% for the combination of galaxy clustering and galaxy-galaxy lensing. The results of the study also show that, in six out of nine tested cases, the reduction in error bars obtained by transforming credible intervals into confidence intervals is equivalent to an increase in the amount of data by a factor of 3.

79 ASTRONOMY AND ASTROPHYSICS↗

Lung Cancer in the Mayak Workers Cohort: Risk Estimation and Uncertainty Analysis

The workers at the Mayak nuclear facility near Ozyorsk, Russia are a primary source of information about exposure to radiation at low-dose rates, since they were subject to protracted exposures to external gamma rays and to internal exposures from plutonium inhalation. Here we re-examine lung cancer mortality rates and assess the effects of external gamma and internal plutonium exposures using recently developed Monte Carlo dosimetry systems. Using individual lagged mean annual lung doses computed from the dose realizations, we fit excess relative risk (ERR) models to the lung cancer mortality data for the Mayak Workers Cohort using risk-modeling software. We then used the corrected information matrix (CIM) approach to widen the confidence intervals of ERR by taking into account the uncertainty in doses represented by multiple realizations from the Monte Carlo dosimetry systems. Findings of this work revealed that there were 930 lung cancer deaths during follow-up. Plutonium lung doses (but not gamma doses) were generally higher in the new dosimetry systems than those used in the previous analysis. This led to a reduction in the risk per unit dose compared to prior estimates. The estimated ERR/Gy for external gamma-ray exposure was 0.19 (95% CI: 0.07 to 0.31) for both sexes combined, while the ERR/Gy for internal exposures based on mean plutonium doses were 3.5 (95% CI: 2.3 to 4.6) and 8.9 (95% CI: 3.4 to 14) for males and females at attained age 60. Accounting for uncertainty in dose had little effect on the confidence intervals for the ERR associated with gamma-ray exposure, but had a marked impact on confidence intervals, particularly the upper bounds, for the effect of plutonium exposure [adjusted 95% CIs: 1.5 to 8.9 for males and 2.7 to 28 for females]. In conclusion, lung cancer rates increased significantly with both external gamma-ray and internal plutonium exposures. Accounting for the effects of dose uncertainty markedly increased the width of the confidence intervals for the plutonium dose response but had little impact on the external gamma dose effect estimate. Adjusting risk estimate confidence intervals using CIM provides a solution to the important problem of dose uncertainty. This work demonstrates, for the first time, that it is possible and practical to use our recently developed CIM method to make such adjustments in a large cohort study.

uncertainty, radiation risk estimation, Mayak Prod↗

Sensitivity and uncertainty of the IFR-1 BISON benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered for adoption by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This paper summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment.For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. Here, the resulting variations in benchmark predictions were normalized to calculate sensitivities. These sensitivities were then used to inform input selections for the UQ.The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of the benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because the confidence intervals associated with benchmark predictions were wide. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Sensitivity and Uncertainty of the IFR-1 BISON Benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This report summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment. For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. The resulting variations in benchmark predictions were normalized to calculate sensitivities. The strongest sensitivities were identified and used to inform input selections for the UQ. The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because of the wide confidence intervals associated with the benchmark predictions. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark development and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quantifying Uncertainty in All-to-All Estimates of Space Object Conjunction Probabilities using U-Statistics

Predicting space object conjunctions is inherently probabilistic due to initial state and orbit model uncertainty. A commonly considered Monte Carlo estimator of the conjunction probability is the ’all-to-all’ estimator. Given independent random samples of the trajectories of both objects, the estimator is the percentage of all pairs of trajectories that result in a conjunction. Intuitively, the all-to-all estimator is the best possible estimator of the conjunction probability since it considers all pairs of Monte Carlo samples. However, its distribution is not available in closed-form, which limits its use in practice and makes this intuition difficult to make rigorous. In this paper, the all-to-all estimator is identified as a U-statistic, which implies that it has several favorable properties. Specifically, the estimator is the minimum variance unbiased estimator of the conjunction probability and is asymptotically Gaussian distributed. An approximate confidence interval for the conjunction probability is obtained from an estimate of the asymptotic Gaussian distribution. We show how to efficiently compute the confidence interval and demonstrate that the interval has the nominal coverage level. The confidence intervals are also seen to be narrower than those based on the commonly-used each-to-each estimator. Furthermore, the all-to-all estimator is shown to allow different Monte Carlo sample sizes, whereas the each-to-each estimator requires equal sample sizes.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A bootstrapping approach to social media quantification

Abstract This work considers the use of classifiers in a downstream aggregation task estimating class proportions, such as estimating the percentage of reviews for a movie with positive sentiment. We derive the bias and variance of the class proportion estimator when taking classification error into account to determine how to best trade off different error types when tuning a classifier for these tasks. Additionally, we propose a method for constructing confidence intervals that correctly adjusts for classification error when estimating these statistics. We conduct experiments on four document classification tasks comparing our methods to prior approaches across classifier thresholds, sample sizes, and label distributions. Prior approaches have focused on providing the most accurate point estimate while this work focuses on the creation of correct confidence intervals that appropriately account for classifier error. Compared to the prior approaches, our methods provide lower error and more accurate confidence intervals.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Grid-based minimization at scale: Feldman-Cousins corrections for light sterile neutrino search

High Energy Physics (HEP) experiments generally employ sophisticated statistical methods to present results in searches of new physics. In the problem of searching for sterile neutrinos, likelihood ratio tests are applied to short-baseline neutrino oscillation experiments to construct confidence intervals for the parameters of interest. The test statistics of the form Δχ2 is often used to form the confidence intervals, however, this approach can lead to statistical inaccuracies due to the small signal rate in the region-of-interest. In this paper, we present a computational model for the computationally expensive Feldman-Cousins corrections to construct a statistically accurate confidence interval for neutrino oscillation analysis. The program performs a grid-based minimization over oscillation parameters and is written in C++. Our algorithms make use of vectorization through Eigen3, yielding a single-core speed-up of 350 compared to the original implementation, and achieve MPI data parallelism by employing DIY. We demonstrate the strong scaling of the application at High-Performance Computing (HPC) sites. We utilize HDF5 along with HighFive to write the results of the calculation to file.

Wospakrik, Marianette↗

Establishing performance metrics for quantitative non-targeted analysis: a demonstration using per- and polyfluoroalkyl substances

Abstract Non-targeted analysis (NTA) is an increasingly popular technique for characterizing undefined chemical analytes. Generating quantitative NTA (qNTA) concentration estimates requires the use of training data from calibration “surrogates,” which can yield diminished predictive performance relative to targeted analysis. To evaluate performance differences between targeted and qNTA approaches, we defined new metrics that convey predictive accuracy, uncertainty (using 95% inverse confidence intervals), and reliability (the extent to which confidence intervals contain true values). We calculated and examined these newly defined metrics across five quantitative approaches applied to a mixture of 29 per- and polyfluoroalkyl substances (PFAS). The quantitative approaches spanned a traditional targeted design using chemical-specific calibration curves to a generalizable qNTA design using bootstrap-sampled calibration values from “global” chemical surrogates. As expected, the targeted approaches performed best, with major benefits realized from matched calibration curves and internal standard correction. In comparison to the benchmark targeted approach, the most generalizable qNTA approach (using “global” surrogates) showed a decrease in accuracy by a factor of ~4, an increase in uncertainty by a factor of ~1000, and a decrease in reliability by ~5%, on average. Using “expert-selected” surrogates ( n = 3) instead of “global” surrogates ( n = 25) for qNTA yielded improvements in predictive accuracy (by ~1.5×) and uncertainty (by ~70×) but at the cost of further-reduced reliability (by ~5%). Overall, our results illustrate the utility of qNTA approaches for a subclass of emerging contaminants and present a framework on which to develop new approaches for more complex use cases. Graphical Abstract

Pu, Shirley (ORCID:0000000201223797)↗

Reference Correlations for the Density and Viscosity of Molten Alkali and Alkaline Earth Fluoride Salts

While there is a significant body of literature pertaining to thermophysical property measurements of molten salts, there is often a wide degree of variability among independent measurements of the same compounds. As such, the scientific community benefits greatly from an unbiased, independent assessment of duplicate datasets, so that reference correlations which describe these thermophysical properties as functions of temperature can be determined and then commonly used by researchers, scientists, and engineers. With regard to molten fluoride compounds, a significant time has elapsed since density and viscosity reference correlations have been determined; Janz conducted the most recent effort, in 1988, to provide reference correlations for the densities and viscosities of molten fluoride compounds via the National Standard Reference Data System coordinated by the National Bureau of Standards. Since then, new data have been published for molten fluoride compounds, and a new precedent has surfaced for putting forth reference correlations that involve fitting to multiple primary datasets. In this work, reference correlations are put forth for molten alkali and alkaline earth fluoride compounds in an effort to provide updated, improved correlations for general use. For molten alkali fluoride densities, estimated uncertainties with a 95% confidence interval are summarized as follows: LiF (0.63%), NaF (0.48%), KF (0.76%), RbF (0.93%), and CsF (0.75%). For molten alkaline earth fluoride densities, an estimated uncertainty was not able to be quantified for BeF 2 because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkaline earth fluorides: MgF 2 (1.5%), CaF 2 (0.92%), SrF 2 (1.6%), and BaF 2 (0.23%). For molten alkali fluoride viscosities, uncertainty was not able to be quantified for RbF and CsF because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkali fluorides: LiF (4.4%), NaF (3.0%), and KF (4.0%). For molten alkaline earth fluoride viscosities, limited consistent data resulted in the recommendation of single datasets (from literature) that are deemed to be the most trustworthy based on the quality of the underlying experimental studies.

Birri, A. [Oak Ridge National Laboratory (ORNL), O↗

Resilience Measurement Framework For Post-deployment Artificial Intelligence (ai) Integrated Systems

Resilience is largely defined as the ability to adapt or recover from adverse conditions, stresses, attacks, or compromises on systems that use or are enabled by digital resources. In Artificial Intelligence Management and Research for Advanced Networked Testbed Hub (AMARANTH), resilience is measured in the amount of time it took from the beginning of a testing period for the model to reach predictions outside of the original 95% confidence interval or using the Kullback-Leibler (KL) divergence theorem, the Population Stability Index (PSI), and traditional methods such as root mean squared error (RMSE) threshold. Artificial Intelligence (AI) model drift is of significant concern when deploying AI-integrated systems into critical and/or secure environments. Drift can impact resilience of the AI-integrated system post-deployment and requires consistent maintenance and upkeep to ensure the model is accurate and precise. To quantify model drift and predict the point when a model's drift becomes unacceptable, we describe using Kullback-Leibler (KL) divergence, Population Stability Index (PSI) and/or confidence interval width estimations to determine the point of failure and time to failure of a model post-deployment. Through simple code functions, the KL-divergence, PSI, confidence interval, and root mean squared (RMSE) point of failures can be used to derive when a model needs to be maintained as well as the impact of adversarial action through statistical means.

Yockey, Patience [Idaho National Laboratory (INL),↗

Drought–induced increase in tree mortality and corresponding decrease in the carbon sink capacity of Canada's boreal forests from 1970 to 2020

Canada's boreal forests, which occupy approximately 30% of boreal forests world wide, play an important role in the global carbon budget. However, there is lit tle quantitative information available regarding the spatiotemporal changes in the drought-induced tree mortality of Canada's boreal forests overall and their associ ated impacts on biomass carbon dynamics. Here, we develop spatiotemporally ex plicit estimates of drought-induced tree mortality and corresponding biomass carbon sink capacity changes in Canada's boreal forests from 1970 to 2020. We show that the average annual tree mortality rate is approximately 2.7%. Approximately 43% of Canada's boreal forests have experienced significantly increasing tree mortality trends (71% of which are located in the western region of the country), and these trends have accelerated since 2002. This increase in tree mortality has resulted in sig nificant biomass carbon losses at an approximate rate of 1.51±0.29 MgC ha -1 year -1 (95% confidence interval) with an approximate total loss of 0.46±0.09 PgC year -1 (95% confidence interval). Under the drought condition increases predicted for this century, the capacity of Canada's boreal forests to act as a carbon sink will be further reduced, potentially leading to a significant positive climate feedback effect

59 BASIC BIOLOGICAL SCIENCES↗

Too Good to Be True? Evaluation of Colonoscopy Sensitivity Assumptions Used in Policy Models

Background: Models can help guide colorectal cancer screening policy. Although models are carefully calibrated and validated, there is less scrutiny of assumptions about test performance. Methods: We examined the validity of the CRC-SPIN model and colonoscopy sensitivity assumptions. Standard sensitivity assumptions, consistent with published decision analyses, assume sensitivity equal to 0.75 for diminutive adenomas (<6 mm), 0.85 for small adenomas (6–10 mm), 0.95 for large adenomas (≥10 mm), and 0.95 for preclinical cancer. We also selected adenoma sensitivity that resulted in more accurate predictions. Targets were drawn from the Wheat Bran Fiber study. In the study, we examined how well the model predicted outcomes measured over a three-year follow-up period, including the number of adenomas detected, the size of the largest adenoma detected, and incident colorectal cancer. Results: Using standard sensitivity assumptions, the model predicted adenoma prevalence that was too low (42.5% versus 48.9% observed, with 95% confidence interval 45.3%–50.7%) and detection of too few large adenomas (5.1% versus 14.% observed, with 95% confidence interval 11.8%–17.4%). Predictions were close to targets when we set sensitivities to 0.20 for diminutive adenomas, 0.60 for small adenomas, 0.80 for 10- to 20-mm adenomas, and 0.98 for adenomas 20 mm and larger. Conclusions: Colonoscopy may be less accurate than currently assumed, especially for diminutive adenomas. Alternatively, the CRC-SPIN model may not accurately simulate onset and progression of adenomas in higher-risk populations. Impact: Misspecification of either colonoscopy sensitivity or disease progression in high-risk populations may affect the predicted effectiveness of colorectal cancer screening. When possible, decision analyses used to inform policy should address these uncertainties.

60 APPLIED LIFE SCIENCES↗

Direct measurement of the contribution of street lighting to satellite observations of nighttime light emissions from urban areas

Nighttime light emissions are increasing in most countries worldwide, but which types of lighting are responsible for the increase remains unknown. Also unknown is what fraction of outdoor light emissions and associated energy use are due to public light sources (i.e. streetlights) or various types of private light sources (e.g. advertising). Here we show that it is possible to measure the contribution of street lighting to nighttime satellite imagery using ‘smart city’ lighting infrastructure. The city of Tucson, USA, intentionally altered its streetlight output over 10 days, and we examined the change in emissions observed by satellite. We find that streetlights operated by the city are responsible for only 13% of the total radiance (in the 500–900 nm band) observed from Tucson from space after midnight (95% confidence interval 10–16%). If Tucson did not dim their streetlights after midnight, the contribution would be 18% (95% confidence interval 15–23%). When streetlights operated by other actors are included, the best estimates rise to 16% and 21%, respectively. Existing energy and lighting policy related to the sustainability of outdoor light use has mainly focused on street lighting. These results suggest an urgent need for consideration of other types of light sources in outdoor lighting policy.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗