Search NASA⌕ Search

SEARCH · Search NASA

Results for “probability and statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Bayesian event categorization matrix approach for explosion monitoring

Current efforts to correctly categorize natural events from suspected explosion sources with data that is collected by ground- or space-based sensors presents historical challenges that remain unaddressed by the Event Categorization Matrix (ECM) model. Smaller historical events (lower yield explosions) may have data available from fewer measurement techniques than are available today, and therefore, a historical event record can lack a complete set of discriminants. The covariance structures can also differ between such observations of event (source-type) categories. Both obstacles are problematic for the classic ECM model. Our work addresses this gap and presents a Bayesian update to the previous ECM model, termed the Bayesian Event Categorization Matrix model, which can be trained on partial observations and does not rely on a pooled covariance structure. We further augment the ECM model with Bayesian Decision Theory so that false negative or false positive rates of an event categorization can be reduced in an intuitive manner. To demonstrate improved categorization rates for the Bayesian Event Categorization Matrix model, we compare an array of Bayesian and classic models with multiple performance metrics using Monte Carlo experiments. We use both synthetic and real data. Our Bayesian models show consistent gains in overall accuracy and lower false negative rates relative to the classic ECM model. Here, we propose future avenues to improve Bayesian Event Categorization Matrix models’ decision making and predictive capability.

58 GEOSCIENCES↗

Seismic moment tensor classification using elliptical distribution functions on the hypersphere

Discrimination of underground explosions from naturally occurring earthquakes and other anthropogenic sources is one of the fundamental challenges of nuclear explosion monitoring. In an operational setting, the number of events that can be thoroughly investigated by analysts is limited by available resources. The capability to rapidly screen out events that can be robustly identified as not being explosions is, therefore, of great potential benefit. Nevertheless, possible mis-classification of explosions as earthquakes currently limits the use of screening methods for verification of test-ban treaties. Moment tensors provide a physics-based classification tool for the characterization of different seismic sources and have enabled the advent of new techniques for discriminating between earthquakes and explosions. Following normalization and projection of their six-degree vectors onto the hypersphere, existing screening approaches use spherically symmetric metrics to determine whether any new moment tensor may have been an explosion. Here, we show that populations of moment tensors for both earthquakes and explosions are anisotropically distributed on the hypersphere. Distributions possessing elliptical symmetry, such as the scaled von Mises–Fisher distribution, therefore provide a better description of these populations than the existing spherically symmetric models. We describe a method that uses these elliptical distributions in combination with a Bayesian classifier to achieve successful classification rates of 99 per cent for explosions and 98 per cent for earthquakes using existing catalogues of events from the western United States. The 1983 May 5 Crowdie underground nuclear test and 2018 July 20 DAG-1 deep-borehole chemical explosion are the only two explosions out of 140 that are incorrectly classified. Application of the method to the 2006–2017 nuclear tests in the Democratic People’s Republic of Korea yields 100 per cent identification rates and we provide a simple routine MTid for general usage. The approach provides a means to rapidly assess the likelihood of an event being an explosion and can be built into monitoring workflows that rely on simultaneously assessing multiple different discrimination metrics.

58 GEOSCIENCES↗

Surface coverage dynamics for reversible dissociative adsorption on finite linear lattices

Dissociative adsorption onto a surface introduces dynamic correlations between neighboring sites not found in non-dissociative absorption. We study surface coverage dynamics where reversible dissociative adsorption of dimers occurs on a finite linear lattice. We derive analytic expressions for the equilibrium surface coverage as a function of the number of reactive sites, N, and the ratio of the adsorption and desorption rates. Using these results, we characterize the finite size effect on the equilibrium surface coverage. For comparable N’s, the finite size effect is significantly larger when N is even than when N is odd. Moreover, as N increases, the size effect decays more slowly in the even case than in the odd case. The finite-size effect becomes significant when adsorption and desorption rates are considerably different. These finite-size effects are related to the number of accessible configurations in a finite system where the odd-even dependence arises from the limited number of accessible configurations in the even case. We confirm our analytical results with kinetic Monte Carlo simulations. We also analyze the surface-diffusion case where adsorbed atoms can hop into neighboring sites. As expected, the odd-even dependence disappears because more configurations are accessible in the even case due to surface diffusion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Search Approaches to Sampling-Based Motion Planning

In this paper, we present a novel formulation of traditional sampling-based motion planners as database-oracle structures that can be solved via quantum search algorithms. We consider two complementary scenarios: for simpler sparse environments, we formulate the Quantum Full Path Search Algorithm (q-FPS), which creates a superposition of full random path solutions, manipulates probability amplitudes with Quantum Amplitude Amplification (QAA), and quantum measures a single obstacle free full path solution. For dense unstructured environments, we formulate the Quantum Rapidly Exploring Random Tree algorithm, q-RRT, that creates quantum superpositions of possible parent-child connections, manipulates probability amplitudes with QAA, and quantum measures a single reachable state, which is added to a tree. As performance depends on the number of oracle calls and the probability of measuring good quantum states, we quantify how these errors factor into the probabilistic completeness properties of the algorithm. We then numerically estimate the expected number of database solutions to provide an approximation of the optimal number of oracle calls in the algorithm. We compare the q-RRT algorithm with a classical implementation and verify quadratic run-time speedup in the largest connected component of a 2D dense random lattice. We conclude by evaluating a proposed approach to limit the expected number of database solutions and thus limit the optimal number of oracle calls to a given number.

97 MATHEMATICS AND COMPUTING↗

Qualitative and Quantitative Evaluation for Representative Human Reliability Analysis Methods

The Korea Institute of Nuclear Safety (KINS) is the regulatory expert organization established by the Korean government to strengthen the nation’s technical capabilities relating to nuclear safety regulation. KINS oversees the technical aspects of nuclear safety regulation, including safety reviews, inspections, education, and safety research—all conducted based on technical knowledge and accumulated regulatory experience. In 2023, KINS requested that Idaho National Laboratory (INL) validates representative human reliability analysis (HRA) methods used throughout the world, thus affording KINS with a basis for determining an HRA method adequate for its domestic regulatory purposes. The present paper mainly examines INL’s efforts in this regard. The resulting INL study covered four representative HRA methods widely used by nuclear utilities and regulatory institutes. These methods were qualitatively evaluated by applying specific evaluation criteria and determining how well each method reflected critical HRA issues. For this assessment, INL benchmarked the Halden International HRA Empirical Study. Using the Halden empirical data, along with information on human failure events (HFEs), the present study employed the selected HRA methods to estimate human error probabilities (HEPs) for the HFEs. It also performed statistical analyses to compare the HEPs predicted via the HRA methods against those from the Halden empirical data.

99 - GENERAL AND MISCELLANEOUS↗

Model averaging approaches to data subset selection

Model averaging is a useful and robust method for dealing with model uncertainty in statistical analysis. Often, it is useful to consider data subset selection at the same time, in which model selection criteria are used to compare models across different subsets of the data. Two different criteria have been proposed in the literature for how the data subsets should be weighted. We compare the two criteria closely in a unified treatment based on the Kullback-Leibler divergence and conclude that one of them is subtly flawed and will tend to yield larger uncertainties due to loss of information. Here, analytical and numerical examples are provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Constraints on the Cosmic Expansion History from GWTC–3

We use 47 gravitational wave sources from the Third LIGO–Virgo–Kamioka Gravitational Wave Detector Gravitational Wave Transient Catalog (GWTC–3) to estimate the Hubble parameter H(z), including its current value, the Hubble constant H0. Each gravitational wave (GW) signal provides the luminosity distance to the source, and we estimate the corresponding redshift using two methods: the redshifted masses and a galaxy catalog. Using the binary black hole (BBH) redshifted masses, we simultaneously infer the source mass distribution and H(z). The source mass distribution displays a peak around 34 M ⊙ , followed by a drop-off. Assuming this mass scale does not evolve with the redshift results in a H(z) measurement, yielding H 0 = $68$$^{+12}_{–8}$km s –1 Mpc –1 (68% credible interval) when combined with the H 0 measurement from GW170817 and its electromagnetic counterpart. This represents an improvement of 17% with respect to the H 0 estimate from GWTC–1. The second method associates each GW event with its probable host galaxy in the catalog GLADE+, statistically marginalizing over the redshifts of each event's potential hosts. Assuming a fixed BBH population, we estimate a value of H 0 = $68$$^{+8}_{–6}$km s –1 Mpc –1 with the galaxy catalog method, an improvement of 42% with respect to our GWTC–1 result and 20% with respect to recent H 0 studies using GWTC–2 events. However, we show that this result is strongly impacted by assumptions about the BBH source mass distribution; the only event which is not strongly impacted by such assumptions (and is thus informative about H 0 ) is the well-localized event GW190814.

79 ASTRONOMY AND ASTROPHYSICS↗

A novel framework for increasing research transparency: Exploring the connection between diversity and innovation

A split sample/dual method research protocol is demonstrated to increase transparency while reducing the probability of false discovery. We apply the protocol to examine whether diversity in ownership teams increases or decreases the likelihood of a firm reporting a novel innovation using data from the 2018 United States Census Bureau’s Annual Business Survey. Transparency is increased in three ways: 1) all specification testing and identifying potentially productive models is done in an exploratory subsample that 2) preserves the validity of hypothesis test statistics fromde novoestimation in the holdout confirmatory sample with 3) all findings publicly documented in an earlier registered report and in this journal publication. Bayesian estimation procedures that leverage information from the exploratory stage included in the confirmatory stage estimation replace traditional frequentist null hypothesis significance testing. In addition to increasing statistical power by using information from the full sample, Bayesian methods directly estimate a probability distribution for the magnitude of an effect, allowing much richer inference. Estimated magnitudes of diversity along academic discipline, race, ethnicity, and foreign-born status dimensions are positively associated with innovation. A maximally diverse ownership team on these dimensions would be roughly six times more likely to report new-to-market innovation than a homophilic team.

Science & Technology - Other Topics↗

An update to the Sandia method for creating Typical Meteorological Years from a limited pool of calendar years

Typical Meteorological Years (TMYs) are essential for the efficient evaluation of energy system performance. Ideally, 30 years of weather data are required to generate TMYs, but significantly fewer years are typically available due to practical limitations. To address this issue, an update to the Sandia method was developed, referred to as the Argonne method, to create TMYs from a limited number of years. Furthermore, this method enhances candidate diversity by systematically shifting original candidate months forward or backward by specific days, creating an expanded pool of candidates. The effectiveness of the Argonne method was validated through statistical testing, comparison of monthly average weather parameters, and numerical simulations. The results demonstrate a high probability of identifying at least one shifted month whose cumulative distribution functions of weather parameters closely align with long-term distributions. In 67 % of all comparisons, the monthly average weather parameters in TMYs generated using the Argonne method exhibit better agreement with long-term averages than TMY3. Moreover, in 74 % of the 318 building simulation cases, the Argonne method outperforms TMY3 in estimating long-term average building heating and cooling demands. Therefore, the Argonne method effectively diversifies the candidate pool and produces typical years that provide more accurate estimations of long-term averages compared to TMY3 when only a limited pool of calendar years (10 years or fewer) is available.

Building energy modeling↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics↗

A global significance evaluation method using simulated events

In High-Energy Physics experiments it is often necessary to evaluate the global statistical significance of apparent resonances observed in invariant mass spectra. One approach to determining significance is to use simulated events to find the probability of a random fluctuation in the background mimicking a real signal. As a high school summer project, we demonstrate a method with Monte Carlo simulated events to evaluate the global significance of a potential resonance with some assumptions. This method for determining significance is general and can be applied, with appropriate modification, to other resonances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An implementation of neural simulation-based inference for parameter estimation in ATLAS

Neural simulation-based inference (NSBI) is a powerful class of machine-learning-based methods for statistical inference that naturally handles high-dimensional parameter estimation without the need to bin data into low-dimensional summary histograms. Such methods are promising for a range of measurements, including at the Large Hadron Collider, where no single observable may be optimal to scan over the entire theoretical phase space under consideration, or where binning data into histograms could result in a loss of sensitivity. This work develops a NSBI framework for statistical inference, using neural networks to estimate probability density ratios, which enables the application to a full-scale analysis. It incorporates a large number of systematic uncertainties, quantifies the uncertainty due to the finite number of events in training samples, develops a method to construct confidence intervals, and demonstrates a series of intermediate diagnostic checks that can be performed to validate the robustness of the method. As an example, the power and feasibility of the method are assessed on simulated data for a simplified version of an off-shell Higgs boson couplings measurement in the four-lepton final states. This approach represents an extension to the standard statistical methodology used by the experiments at the Large Hadron Collider, and can benefit many physics analyses.

frequentist statistics↗

Production of alternate realizations of DESI fiber assignment for unbiased clustering measurement in data and simulations

A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi & Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Computing the Critical Temperature of the Affine-Transformed $D=3$ Ising Model Using Masked Autoregressive Flow

The simple Ising model provides a rich environment to build and study lattice field theories. As part of an ongoing project to construct a conformal field theory (CFT) on an arbitrarily curved manifold, in this work we develop methods to measure the critical temperature $β_c$ of the affine-transformed Ising model on the face-centered cubic (FCC) lattice. The main challenge in this endeavor is finding a computationally efficient and accurate method of interpolating and extrapolating Monte Carlo observables with respect to coupling coefficients and temperature. Herein, we compare two such methods. A traditional statistical approach uses the multiple histogram (MH) method, while a newer machine learning approach uses a masked autoregressive flow (MAF) to estimate the underlying probability density function of a set of observables. While the MH method is specifically designed to interpolate and extrapolate Monte Carlo observables, we find that MAF is a viable alternative for measuring $β_c$ with a computational cost that scales more favorably. Furthermore, we comment on additional advantages of MAF relevant to our work, such as extrapolating in system volume.

Svenson, Kai [Texas U.]↗

Generating high-resolution total canopy SIF emission from TROPOMI data: Algorithm and application

Solar-induced chlorophyll fluorescence (SIF) is a rapidly advancing front in modeling global terrestrial gross primary production (GPP). Canopy total SIF emissions (SIF total ) are mechanistically linked to the plant photosynthesis, and can be estimated from satellite observed SIF (SIF obs ) through radiative transfer modeling. However, the current satellite SIF obs and thus SIF total are available only at coarse spatial resolutions from several kilometers to tens of kilometers, inhibiting the application at fine spatial scales. Here, in this work, we proposed an algorithm to generate both global high-resolution SIF total (HSIF total ) and high-resolution SIF obs (HSIF obs ) at 1 km from low-resolution SIF obs (LSIF obs ) from the TROPOspheric Monitoring Instrument (TROPOMI), which has a spatial resolution at nadir of 3.5 km by 5.6–7 km. Our statistical method is based on the law of energy conservation and uses satellite derived fraction of absorbed photosynthetically active radiation, fluorescence efficiency, and the escape probability of fluorescence. We evaluated the accuracy of our HSIF total using the Orbiting Carbon Observatory-2 SIF (R 2 = 0.78). We found that the spatial resolution had clear effects on the relationship between HSIF total and GPP. We also compared HSIF total to 8-day averaged tower GPP from 135 flux sites and found that they were better correlated when HSIF total was averaged over a 1-km radius around the tower than when averaged over a larger radius. Our study provided a unique high-resolution HSIF total product, which will advance the estimation of GPP by extrapolating site-level relationships to the global scale.

54 ENVIRONMENTAL SCIENCES↗