Search NASA⌕ Search

SEARCH · Search NASA

Results for “Probability and statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Experimental Validation of a Module Cell Cracking Model

The What's Cracking app can predict how changes in crystalline silicon photovoltaic (PV) module materials, design, and mounting affect its susceptibility for cell fracture under uniform loading. This work has experimentally validated the app. A set of commercial crystalline silicon PV modules was obtained for this study. The modules were uniformly loaded at three different mounting points, and their subsequent cell fractures were recorded. A large sample size allowed for the development of an experimental statistical model for cell fracture. Here, the comparison of the experiment to predictions from the app is in excellent agreement. Both experimental and modeling results also elucidate how moving the module mounting points toward the center of the module increases the probability of cell fracture.

14 SOLAR ENERGY↗

Data-driven upper bounds and event attribution for unprecedented heatwaves

The last decade has seen numerous record-shattering heatwaves in all corners of the globe. In the aftermath of these devastating events, there is interest in identifying worst-case thresholds or upper bounds that quantify just how hot temperatures can become. Generalized Extreme Value theory provides a data-driven estimate of extreme thresholds; however, upper bounds may be exceeded by future events, which undermines attribution and planning for heatwave impacts. Here, we show how the occurrence and relative probability of observed yet unprecedented events that exceed a priori upper bound estimates, so-called “impossible” temperatures, has changed over time. We find that many unprecedented events are actually within data-driven upper bounds, but only when using modern spatial statistical methods. Furthermore, there are clear connections between anthropogenic forcing and the “impossibility” of the most extreme temperatures. Robust understanding of heatwave thresholds provides critical information about future record-breaking events and how their extremity relates to historical measurements.

54 ENVIRONMENTAL SCIENCES↗

Turbulence statistical analysis of the L-H transition and RMPs in KSTAR

Here, we investigate the turbulence statistics associated with low-to-high confinement (L-H) transitions and externally applied resonant magnetic perturbations (RMPs) in KSTAR. Time-series fluctuations of electron density n e , electron temperature T e , and the time derivative of the poloidal magnetic field dB θ /dt (Mirnov coils) are analysed using information-geometric measures (information rate Γ and information length $\mathcal{L}$ = ∫ Γ dt), together with kurtosis κ and variance σ 2 . In low-density upper single-null plasmas (n e ~ 1.2 x 10 19 m -3 ), a ~80 kHz magnetic mode coupling n e , T e , dB θ /dt and emerges prior to the L-H transition and persists into the edge-localised modes H-mode. Edge-localised RMPs (ERMPs) suppress this coherent mode but enhance intermittency, producing frequent bursts that abruptly reshape the time-dependent probability density functions (PDFs) and generate large spikes in Γ (with smaller changes in κ), signalling ERMP-driven departures from quasi-stationarity. The impact of ERMPs on background fluctuation levels depends on density, radial location, and the fluctuating variable itself ($\tilde{n}$, $\tilde{T}$, $\dot{B}$ θ ), whereas $\mathcal{L}$ provides a robust, regime-agnostic measure of cumulative statistical reorganisation and spatial decorrelation. In particular, at low density we observe weaker coupling between $\tilde{n}$ and $\tilde{T}$, along with a tendency toward decreased radial correlation-most clearly for $\tilde{T}$-under ERMPs. Overall, information geometry cleanly captures intermittent events, quantifies non-equilibrium PDF evolution, and offers a compact, cross-diagnostic metric for assessing resonant magnetic perturbation effects on edge transport and correlation across densities, radial locations, and confinement states.

Kim, Eun-jin [Coventry Univ. (United Kingdom); Seo↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Combining Machine Learning and Comparative Effectiveness Methodology to Study Primary Care Pharmacotherapy Pathways for Veterans With Depression

Our objective is to demonstrate an innovative method combining machine learning with comparative effectiveness research techniques and to investigate a hitherto unstudied question about the effectiveness of common prescribing patterns. For Operation Enduring Freedom/Operation Iraqi Freedom veterans with major depressive disorder, we generate pharmacotherapy pathways (of antidepressants) using process mining and machine learning. We select the medication episodes that were started at subtherapeutic doses by the first assigned primary care physician and observe the paths that those medication episodes follow. Using 2-stage least squares, we test the effectiveness of starting at a low dose and staying low for longer versus ramping up fast while balancing observable and unobservable characteristics of patients and providers through instrumental variables. We leverage predetermined provider practice patterns as instruments. We collected outpatient pharmacy data for selective serotonin reuptake inhibitors and selective norepinephrine reuptake inhibitors, patient and provider characteristics (as control variables), and the instruments for our cohort. All data were extracted for the period between 2006 and 2020. There is a statistically significant positive effect (0.68, 95% CI 0.11–1.25) of “ramping up fast” on engagement in care. When we examine the effect of “ramping up slow”, we see an insignificant negative impact on engagement in care (−0.82, 95% CI −1.89 to 0.25). As expected, the probability of drop-out also seems to have a negative effect on engagement in care (−0.39, 95% CI −0.94 to 0.17). We further validate these results by testing with medication possession ratios calculated periodically as an alternative engagement in care metric. Our findings contradict the “Start low, go slow” adage, indicating that ramping up the dose of an antidepressant faster has a significantly positive effect on engagement in care for our population.

60 APPLIED LIFE SCIENCES↗

Enabling probabilistic learning on manifolds through double diffusion maps

Here, we present a generative learning framework for probabilistic sampling that extends Probabilistic Learning on Manifolds (PLoM), which is designed to generate statistically consistent realizations of a random vector in a finite-dimensional Euclidean space, informed by a (representative) set of observations. In its original form, PLoM constructs a reduced-order probabilistic model by combining three main components: (a) kernel density estimation to approximate the underlying probability measure, (b) Diffusion Maps to characterize the manifold of the data, and (c) a reduced-order Itô Stochastic Differential Equation (ISDE) to sample from the learned distribution. However, its sampling dynamics are posed in the ambient space and the retained number of reduced coordinates is chosen by projection-reconstruction error. In practice, this often (i) requires more coordinates than the data’s intrinsic dimension to achieve stable sampling and (ii) lacks a smooth, basis-independent lifting back to the data domain; moreover, standard Diffusion Maps emphasize harmonic eigenfunctions and can miss non-harmonic latent structure. We address these limitations by decoupling geometry learning from sampling: a first Diffusion Maps pass identifies non-harmonic coordinates on which we formulate a full-order ISDE directly in the latent space, while Double Diffusion Maps captures multiscale geometric features and Geometric Harmonics (GH) learns a smooth lifting map to the ambient variables that is independent of the particular diffusion basis. This hybrid design preserves the system’s dynamical richness with a compact geometric representation and enables principled out-of-sample inference. The effectiveness and robustness of the proposed method are illustrated through two numerical studies: one based on data generated from two-dimensional Hermite polynomial functions and another based on high-fidelity simulations of a detonation wave in a reactive flow.

Double diffusion maps↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

Statistical inference of anomalous thermal transport with uncertainty quantification for interpretive 2D SOL models

The critical task of inferring anomalous cross-field transport coefficients is addressed in simulations of boundary plasmas with fluid models. A workflow for parameter inference in the UEDGE fluid code is developed using Bayesian optimization with parallelized sampling and integrated uncertainty quantification. In this workflow, transport coefficients are inferred by maximizing their posterior probability distribution, which is generally multidimensional and non-Gaussian. Uncertainty quantification is integrated throughout the optimization within the Bayesian framework that combines diagnostic uncertainties and model limitations. As a concrete example, we infer the anomalous electron thermal diffusivity $\chi_\perp$ from an interpretive 2D model describing electron heat transport in the conduction-limited region with radiative power loss. The workflow is first benchmarked against synthetic data and then tested on H-, L-, and I-mode discharges to match their midplane temperature and divertor heat flux profiles. We demonstrate that the workflow efficiently infers diffusivity and its associated uncertainty, generating 2D profiles that match 1D measurements. Future efforts will focus on incorporating more complicated fluid models and analyzing transport coefficients inferred from a large database of experimental results.

Bayesian optimization↗

HFIR Steady State Heat Transfer Code (HSSHTC) Statistical Uncertainty Analysis

HSSHTC, the safety basis steady state TH code for HFIR, uses a highly conservative approach in which all input and calculation uncertainties are resolved simultaneously at their most limiting setting. This results in excessive conservatism which does not account for the high unlikelihood of such simultaneous worst-case conditions. The present study explores an alternative approach, BEPU, in which reasonable working assumptions for the probability distribution of each input uncertainty are used to determine a relationship between burnout power margin and core fuel failure probability. This was performed under a philosophy of perturbing uncertainty parameters already defined within the HSSHTC methodology while preserving the HSSHTC calculation approach and solution methodology itself. Based on the assumptions employed in this study, the BEPU approach resulted in a 0.29 increase in burnout power ratio (25 MW increase in burnout power) compared to the latest HSSHTC calculations of C-HFIR-2026-004. The study can be refined in the future by employing fuel fabrication data to provide more realistic input distributions. Future changes to the HSSHTC methodology would potentially allow a more comprehensive treatment of uncertainties which may further increase the burnout power ratio.

Wysocki, Aaron [ORNL] (ORCID:0000000222043779)↗

Optimal Zeno Dragging for Quantum Control: A Shortcut to Zeno with Action-Based Scheduling Optimization

The quantum Zeno effect asserts that quantum measurements inhibit simultaneous unitary dynamics when the “collapse” events are sufficiently strong and frequent. This applies in the limit of strong continuous measurement or dissipation. It is possible to implement a dissipative control that is known as “Zeno dragging” by dynamically varying the monitored observable, and hence also the eigenstates, which are attractors under Zeno dynamics. This is similar to adiabatic processes, in that the Zeno-dragging fidelity is highest when the rate of eigenstate change is slow compared to the measurement rate. We demonstrate here two theoretical methods for using such dynamics to achieve control of quantum systems. The first, which we shall refer to as “shortcut to Zeno,” is analogous to the shortcuts to adiabaticity (counterdiabatic driving) that are frequently used to accelerate unitary adiabatic evolution. In the second approach, we apply the Chantasri-Dressel-Jordan stochastic action [PRA 88, 042110 (2013)], and demonstrate that the extremal-probability readout paths derived from this are well suited to setting up a Pontryagin-style optimization of the Zeno-dragging schedule. A fundamental contribution of the latter approach is to show that an action suitable for measurement-driven control optimization can be derived quite generally from statistical arguments. Implementing these methods on the Zeno dragging of a qubit, we find that both approaches yield the same solution, namely, that the optimal control is a unitary that matches the motion of the Zeno-monitored eigenstate. We then show that such a solution can be more robust than a unitary-only operation and we comment on solvable generalizations of our qubit example embedded in larger systems. These methods open up new pathways toward systematically developing dynamic control of Zeno subspaces to realize dissipatively stabilized quantum operations. Published by the American Physical Society 2024

Physics↗

Huge ensembles – Part 2: Properties of a huge ensemble of hindcasts generated with spherical Fourier neural operators

Abstract. In Part 1, we created an ensemble based on spherical Fourier neural operators. As initial condition perturbations, we used bred vectors, and as model perturbations, we used multiple checkpoints trained independently from scratch. Based on diagnostics that assess the ensemble's physical fidelity, our ensemble has comparable performance to operational weather forecasting systems. However, it requires orders-of-magnitude fewer computational resources. Here in Part 2, we generate a huge ensemble (HENS), with 7424 members initialized each day of summer 2023. We enumerate the technical requirements for running huge ensembles at this scale. HENS precisely samples the tails of the forecast distribution and presents a detailed sampling of internal variability. HENS has two primary applications: (1) as a large dataset with which to study the statistics and drivers of extreme weather and (2) as a weather forecasting system. For extreme climate statistics, HENS samples events 4σ away from the ensemble mean. At each grid cell, HENS increases the skill of the most accurate ensemble member and enhances coverage of possible future trajectories. As a weather forecasting model, HENS issues extreme weather forecasts with better uncertainty quantification. It also reduces the probability of outlier events, in which the verification value lies outside the ensemble forecast distribution.

Mahesh, Ankur↗

First 𝛽-Delayed Two-Neutron Spectroscopy of the 𝑟-Process Nucleus 134 In and Observation of the 𝑖 13/2 Single-Particle Neutron State in 133 Sn

This manuscript reports on the direct observation of a 𝛽-delayed two-neutron emission in a study of 134 In at the ISOLDE Decay Station using neutron spectroscopy. We also report on the first measurement in 𝛽 − decay of the long-sought 13/2 + excited state in 133 Sn, attributed to be the neutron single-particle 𝑖 13/2 orbital. The observation of sequential neutron emission is used to extract the relative population of the 𝑖 13/2 state, which was found to be much smaller than the predictions of the statistical model. The experiment was possible because of the innovative use of a neutron array with neutron discrimination and interaction tracking capabilities. This is the first study of the details of the two-neutron emission for a nucleus, which belongs to the 𝑟-process path. Understanding 𝛽-delayed two-neutron emission probabilities is essential to validate models used in astrophysical 𝑟-process nucleosynthesis calculations. Observing two-neutron emissions in 𝛽 − decay paves the way for new experiments to study energy and angular correlations for 𝛽-delayed multineutron emitters.

Beta decay↗

Clustering and Cliques in Preferential Attachment Random Graphs with Edge Insertion

In this paper, we investigate the global clustering coefficient (a.k.a transitivity) and clique number of graphs generated by a preferential attachment random graph model with an additional feature of allowing edge connections between existing vertices. Specifically, at each time step t, either a new vertex is added with probability f(t), or an edge is added between two existing vertices with probability 1 – f(t). We establish concentration inequalities for the global clustering and clique number of the resulting graphs under the assumption that f(t) is a regularly varying function at infinity with index of regular variation –$\gamma$, where $\gamma$ $\in$ [0, 1). Finally, we also demonstrate an inverse relation between these two statistics: the clique number is essentially the reciprocal of the global clustering coefficient.

97 MATHEMATICS AND COMPUTING↗

Field testing and validation of a low-cost MPC for demand flexibility for grid-interactive K-12 schools

K-12 school buildings account for the highest energy consumption within the public sector. Implementing advanced HVAC controls in grid-interactive K-12 schools could bring substantial economic advantages and grid flexibility. Our previous study demonstrated that a low-cost model predictive control (MPC) solution, which coordinates multiple packaged units, can enable demand flexibility without major hardware upgrades. However, a significant gap remains between academic pilots and market-ready scalable solutions. This paper extends the previous single-site pilot to a multi-site demonstration involving three school campuses (95 total units) through a commercial technology transfer process. Addressing the challenge of verifying performance with sparse field data, we present a new statistical approach using Bayesian methods to estimate the MPC’s effect on peak demand. Unlike traditional methods, this approach robustly quantifies uncertainty in non-normal, limited datasets. The results confirm the solution’s replicability, achieving a 21.6–38.9% reduction in HVAC peak demand (10.8–22.1% at the site-level) with > 98% probability across diverse locations. Finally, we document critical barriers to scaling software-as-a-service (SaaS) solutions–such as API instability and diverse legacy systems–and offer practical strategies to accelerate the commercial adoption of grid-interactive efficient buildings.

Ham, Sang Woo↗

Predicting the Evolution of Shallow Cumulus Clouds With a Lotka‐Volterra Like Model

Abstract In numerical weather prediction and climate models, boundary‐layer clouds are controlled by a wide range of subgrid‐scale processes. However, understanding the nature of these processes and their role in the evolution of the cloud size distribution as a whole has been elusive. To address this issue, we adopt a novel empirical framework from the field of population dynamics to model the evolution of cloud size statistics by using the shallow cumulus properties obtained from a large‐eddy simulation (LES). Our approach involves representing the cloud size distribution and the total cloud area using a revised Lotka‐Volterra model and ridge linear model, respectively. The physical interpretation of the total cloud area and coefficients obtained from the optimization of the models reveals three stages probably interpreted by dominant processes: the formation of new clouds, the growth of single clouds, and a steady state with organized transitions involving the growth and decay of multiple clouds. Furthermore, we showcase the potential of this framework to serve as a component of scale‐aware parameterizations of shallow‐convective clouds in atmospheric models.

54 ENVIRONMENTAL SCIENCES↗

Statistical evaluation of microscale stress conditions leading to void nucleation in the weak shock regime

Here, we investigate the heterogeneity of the stress state driven by anisotropic deformation response at the single crystal level through five statistical volume element (SVE) calculations of polycrystalline BCC tantalum. This work focuses on grain boundaries as a prominent material defect type prone to void nucleation based upon experimental observations of predominantly intergranular void nucleation in this material. The SVEs are constructed to be statistically representative of larger volumes of material and are meshed such that mean and standard deviation of grain size and orientation information is reconstructed. The computational meshes feature hexahedral (brick) elements and smooth conformal grain boundaries where significant stress concentration is known to occur, a tail effect of interest in the extreme events process of dynamic ductile damage. An existing micromechanical crystallographic plasticity model shown to capture the single crystal behavior of BCC tantalum well is used to perform the polycrystal calculations. The model includes representation of the non-Schmid effect of non-planar screw dislocation kinetics in tantalum. A three-dimensional stress state time profile predicted by damage modeling of a flyer plate impact experiment is applied as boundary conditions to each SVE. Resulting grain boundary stress state statistics are strongly non-Gaussian. Significant structural evolution is observed within the compressive hold before unloading into tension in the stress profile. Strong angular dependence of grain boundary traction magnitude with shock direction is observed. Non-Schmid effects continue to suggest their influence on propensity of microstructural defect types to nucleate voids. A general void nucleation criterion is proposed using probability theory. The general framework is specified to polycrystalline BCC tantalum in the weak shock regime to include the SVE calculations and literature molecular dynamics calculations of grain boundary void nucleation strength. Probability density functions (PDFs) are used to describe the interaction between the local stress state heterogeneity and the distributed grain boundary void nucleation strength state. A causation entropy maximization procedure removes the requirement for ad hoc selection of a PDF functional form and provides a rigorous procedure for data-based PDF determination. The resulting physically informed PDF describes the spatial appearance frequency of nucleated voids as a function of applied macroscale pressure. Lower length scale physics are thus packaged in a precise and computationally efficient way to provide computational plasticity insight to macroscale dynamic ductile damage models.

36 MATERIALS SCIENCE↗

Breaking the curse of dimensionality: Solving configurational integrals for crystalline solids by tensor networks

Accurately evaluating configurational integrals for dense solids remains a central and difficult challenge in the statistical mechanics of condensed systems. Here, we present a tensor network approach that reformulates the high-dimensional configurational integral for identical-particle crystals into a sequence of computationally efficient summations. We represent the integrand as a high-dimensional tensor and apply tensor-train (TT) decomposition together with a custom TT-cross interpolation. This approach circumvents the need to explicitly construct the full tensor. We introduce tailored rank-1 and rank-2 schemes optimized for sharply peaked Boltzmann probability densities, typical for identical-particle crystals. When applied to the calculation of internal energy and pressure-temperature curves for crystalline Cu and Ar at high (GPa) pressures, as well as the alpha-to-beta phase transition diagram of Sn, our method accurately reproduces molecular dynamics simulation results using tight-binding, machine learning, hierarchical interacting particle–neural network, and modified embedded atom method potentials,all within seconds of computation time.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗