Search NASASearch

SEARCH · Search NASA

Results for “Probabilities Mathematics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Practical Probabilistic Programming

Recent advances in probabilistic programming languages (PPLs) have provided the capability for exact inference: computing a closed-form probability distribution for a given probabilistic program. In particular, the new language Roulette uses a language oriented programming (LOP) approach, wherein analysts build new programming languages on top of a set of primitives provided by Roulette, which then translates these structures into a weighted model counting problem which can be solved by automated reasoning tools. However, because Roulette provides few convenience features, developing these new languages is challenging even for expert users. We developed a standard library of common probability functions for Roulette with the goal of improved usability. This included approximation of continuous probability density functions using discrete probability mass functions. We demonstrated this approach by modeling a cosmic ray striking a RAM controller. We found that Roulette provides a powerful interface for highly expressive probabilistic programs to be generated. In collaboration with the NNSA Advanced Simulation and Computing program, which resulted in development of a tool called Circulette, we were able to model complex circuits expressed in Verilog using probabilistic programs with an expressivity not previously possible. Our research question that motivated the development of a Roulette standard library was to determine whether non-experts could use a PPL to model relevant problems regarding radiation effects on microelectronics. This standard library improved the expressivity of Roulette by implementing common probability density functions, mathematical operators on distributions, and support for empirical distributions. While Roulette is a powerful modeling language, the untyped, LOP approach makes error messages difficult to understand and requires expert aid. We recommend further research on Roulette, especially with its error messages, to enable improved usability. At the same time, this project demonstrated that for users familiar with Roulette and the LOP approach, Roulette provides powerful new capabilities that can be integrated with other Sandia modeling capabilities.

97 MATHEMATICS AND COMPUTING

Conditional Pseudo-Reversible Normalizing Flow for Surrogate Modeling in Quantifying Uncertainty Propagation

We introduce a conditional pseudo-reversible normalizing flow (PR-NF) that directly learns conditional probability distributions from noisy physical models to efficiently quantify both forward and inverse uncertainty propagation. Traditional surrogate modeling approaches approximate only the deterministic component of physical models, requiring separate noise characterization and computationally expensive sampling methods for inverse problems. Here, in this work, we develop the conditional PR-NF model to directly learn and efficiently generate samples from the conditional probability density functions (PDFs). The training process utilizes dataset consisting of input-output pairs without requiring prior knowledge about the noise and the function. Once trained, our model efficiently generates samples from conditional PDFs for any input within the training domain. Moreover, the pseudo-reversibility feature allows for the use of fully connected neural network architectures, which simplifies the implementation and enables theoretical analysis. We provide a rigorous convergence analysis of the conditional PR-NF model, showing its ability to converge to the target conditional PDF using the Kullback−Leibler divergence. To demonstrate the effectiveness of our method, we apply it to several benchmark tests and a real-world geologic carbon storage problem.

97 MATHEMATICS AND COMPUTING

Ergodic Lagrangian dynamics in a superhero universe

We present a fictional scenario that, while undeniably whimsical, provides the foundation for a unique exercise in extended problem solving, physics analysis, and quantitative model development. Starting with the foundational premise of the Wild Cards shared-world superhero universe, we demonstrate how a variety of concepts appropriate to the advanced undergraduate level—ergodicity, functional analysis, Lagrangian mechanics, and the ever-important simplifying approximation—can be combined into a rich, coherent mathematical model. The goal of this case study is to develop a useful pedagogical exercise in exploring an open-ended research question that presents, at first glance, no clear path forward. Being both eclectic and lengthy, this exercise offers a unique way for students to apply their core physics and mathematics education. It is perhaps best used within a senior honors seminar or within a brief (e.g., January term) elective class.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Structural Aspects of Neutron Survival Probabilities

The neutron survival probability (and related quantities including probabilities of extinction and initiation) is a central element of the broader stochastic theory of neutron populations and finds application in fields including reactor start-up, analysis of reactor power bursts and criticality accidents, and safeguards. In a full neutron transport formulation, the equation governing the single-neutron survival probability is a backward or adjoint-like integro-partial differential equation with the added complexity of being highly nonlinear. Analogous formulations of this equation exist in the context of many approximate theories of neutron transport, with the point kinetics formulation having received significant theoretical attention since the 1940s. This work continues this tradition by providing a novel analysis of the single-neutron survival probability equation using the tools of boundary layer theory. The analysis reveals that the “fully dynamic” solution of the single-neutron survival probability equation—and some key probability distributions derived from it—may be cast as a singular perturbation around the underlying quasi-static single-neutron probability of initiation. In this perturbation solution, the expansion parameter is the ratio of the neutron generation time to a macroscopic time scale characterizing the overall system evolution; this interpretation illuminates some of the fundamental structural aspects of neutron survival phenomena.

97 MATHEMATICS AND COMPUTING

Stability and Convergence of Solutions to Stochastic Inverse Problems Using Approximate Probability Densities

Data-consistent inversion is designed to solve a class of stochastic inverse problems where the solution is a pullback of a probability measure specified on the outputs of a quantities of interest (QoI) map. Here, this work presents stability and convergence results for the case where finite QoI data result in an approximation of the solution as a density. Given their popularity in the literature, separate results are proven for three different approaches to measuring discrepancies between probability measures: f-divergences, integral probability metrics, and L p metrics. In the context of integral probability metrics, we also introduce a pullback probability metric that is well-suited for data-consistent inversion. This fills a theoretical gap in the convergence and stability results for data-consistent inversion that have mostly focused on convergence of solutions associated with approximate maps. Numerical results are included to illustrate key theoretical results with intuitive and reproducible test problems that include a demonstration of convergence in the measure-theoretic "almost" sense.

97 MATHEMATICS AND COMPUTING

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING

Asymptotic consistency of the WSINDy algorithm in the limit of continuum data

In this work we study the asymptotic consistency of the weak-form sparse identification of nonlinear dynamics algorithm (WSINDy) in the identification of differential equations from noisy samples of solutions. We prove that the WSINDy estimator is unconditionally asymptotically consistent for a wide class of models that includes the Navier–Stokes, Kuramoto–Sivashinsky and Sine–Gordon equations. We thus provide a mathematically rigorous explanation for the observed robustness to noise of weak-form equation learning. Conversely, we also show that, in general, the WSINDy estimator is only conditionally asymptotically consistent, yielding discovery of spurious terms with probability one if the noise level exceeds a critical threshold σ c . We provide explicit bounds on σ c in the case of Gaussian white noise and we explicitly characterize the spurious terms that arise in the case of trigonometric and/or polynomial libraries. Furthermore, we show that, if the data is suitably denoised (a simple moving average filter is sufficient), then asymptotic consistency is recovered for models with locally-Lipschitz, polynomial-growth nonlinearities. Our results reveal important aspects of weak-form equation learning, which may be used to improve future algorithms. We demonstrate our findings numerically using the Lorenz system, the cubic oscillator, a viscous Burgers-growth model and a Kuramoto–Sivashinsky-type high-order PDE.

asymptotic consistency

Representing Complex Systems as Graphs for Debugging and Predictive Maintenance-Preliminary Thoughts

Representing complex systems as graphs enables use of mathematical tools to identify faults or predict failures. Graph nodes correspond to individual modules or subsystems, and edges link coupled system parts. ‘Probes’ measure the node outputs, monitoring the system health for unexpected behavior. Assuming one cannot probe every point, within a system, the fault correlates to a region—not necessarily the specific location. Bayesian networks trained to understand fault patterns can accurately identify the source. The diagnostic tool described aides debugging by pinpointing system failure causes. For predictive maintenance, probe data develop probability distribution functions describing subsystem mean time to failure. Unit lifetime can be estimated through these probability distributions. Two approaches include using Bayesian classifiers to infer the system failure source and developing maintenance schedules by treating systems as collections of random variables. When failure behavior does not follow a closed form function, use of similarity models is proposed.

97 MATHEMATICS AND COMPUTING

Targeted Adaptive Design

Modern advanced manufacturing and advanced materials design often require searches of relatively high-dimensional process control parameter spaces for settings that result in optimal structure, property, and performance parameters. The mapping from the former to the latter must be determined from noisy experiments or from expensive simulations. Here, we abstract this problem to a mathematical framework in which an unknown function from a control space to a design space must be ascertained by means of expensive noisy measurements, which locate control settings generating desired design features within specified tolerances, with quantified uncertainty. We describe targeted adaptive design (TAD), a new algorithm that performs this sampling task efficiently. TAD creates a Gaussian process surrogate model of the unknown mapping at each iterative stage, proposing a new batch of control settings to sample experimentally and optimizing the updated expected log-predictive probability density of the target design. TAD either stops upon locating a solution with uncertainties that fit inside the tolerance box or uses a measure of expected future information to determine that the search space has been exhausted with no solution. TAD thus embodies the exploration-exploitation tension in a manner that recalls, but is essentially different from, Bayesian optimization and optimal experimental design.

97 MATHEMATICS AND COMPUTING

Mechanistic within-host mathematical model of inhalational anthrax

We present a mathematical model of the dynamics of Bacillus anthracis bacteria within the lymph nodes and blood of a host, following inhalation of an initial dose of spores. We also incorporate the dynamics of protective antigen, which is the binding component of the anthrax toxin produced by the bacteria. The model offers a mechanistic description of the early infection dynamics of inhalational anthrax, while its stochastic nature allows us to study the probabilities of different outcomes (for example, how likely it is that the infection will be cleared for a given inhaled dose of spores) in order to explain dose-response data for inhalational anthrax. The model is calibrated via a Bayesian approach, using in vivo data from New Zealand white rabbit and guinea pig infection studies, enabling within-host parameters to be estimated. We also leverage incubation-period data from the Sverdlovsk 1979 anthrax outbreak to show that the model can accurately describe human time-to-symptoms data under reasonable parameter regimes. Finally, we derive a simple approximate formula for the probability of symptom onset before time t, assuming that the number of inhaled spores has a Poisson distribution.

59 BASIC BIOLOGICAL SCIENCES

Exact block encoding of imaginary time evolution with universal quantum neural networks

We develop a constructive approach to generate quantum neural networks capable of representing the exact thermal states of all many-body qubit Hamiltonians. The Trotter expansion of the imaginary time propagator is implemented through an exact block encoding by means of a unitary, restricted Boltzmann machine architecture. Marginalization over the hidden-layer neurons (auxiliary qubits) creates the nonunitary action on the visible layer. Then, we introduce a unitary deep Boltzmann machine architecture in which the hidden-layer qubits are allowed to couple laterally to other hidden qubits. We prove that this wave-function is closed under the action of the imaginary time propagator and, more generally, can represent the action of a universal set of quantum gate operations. We provide analytic expressions for the coefficients for both architectures, thus enabling exact network representations of thermal states without stochastic optimization of the network parameters. In the limit of large imaginary time, the yields the ground state of the system. The number of qubits grows linearly with the number of interactions and total imaginary time for a fixed interaction order. Both networks can be readily implemented on quantum hardware via midcircuit measurements of auxiliary qubits. If only one auxiliary qubit is measured and reset, the circuit depth scales linearly with imaginary time and number of interactions, while the width is constant. Alternatively, one can employ a number of auxiliary qubits linearly proportional to the number of interactions, and circuit depth grows linearly with imaginary time only. Every midcircuit measurement has a postselection success probability, and the overall success probability is equal to the product of the probabilities of the midcircuit measurements.

97 MATHEMATICS AND COMPUTING

Model-Based Detection of Coordinated Attacks (DCA) in Distribution Systems

The fast-paced growth in digitization of smart grid components enhances system observability and remote-control capabilities through efficient communication. However, enhanced connectivity results in heightened system vulnerability towards cybersecurity risks in the cyber-physical power system. Coordinated cyber-attacks (CCA), when undetected, lead to system-wide impact in terms of large disturbances or widespread outages. Detecting CCA in the cyber layer is critical to thwart cyber-attacks in real-time before the attack impacts the physical system. The challenge of locating CCA stems from the complex grid dynamics, making it difficult to distinguish between normal operational variations and cyber-attack impact. CCA often employs multiple attack vectors targeting geographically distributed components, further complicating CCA identification. Existing research in intrusion detection is primarily focused on the transmission network and limited to detecting individual attacks. In this paper, a novel proactive DCA strategy is proposed for early detection of CCA by establishing correlations among distinct attack events through model-based reinforcement learning that utilizes abductive reasoning to conclude the attacker goal. The solution includes understanding the system model, learning the system dynamics, and correlating individual cyber-attacks to extract the attacker’s objective. The developed learning algorithm identifies the most probable attack path to reach the attacker’s objective by predicting the next attack steps. A DNP3-based cyber-physical co-simulation testbed is developed to test the proposed algorithm using the IEEE 13-node test feeder.

24 POWER TRANSMISSION AND DISTRIBUTION

Uncertainty Quantification Enabled by Automatic Differentiation for Hydrodynamic Simulation of Shock‐to‐Detonation Transition in High Explosives

Quantifying the effects of uncertainty in a reactive burn model on the run-to-detonation time in high explosives (HEs) provides a robust methodology for assessing the probability of an HE failing the IHE qualification standard. Moreover, uncertainty quantification helps evaluate whether the model calibration accurately represents data outside the calibration set. This study uses a specialized hydrodynamic simulation code for modeling detonation to determine the run-to-detonation time of the HE PBX 9502 for various impact velocities. To quickly approximate uncertainties in the model, a surrogate was constructed using a Taylor series expansion centered at the mean of the input parameters. To obtain the sensitivities required for constructing the Taylor series, HYP-percomplex Automatic Differentiation (HYPAD) was implemented. HYPAD is a methodology for infusing existing codes with automatic differentiation capabilities by augmenting variables with one or more imaginary units to compute step-size independent partial derivatives. These derivatives are accurate to machine precision with respect to the implemented numerical algorithm, meaning their accuracy reflects that of the underlying method (e.g., integration or discretization schemes). Using reduced order modeling techniques, the mean and standard deviation of the run-to-detonation time of a shock within PBX 9502 were computed for a number of initial impact velocities. A weighted least squares regression was then performed to obtain a best fit curve and prediction interval for the computed statistics. Historical data points from explosively driven wedge tests were utilized to validate the prediction interval, ensuring its reliability in predicting future outcomes. With this prediction interval and a known safety constraint curve, the most probable point of failure and the probability of failure for the HE PBX 9502 were determined.

97 MATHEMATICS AND COMPUTING

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat

PyFaults: A Python Toolkit for Stacking Fault Screening

PyFaults is an open-source Python library designed to model stacking fault disorder in crystalline materials and qualitatively assess the characteristic selective broadening effects in powder X-ray diffraction (PXRD). Here, the main capabilities of PyFaults are presented, including unit cell and supercell model construction, PXRD pattern calculation, assessment against experimental PXRD, and methods for rapid screening of candidate models within a set of possible stacking vectors and fault occurrence probabilities. This program aims to serve as a computationally inexpensive tool for identifying and screening potential stacking fault models in materials with planar disorder. Three diverse case studies, involving GaN, Li2MnO3 and Li3YCl6, are presented to illustrate the program functionality across a range of structure types and stacking fault modalities.

MATHEMATICS AND COMPUTING

Understanding Pore Filling Processes and Adsorption/Desorption Hysteresis in Nanoporous Metal–Organic Frameworks: Insights from Grand Canonical Monte Carlo Simulations and Free Energy Calculations

Grand canonical Monte Carlo (GCMC) simulations were used to investigate pore filling and hysteresis in nanoporous metal-organic frameworks (MOFs). Adsorption and desorption isotherms were calculated for argon at 87 K in 1866 MOFs from the CoRE MOF database and for short n-alkanes in selected MOFs, keeping the adsorbent structure rigid. Analysis of the molecular configurations showed two different mechanisms and origins of hysteresis: one involving a transition of the adsorbate arrangement in the pores similar to a gas-to-liquid transition associated with a large change in the loading and one more similar to a liquid-to-solid transition associated with a relatively small change in the loading. Our GCMC simulations in MOFs with diverse pore topologies indicate exceptions to an empirical relationship for the minimum diameter of a cylindical pore required for hysteresis as a function of the adsorbate diameter and reduced temperature. The simulations reveal some structures where isotherms exhibit two steps in the adsorption branch and only one step in the desorption branch. Hysteresis loops with a different number of adsorption and desorption steps are not common. Here, to better understand why hysteresis is observed in the GCMC simulations, the concept of the transition probability for observing a step in the adsorption isotherm at a given pressure in a GCMC simulation is introduced. We used two different methods to calculate the transition probabilities and find that these yield comparable results. Furthermore, the transition probability provides a measure for the length of GCMC simulations to yield reliable results.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Accelerating multilevel Markov Chain Monte Carlo using machine learning models

Here, this work presents an efficient approach for accelerating multilevel Markov Chain Monte Carlo (MCMC) sampling for large-scale problems using low-fidelity machine learning models. While conventional techniques for large-scale Bayesian inference often substitute computationally expensive high-fidelity models with machine learning models, thereby introducing approximation errors, our approach offers a computationally efficient alternative by augmenting high-fidelity models with low-fidelity ones within a hierarchical framework. The multilevel approach utilizes the low-fidelity machine learning model (MLM) for inexpensive evaluation of proposed samples thereby improving the acceptance of samples by the high-fidelity model. The hierarchy in our multilevel algorithm is derived from geometric multigrid hierarchy. We utilize an MLM to accelerate the coarse level sampling. Training machine learning model for the coarsest level significantly reduces the computational cost associated with generating training data and training the model. We present an MCMC algorithm to accelerate the coarsest level sampling using MLM and account for the approximation error introduced. We provide theoretical proofs of detailed balance and demonstrate that our multilevel approach constitutes a consistent MCMC algorithm. Additionally, we derive the expression for cost reduction due to machine learning model to facilitate cost analysis of the hierarchical sampling algorithm. Our technique is demonstrated on a standard benchmark inference problem in groundwater flow, where we estimate the probability density of a quantity of interest using a four-level MCMC algorithm. Our proposed algorithm accelerates multilevel sampling by a factor of two while achieving similar accuracy compared to sampling using the standard multilevel algorithm.

97 MATHEMATICS AND COMPUTING