Search NASA⌕ Search

SEARCH · Search NASA

Results for “Probabilistic Graphical Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Software for Data Analysis with Graphical Models

Probabilistic graphical models are being used widely in artificial intelligence and statistics, for instance, in diagnosis and expert systems, as a framework for representing and reasoning with probabilities and independencies. They come with corresponding algorithms for performing statistical inference. This offers a unifying framework for prototyping and/or generating data analysis algorithms from graphical specifications. This paper illustrates the framework with an example and then presents some basic techniques for the task: problem decomposition and the calculation of exact Bayes factors. Other tools already developed, such as automatic differentiation, Gibbs sampling, and use of the EM algorithm, make this a broad basis for the generation of data analysis software.

Buntine, Wray L.↗

Representing Learning With Graphical Models

Probabilistic graphical models are being used widely in artificial intelligence, for instance, in diagnosis and expert systems, as a unified qualitative and quantitative framework for representing and reasoning with probabilities and independencies. Their development and use spans several fields including artificial intelligence, decision theory and statistics, and provides an important bridge between these communities. This paper shows by way of example that these models can be extended to machine learning, neural networks and knowledge discovery by representing the notion of a sample on the graphical model. Not only does this allow a flexible variety of learning problems to be represented, it also provides the means for representing the goal of learning and opens the way for the automatic development of learning algorithms from specifications.

Buntine, Wray L.↗

A probabilistic graphical model foundation for enabling predictive digital twins at scale

A unifying mathematical formulation is needed to move from one-off digital twins built through custom implementations to robust digital twin implementations at scale. This work proposes a probabilistic graphical model as a formal mathematical representation of a digital twin and its associated physical asset. We create an abstraction of the asset–twin system as a set of coupled dynamical systems, evolving over time through their respective state spaces and interacting via observed data and control inputs. The formal definition of this coupled system as a probabilistic graphical model enables us to draw upon well-established theory and methods from Bayesian statistics, dynamical systems and control theory. The declarative and general nature of the proposed digital twin model make it rigorous yet flexible, enabling its application at scale in a diverse range of application areas. Here, we demonstrate how the model is instantiated to enable a structural digital twin of an unmanned aerial vehicle (UAV). The digital twin is calibrated using experimental data from a physical UAV asset. Its use in dynamic decision-making is then illustrated in a synthetic example where the UAV undergoes an in-flight damage event and the digital twin is dynamically updated using sensor data. The graphical model foundation ensures that the digital twin calibration and updating process is principled, unified and able to scale to an entire fleet of digital twins.

42 ENGINEERING↗

Active learning of chemical reaction networks via probabilistic graphical models and Boolean reaction circuits

Discerning networks of many reactions among multiple interconverting species is challenging. Here, we present a reaction network identification methodology. Our methodology enumerates all stoichiometrically and chemically feasible reactions and requires statistical evidence from effluent concentrations for the inclusion or exclusion of each from the reaction network, contrasting with the commonly seen incremental approach and other work of relying heavily upon chemical intuition and assuming the reactions occurring. Using graph theory alongside an active learning design of experiments that propose maximally informative feeds, we identify the underlying reaction network with minimal laboratory runs. Here, we introduce chemistry-probabilistic graphical modeling and Boolean reaction circuits to statistically quantify which reactions occur from effluent concentrations. Our methodology accurately discerns active reactions, as showcased upon a laboratory network of cross-ketonization of furoic and lauric acid and validated upon simulated networks of thermal and CO 2 -assisted ethane dehydrogenation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multisource Data Fusion Outage Location in Distribution Systems via Probabilistic Graphical Models

Efficient outage location is critical to enhancing the resilience of power distribution systems. However, accurate outage location requires combining massive evidence received from diverse data sources, including smart meter (SM) last gasp signals, customer trouble calls, social media messages, weather data, vegetation information, and physical parameters of the network. This is a computationally complex task due to the high dimensionality of data in distribution grids. In this paper, we propose a multi-source data fusion approach to locate outage events in partially observable distribution systems using Bayesian networks (BNs). A novel aspect of the proposed approach is that it takes multi-source evidence and the complex structure of distribution systems into account using a probabilistic graphical method. Our method can radically reduce the computational complexity of outage location inference in high-dimensional spaces. The graphical structure of the proposed BN is established based on the network’s topology and the causal relationship between random variables, such as the states of branches/customers and evidence. Utilizing this graphical model, accurate outage locations are obtained by leveraging a Gibbs sampling (GS) method, to infer the probabilities of de-energization for all branches. Compared with commonly-used exact inference methods that have exponential complexity in the size of the BN, GS quantifies the target conditional probability distributions in a timely manner. As a result, a case study of several real-world distribution systems is presented to validate the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Analysis with Graphical Models: Software Tools

Probabilistic graphical models (directed and undirected Markov fields, and combined in chain graphs) are used widely in expert systems, image processing and other areas as a framework for representing and reasoning with probabilities. They come with corresponding algorithms for performing probabilistic inference. This paper discusses an extension to these models by Spiegelhalter and Gilks, plates, used to graphically model the notion of a sample. This offers a graphical specification language for representing data analysis problems. When combined with general methods for statistical inference, this also offers a unifying framework for prototyping and/or generating data analysis algorithms from graphical specifications. This paper outlines the framework and then presents some basic tools for the task: a graphical version of the Pitman-Koopman Theorem for the exponential family, problem decomposition, and the calculation of exact Bayes factors. Other tools already developed, such as automatic differentiation, Gibbs sampling, and use of the EM algorithm, make this a broad basis for the generation of data analysis software.

Buntine, Wray L.↗

ML for microbiomes

The software provides machine learning analysis and visualization to detect patterns in microbiome data, including topic modeling, probabilistic graphical modeling, conventional machine learning methods, and deep learning. The software is written in python and R, it uses some python and R libraries as well as big open-source libraries like sklearn, networkX, pytorch (python), pgmpy (python), and bnlearn (R). It also has a script to use for MALLET and DTM (open-source packages for topic modeling, written in Java).

Kim, Anastasiia↗

A Guide to the Literature on Learning Graphical Models

This literature review discusses different methods under the general rubric of learning Bayesian networks from data, and more generally, learning probabilistic graphical models. Because many problems in artificial intelligence, statistics and neural networks can be represented as a probabilistic graphical model, this area provides a unifying perspective on learning. This paper organizes the research in this area along methodological lines of increasing complexity.

Buntine, Wray L.↗

Conin

SAND2025-07645O Conin is a Python library that supports constrained analysis of probabilistic graphical models (PGMs). It enables constrained inference and learning for hidden Markov models, Bayesian networks, dynamic Bayesian networks, and Markov networks. Conin interfaces with the pgmpy library to specify general probabilistic graphical models with a variety of optimization solvers to support learning and inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, William [Sandia National Lab. (SNL-CA), Live↗

Graph-learning approach to combine multiresolution seismic velocity models

SUMMARY The resolution of velocity models obtained by tomography varies due to multiple factors and variables, such as the inversion approach, ray coverage, data quality, etc. Combining velocity models with different resolutions can enable more accurate ground motion simulations. Toward this goal, we present a novel methodology to fuse multiresolution seismic velocity maps with probabilistic graphical models (PGMs). The PGMs provide segmentation results, corresponding to various velocity intervals, in seismic velocity models with different resolutions. Further, by considering physical information (such as ray path density), we introduce physics-informed probabilistic graphical models (PIPGMs). These models provide data-driven relations between subdomains with low (LR) and high (HR) resolutions. Transferring (segmented) distribution information from the HR regions enhances the details in the LR regions by solving a maximum likelihood problem with prior knowledge from HR models. When updating areas bordering HR and LR regions, a patch-scanning policy is adopted to consider local patterns and avoid sharp boundaries. To evaluate the efficacy of the proposed PGM fusion method, we tested the fusion approach on both a synthetic checkerboard model and a fault zone structure imaged from the 2019 Ridgecrest, CA, earthquake sequence. The Ridgecrest fault zone image consists of a shallow (top 1 km) high-resolution shear-wave velocity model obtained from ambient noise tomography, which is embedded into the coarser Statewide California Earthquake Center Community Velocity Model version S4.26-M01. The model efficacy is underscored by the deviation between observed and calculated traveltimes along the boundaries between HR and LR regions, 38 per cent less than obtained by conventional Gaussian interpolation. The proposed PGM fusion method can merge any gridded multiresolution velocity model, a valuable tool for computational seismology and ground motion estimation.

Geochemistry & Geophysics↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Risk-Informed Condition Evaluation of Solar-centered Energy Generation and Distribution Networks through Bayesian Learning and Inference

We develop a methodology based on Bayesian inference over Probabilistic Graphical Models (PGMs) to understand and quantify risk in solar-centered grids using targeted measurements and learned system behavior. Being non-prescriptive but, rather, able to infer system behavior and, ultimately, address risk queries from data, our machine learning-type paradigm is tailored for diverse topologies and threat scenarios often associated with distributed energy generation and photovoltaic distributed energy resources (PV-DERs) in particular. We describe algorithmic processes for: (i) learning the structure of PGMs that result from attack-prone PV-DER-proliferated distribution systems, (ii) quantifying cause-effect relationships, and (iii) evaluating risk queries based on diverse evidence. The contributions are illustrated on a residential grid subject to output impairment attacks on its PV-DER infrastructure.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Markov chain Monte Carlo with digital dissipative dynamics on quantum computers

Modeling the dynamics of a quantum system connected to the environment is critical for advancing our understanding of complex quantum processes, as most quantum processes in nature are affected by an environment. Modeling a macroscopic environment on a quantum simulator may be achieved by coupling independent ancilla qubits that facilitate energy exchange in an appropriate manner with the system and mimic an environment. This approach requires a large, and possibly exponential number of ancillary degrees of freedom which is impractical. In contrast, we develop a digital quantum algorithm that simulates interaction with an environment using a small number of ancilla qubits. By combining periodic modulation of the ancilla energies, or spectral combing, with periodic reset operations, we are able to mimic interaction with a large environment and generate thermal states of interacting many-body systems. We evaluate the algorithm by simulating preparation of thermal states of the transverse Ising model. Our algorithm can also be viewed as a quantum Markov chain Monte Carlo process that allows sampling of the Gibbs distribution of a multivariate model. To illustrate this we evaluate the accuracy of sampling Gibbs distributions of simple probabilistic graphical models using the algorithm.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Methods for Probabilistic Fault Diagnosis: An Electrical Power System Case Study

Health management systems that more accurately and quickly diagnose faults that may occur in different technical systems on-board a vehicle will play a key role in the success of future NASA missions. We discuss in this paper the diagnosis of abrupt continuous (or parametric) faults within the context of probabilistic graphical models, more specifically Bayesian networks that are compiled to arithmetic circuits. This paper extends our previous research, within the same probabilistic setting, on diagnosis of abrupt discrete faults. Our approach and diagnostic algorithm ProDiagnose are domain-independent; however we use an electrical power system testbed called ADAPT as a case study. In one set of ADAPT experiments, performed as part of the 2009 Diagnostic Challenge, our system turned out to have the best performance among all competitors. In a second set of experiments, we show how we have recently further significantly improved the performance of the probabilistic model of ADAPT. While these experiments are obtained for an electrical power system testbed, we believe they can easily be transitioned to real-world systems, thus promising to increase the success of future NASA missions.

Ricks, Brian W.↗

Use of Machine Learning Techniques for Iidentification of Robust Teleconnections to East African Rainfall Variability in Observations and Models

Providing advance warning of East African rainfall variations is a particular focus of several groups including those participating in the Famine Early Warming Systems Network. Both seasonal and long-term model projections of climate variability are being used to examine the societal impacts of hydrometeorological variability on seasonal to interannual and longer time scales. The NASA / USAID SERVIR project, which leverages satellite and modeling-based resources for environmental decision making in developing nations, is focusing on the evaluation of both seasonal and climate model projections to develop downscaled scenarios for using in impact modeling. The utility of these projections is reliant on the ability of current models to capture the embedded relationships between East African rainfall and evolving forcing within the coupled ocean-atmosphere-land climate system. Previous studies have posited relationships between variations in El Niño, the Walker circulation, Pacific decadal variability (PDV), and anthropogenic forcing. This study applies machine learning methods (e.g. clustering, probabilistic graphical model, nonlinear PCA) to observational datasets in an attempt to expose the importance of local and remote forcing mechanisms of East African rainfall variability. The ability of the NASA Goddard Earth Observing System (GEOS5) coupled model to capture the associated relationships will be evaluated using Coupled Model Intercomparison Project Phase 5 (CMIP5) simulations.

Roberts, J. Brent↗

Artificial Reasoning System for Symptom-Based Conditional Failure Probability Estimation Using Bayesian Network

Advances in nuclear power technologies require enhanced capabilities for operator advice and autonomous control. One of the first tasks in the development of such capabilities is the formulation of symptom-based conditional failure probabilities for structures, systems, and components (SSCs) of interest, for which the primary goal is to aid plant personnel in deducing the probabilistic performance status of the monitored SSCs and in detecting impending faults/failure. The task of conditional failure probability estimation is a bidirectional inference problem and shall be logically tackled by the Bayesian network (BN) approach. As a knowledge-based artificial intelligence tool and a probabilistic graphical model, BN offers the capability of reasoning under uncertainty and graphical representation emulating the physical behavior of the target SSC. This paper provides a systematic overview of the BN technique and the software tools for handling implementation of BN models, along with the associated knowledge representation and reasoning paradigm. Both operational data and expert judgement can be readily incorporated into the knowledge base of a BN model. The challenges with data availability are highlighted, and the general approach to target SSC identification is presented. Our focus is upon failure-prone and risk-important balance of plant assets, especially cases having strong operator involvement. An exemplary case study on the failure of a motor-driven centrifugal pump is also conducted to demonstrate the usefulness and technical feasibility of the proposed artificial reasoning system using an expert system shell.

Zhao, Xingang↗

How to Obtain the Redshift Distribution from Probabilistic Redshift Estimates

Abstract A reliable estimate of the redshift distribution n ( z ) is crucial for using weak gravitational lensing and large-scale structures of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo- z ) probability density functions (PDFs) the next best alternative for comprehensively encapsulating the nontrivial systematics affecting photo- z point estimation. The established stacked estimator of n ( z ) avoids reducing photo- z PDFs to point estimates but yields a systematically biased estimate of n ( z ) that worsens with a decreasing signal-to-noise ratio, the very regime where photo- z PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts ( CHIPPR ), a statistically rigorous probabilistic graphical model of redshift-dependent photometry that correctly propagates the redshift uncertainty information beyond the best-fit estimator of n ( z ) produced by traditional procedures and is provably the only self-consistent way to recover n ( z ) from photo- z PDFs. We present the chippr prototype code, noting that the mathematically justifiable approach incurs computational cost. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo- z techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo- z community focus on developing methodologies that enable the recovery of photo- z likelihoods with support over all redshifts, either directly or via a known prior probability density.

79 ASTRONOMY AND ASTROPHYSICS↗