Random Graphs and Probability Models
Preliminary results concerning random graphs, digraphs, time dependent processes, and probability models
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Preliminary results concerning random graphs, digraphs, time dependent processes, and probability models
In this paper, we investigate the global clustering coefficient (a.k.a transitivity) and clique number of graphs generated by a preferential attachment random graph model with an additional feature of allowing edge connections between existing vertices. Specifically, at each time step t, either a new vertex is added with probability f(t), or an edge is added between two existing vertices with probability 1 – f(t). We establish concentration inequalities for the global clustering and clique number of the resulting graphs under the assumption that f(t) is a regularly varying function at infinity with index of regular variation –$\gamma$, where $\gamma$ $\in$ [0, 1). Finally, we also demonstrate an inverse relation between these two statistics: the clique number is essentially the reciprocal of the global clustering coefficient.
Many social networks contain sensitive relational information. One approach to protect the sensitive relational information while offering flexibility for social network research and analysis is to release synthetic social networks at a pre-specified privacy risk level, given the original observed network. In this work, we propose the DP-ERGM procedure that synthesizes networks that satisfy the differential privacy (DP) via the exponential random graph model (EGRM). We apply DP-ERGM to a college student friendship network and compare its original network information preservation in the generated private networks with two other approaches: differentially private DyadWise Randomized Response (DWRR) and Sanitization of the Conditional probability of Edge given Attribute classes (SCEA). The results suggest that DP-EGRM preserves the original information significantly better than DWRR and SCEA in both network statistics and inferences from ERGMs and latent space models. In addition, DP-ERGM satisfies the node DP, a stronger notion of privacy than the edge DP that DWRR and SCEA satisfy.
Abstract Network data often contain sensitive relational information. One approach to protecting sensitive information while offering flexibility for network analysis is to share synthesized networks based on the information in originally observed networks. We employ differential privacy (DP) and exponential random graph models (ERGMs) and propose the DP-ERGM method to synthesize network data. We apply DP-ERGM to two real-world networks. We then compare the utility of synthesized networks generated by DP-ERGM, the DyadWise Randomized Response (DWRR) approach, and the Synthesis through Conditional distribution of Edge given nodal Attribute (SCEA) approach. In general, the results suggest that DP-ERGM preserves the original information significantly better than two other approaches in network structural statistics and inference for ERGMs and latent space models. Furthermore, DP-ERGM satisfies node DP through modeling the global network structure with ERGM, a stronger notion of privacy than the edge DP under which DWRR and SCEA operate.
Abstract not provided.
Abstract not provided.
SAND2025-04671O Feature Pathway Graphs using Random Forest Regressors is a software tool that uses machine learning to determine pathways of influence between features in data sets. It can be used as a surrogate method for casual discovery. The output creates pathway graphs between features of interest. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
Highlights: • Anderson transition from ergodicity to localization on random regular graphs (RRG). • Analytical, pool method, and exact-diagonalization study of correlations on RRG. • Many-body localization (MBL) to ergodicity transition: quantum dots and spin chains. • Dynamical eigenstate correlation functions in RRG and MBL problems. • Anderson localization on RRG as a toy model for MBL. The article reviews the physics of Anderson localization on random regular graphs (RRG) and its connections to many-body localization (MBL) in disordered interacting systems. Properties of eigenstate and energy level correlations in delocalized and localized phases, as well at criticality, are discussed. In the many-body part, models with short-range and power-law interactions are considered, as well as the quantum-dot model representing the limit of the “most long-range” interaction. Central themes – which are common to the RRG and MBL problems – include ergodicity of the delocalized phase, localized character of the critical point, strong finite-size effects, and fractal scaling of eigenstate correlations in the localized phase.
Local structure graph models (LSGMs) describe random graphs and networks as a Markov random field (MRF)—each graph edge has a specified conditional distribution dependent on explicit neighbourhoods of other graph edges. Centered parameterizations of LSGMs allow for direct control and interpretation of parameters for large- and small-scale structures (e.g., marginal means vs. dependence). Here, we extend this parameterization to account for triples of dependent edges and illustrate the importance of centered parameterizations for incorporating covariates and interpreting parameters. Using a MRF framework, common exponential random graph models are also shown to induce conditional distributions without centered parameterizations and thereby have undesirable features. This work attempts to advance graph models through conditional model specifications with modern parameterizations, covariates and higher-order dependencies.
The Katz centrality of a node in a complex network is a measure of the node’s importance as far as the flow of information across the network is concerned. For ensembles of locally tree-like undirected random graphs, this observable is a random variable. Its full probability distribution is of interest but difficult to handle analytically because of its “global” character and its definition in terms of a matrix inverse. Leveraging a fast Gaussian Belief Propagation-Cavity algorithm to solve linear systems on tree-like structures, we show that i) the Katz centrality of a single instance can be computed recursively in a very fast way, and ii) the probability P ( K ) that a random node in the ensemble of undirected random graphs has centrality K satisfies a set of recursive distributional equations, which can be analytically characterized and efficiently solved using a population dynamics algorithm. We test our solution on ensembles of Erdős-Rényi and Scale Free networks in the locally tree-like regime, with excellent agreement. The analytical distribution of centrality for the configuration model conditioned on the degree of each node can be employed as a benchmark to identify nodes of empirical networks with over- and underexpressed centrality relative to a null baseline. We also provide an approximate formula based on a rank- 1 projection that works well if the network is not too sparse, and we argue that an extension of our method could be efficiently extended to tackle analytical distributions of other centrality measures such as PageRank for directed networks in a transparent and user-friendly way.
While classical spin systems in random networks have been intensively studied, much less is known about quantum magnets in random graphs. In this work, we investigate interacting quantum spins on small-world networks, building on mean-field theory and extensive quantum Monte Carlo simulations. Starting from one-dimensional (1D) rings, we consider two situations: All-to-all interacting and long-range interactions randomly added. The effective infinite dimension of the lattice leads to a magnetic ordering at finite temperature Tc with mean-field criticality. Nevertheless, in contrast to the classical case, we find two distinct power-law behaviors for Tc versus the average strength of the extra couplings. This is controlled by a competition between a characteristic length scale of the random graph and the thermal correlation length of the underlying 1D system, thus challenging mean-field theories. We also investigate the fate of a gapped 1D spin chain against the small-world effect.
Female reproductive disorders (FRDs) are common health conditions that may present with significant symptoms. Diet and environment are potential areas for FRD interventions. We utilized a knowledge graph (KG) method to predict factors associated with common FRDs (for example, endometriosis, ovarian cyst, and uterine fibroids). We harmonized survey data from the Personalized Environment and Genes Study (PEGS) on internal and external environmental exposures and health conditions with biomedical ontology content. We merged the harmonized data and ontologies with supplemental nutrient and agricultural chemical data to create a KG. We analyzed the KG by embedding edges and applying a random forest for edge prediction to identify variables potentially associated with FRDs. We also conducted logistic regression analysis for comparison. Across 9765 PEGS respondents, the KG analysis resulted in 8535 significant or suggestive predicted links between FRDs and chemicals, phenotypes, and diseases. Amongst these links, 32 were exact matches when compared with the logistic regression results, including comorbidities, medications, foods, and occupational exposures. Mechanistic underpinnings of predicted links documented in the literature may support some of our findings. Our KG methods are useful for predicting possible associations in large, survey-based datasets with added information on directionality and magnitude of effect from logistic regression. These results should not be construed as causal but can support hypothesis generation. This investigation enabled the generation of hypotheses on a variety of potential links between FRDs and exposures. Future investigations should prospectively evaluate the variables hypothesized to impact FRDs.
Information-theoretic concepts in theory of random graphs - entropy functions for probability distributions and Markov chains
Motivation: Graph representation learning is a family of related approaches that learn low-dimensional vector representations of nodes and other graph elements called embeddings. Embeddings approximate characteristics of the graph and can be used for a variety of machine-learning tasks such as novel edge prediction. For many biomedical applications, partial knowledge exists about positive edges that represent relationships between pairs of entities, but little to no knowledge is available about negative edges that represent the explicit lack of a relationship between two nodes. For this reason, classification procedures are forced to assume that the vast majority of unlabeled edges are negative. Existing approaches to sampling negative edges for training and evaluating classifiers do so by uniformly sampling pairs of nodes. Results: We show here that this sampling strategy typically leads to sets of positive and negative examples with imbalanced node degree distributions. Using representative heterogeneous biomedical knowledge graph and random walk-based graph machine learning, we show that this strategy substantially impacts classification performance. If users of graph machine-learning models apply the models to prioritize examples that are drawn from approximately the same distribution as the positive examples are, then performance of models as estimated in the validation phase may be artificially inflated. We present a degree-aware node sampling approach that mitigates this effect and is simple to implement. Availability and implementation: Our code and data are publicly available at https://github.com/monarch-initiative/negativeExampleSelection.
Synthetically generated, large graph networks serve as useful proxies to real-world networks for many graph-based applications. The ability to generate such networks helps overcome several limitations of real-world networks regarding their number, availability, and access. Here, we present the design, implementation, and performance study of a novel network generator that can produce very large graph networks conforming to any desired degree distribution. The generator is designed and implemented for efficient execution on modern graphics processing units (GPUs). Given an array of desired vertex degrees and number of vertices for each desired degree, our algorithm generates the edges of a random graph that satisfies the input degree distribution. Multiple runtime variants are implemented and tested: 1) a uniform static work assignment using a fixed thread launch scheme, 2) a load-balanced static work assignment also with fixed thread launch but with cost-aware task-to-thread mapping, and 3) a dynamic scheme with multiple GPU kernels asynchronously launched from the CPU. The generation is tested on a range of popular networks such as Twitter and Facebook, representing different scales and skews in degree distributions. Results show that, using our algorithm on a single modern GPU (NVIDIA Volta V100), it is possible to generate large-scale graph networks at rates exceeding 50 billion edges per second for a 69 billion-edge network. GPU profiling confirms high utilization and low branching divergence of our implementation from small to large network sizes. For networks with scattered distributions, we provide a coarsening method that further increases the GPU-based generation speed by up to a factor of 4 on tested input networks with over 45 billion edges.
Graph and network theory play a fundamental role in quantum computer sciences, including quantum information and computation. Random graphs and complex network theory are pivotal in predicting novel quantum phenomena, where entangled links are represented by edges. Quantum algorithms have been developed to enhance solutions for various network problems, giving rise to quantum graph computing and quantum graph learning (QGL). Here, in this review, we explore graph theory and graph learning methods as powerful tools for quantum computers to generate efficient solutions to problems beyond the reach of classical systems. We delve into the development of quantum complex network theory and its applications in quantum computation, materials discovery, and research. We also discuss quantum machine learning (QML) methodologies for effective image classification using qubits, quantum gates, and quantum circuits. Additionally, the paper addresses the challenges of QGL and algorithms, emphasizing the steps needed to develop flexible QGL solvers. This review presents a comprehensive overview of the fields of QGL and QML, highlights recent advancements, and identifies opportunities for future research.
Many applications require to learn, mine, analyze and visualize large-scale graphs. These graphs are often too large to be addressed efficiently using conventional graph processing technologies. Fortunately, recent research efforts find out graph sampling and random walk, which significantly reduce the size of original graphs, can benefit the tasks of learning, mining, analyzing and visualizing large graphs by capturing the desirable graph properties. This paper introduces C-SAW, the first framework that accelerates Sampling and Random Walk framework on GPUs. Particularly, C-SAW makes three contributions: First, our framework provides a generic API which allows users to implement a wide range of sampling and random walk algorithms with ease. Second, offloading this framework on GPU, we introduce warp-centric parallel selection, and two novel optimizations for collision migration. Third, towards supporting graphs that exceed the GPU memory capacity, we introduce efficient data transfer optimizations for out-of-memory and multi-GPU sampling, such as workload-aware scheduling and batched multi-instance sampling. Taken together, our framework constantly outperforms the state of the art projects in addition to the capability of supporting a wide range of sampling and random walk algorithms.