Search NASA⌕ Search

SEARCH · Search NASA

Results for “common random numbers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Stochastic Quasi-Newton Method in the Absence of Common Random Numbers

We present Q-SASS, a quasi-Newton method for unconstrained stochastic optimization that does not rely on common random numbers. Most existing quasi-Newton approaches leverage common random numbers to construct second-order updates. However, motivated by challenges in variational quantum algorithms—where such coordination is not possible—we consider the setting in which function values and gradients are accessible only through noisy probabilistic zeroth- and first-order oracles, and no common random numbers can be exploited. We derive high-probability tail bounds on the iteration complexity of our algorithm for nonconvex, convex, and strongly convex (more generally, those satisfying the PL condition) objective functions. Finally, we demonstrate the empirical benefits of our quasi-Newton updating scheme on both synthetic and quantum chemistry problems.

Complexity bound↗

Derivative-free stochastic optimization via adaptive sampling strategies

In this paper, we present a novel derivative-free framework for solving unconstrained stochastic optimization problems. Many problems in fields ranging from simulation optimization to reinforcement learning to quantum computing involve settings where only stochastic function values are obtained via a zeroth-order oracle, which has no available gradient information and necessitates the usage of derivative-free optimization methodologies. Our approach includes estimating gradients using stochastic function evaluations and integrating adaptive sampling techniques to control the accuracy in these stochastic approximations. Our framework encapsulates several gradient estimation techniques, including standard finite-difference, Gaussian smoothing, sphere smoothing, randomized coordinate finite-difference, and randomized subspace finite-difference methods. We provide theoretical convergence guarantees for our framework and analyze the worst-case iteration and sample complexities associated with each gradient estimation method. Finally, we demonstrate the empirical performance of the methods on logistic regression and nonlinear least squares problems.

Adaptive sampling↗

A mixture of grass–legume cover crop species may ameliorate water stress in a changing climate

Climate change models predict increasing precipitation variability in the mid-latitude regions of Earth, generating a need to reduce the negative impacts of these changes on crop production. Despite considerable research on how cover crops support agriculture in a changing climate, understanding is limited of how climate change influences the growth of cover crops. We investigated the early development of two common cover crop species—crimson clover (Trifolium incarnatum) and rye (Secale cereale)—and hypothesized that growing them in the mixture would ameliorate stress from drought or waterlogging. This hypothesis was tested in a 25-day greenhouse experiment, where the two factors (species number and water stress) were fully crossed in randomized blocks, and plant responses were quantified through survival, growth rate, biomass production and root morphology. Water stress negatively influenced the early growth of these two species in contrasting ways: crimson clover was susceptible to drought while rye performed poorly under waterlogging. Per-plant biomass in rye was always greater in mixture than in monoculture, while per-plant biomass of crimson clover was greater in mixture under drought. Both species grew longer roots in mixture than in monoculture under drought, and total biomass of mixtures did not differ significantly from the more-productive monoculture (rye) in any water condition. In the face of increasingly variable precipitation, growing crimson clover and rye together has potential to ameliorate water stress, a possibility that should be further investigated in field experiments.

54 ENVIRONMENTAL SCIENCES↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Improved multifidelity Monte Carlo estimators based on normalizing flows and dimensionality reduction techniques

Here, we study the problem of multifidelity uncertainty propagation for computationally expensive models. In particular, we consider the general setting where the high-fidelity and low-fidelity models have a dissimilar parameterization both in terms of number of random inputs and their probability distributions, which can be either known in closed form or provided through samples. We derive novel multifidelity Monte Carlo estimators which rely on a shared subspace between the high-fidelity and low-fidelity models where the parameters follow the same probability distribution, i.e., a standard Gaussian. We build the shared space employing normalizing flows to map different probability distributions into a common one, together with linear and nonlinear dimensionality reduction techniques, active subspaces and autoencoders, respectively, which capture the subspaces where the models vary the most. We then compose the existing low-fidelity model with these transformations and construct modified models with an increased correlation with the high-fidelity model, which therefore yield multifidelity estimators with reduced variance. A series of numerical experiments illustrate the properties and advantages of our approaches.

97 MATHEMATICS AND COMPUTING↗

Sign Problem in Tensor-Network Contraction

We investigate how the computational difficulty of contracting tensor networks depends on the sign structure of the tensor entries. Using results from computational complexity, we observe that the approximate contraction of tensor networks with only positive entries has lower computational complexity as compared to tensor networks with general real or complex entries. This raises the question of how this transition in computational complexity manifests itself in the hardness of different tensor-network-contraction schemes. We pursue this question by studying random tensor networks with varying bias toward positive entries. First, we consider contraction via Monte Carlo sampling and find that the transition from hard to easy occurs when the tensor entries become predominantly positive; this can be understood as a tensor-network manifestation of the well-known negative-sign problem in quantum Monte Carlo. Second, we analyze the commonly used contraction based on boundary tensor networks. The performance of this scheme is governed by the number of correlations in contiguous parts of the tensor network (which by analogy can be thought of as entanglement). Remarkably, we find that the transition from hard to easy—i.e., from a volume-law to a boundary-law scaling of entanglement—already occurs for a slight bias of the tensor entries toward a positive mean, scaling inversely with the bond dimension D , and thus the problem becomes easy the earlier the larger D occurs. This is in contrast both to expectations and to the behavior found in Monte Carlo contraction, where the hardness at fixed bias increases with the bond dimension. To provide insight into this early breakdown of computational hardness and the accompanying entanglement transition, we construct an effective classical statistical-mechanical model that predicts a transition at a bias of the tensor entries of 1 / D , confirming our observations. We conclude by investigating the computational difficulty of computing expectation values of tensor-network wave functions (projected entangled-pair states, PEPSs) and find that in this setting, the complexity of entanglement-based contraction always remains low. We explain this by providing a local transformation that maps PEPS expectation values to a positive-valued tensor network. This not only provides insight into the origin of the observed boundary-law entanglement scaling but also suggests new approaches toward PEPS contraction based on positive decompositions. Published by the American Physical Society 2025

Chen, Jielun (ORCID:0000000178411545)↗

Structured interactions drive abrupt transitions in the spatial organization of microbial communities

Bacteria possess diverse mechanisms to regulate their motility in response to environmental and physiological signals, enabling them to navigate complex habitats and adapt their behavior. Some of these mechanisms are species specific and enable cells to modulate their movement based on the ecological identity of neighboring species. Here, we introduce a model in which bacteria interact via local signals that either enhance or suppress the motility of neighboring cells depending on species type. Through large-scale simulations and a coarse-grained stochastic model, we demonstrate the emergence of a sharp transition driven by nucleation processes: increasing the density of motility-suppressing interactions drives the system from a fully mixed, motile phase to a state characterized by large, stationary bacterial clusters. Remarkably, in systems with a large number of interacting species, this transition can be triggered solely by altering the structure of the motility-regulation interaction matrix while maintaining species and interaction densities constant. In particular, we find that heterogeneous and modular interactions promote the transition more readily than homogeneous random ones. These findings add a dimension to the theory of motility-induced phase separation and contribute to the ongoing effort to understand microbial interactions, suggesting that structured, nonrandom ones may be key to reproducing commonly observed spatial patterns in microbial communities.

bacterial communities↗

Machine learning reveals strong grid-scale dependence in the satellite N d –LWP relationship

The relationship between cloud droplet number concentration ( N d ) and liquid water path (LWP) is highly uncertain yet crucial for determining the impact of aerosol-cloud interactions (ACI) on Earth's radiation budget. The N d -LWP relationship is examined using a machine learning (ML) random forest model applied to five years of satellite data at grid resolutions ranging from 10° to 0.05° in 12 distinct regions. In the subtropics, the shape of the N d -LWP relationship switches from an inverted-V at 1° grid-resolution to an “M” shape at 0.1° resolution with decreased $\frac{\textrm{dln⁡LWP}}{\textrm{dln⁡}N_d}$ sensitivity. Tropical and midlatitude regions generally show a more positive sensitivity. Cloud sampling and filtering also influence this slope, wherein the exclusion of thin clouds, as commonly performed to reduce retrieval uncertainty, leads to strongly negative sensitivity across all regions. Precipitation is primarily responsible for driving the strength of the sensitivity, with strong positive slopes in raining clouds and negative and/or neutral responses found in non-raining clouds. A new method to compute radiative forcing from the ML model shows a robust Twomey radiative forcing across all regions and grid resolutions. However, LWP and cloud fraction adjustments to the radiative forcing, which are ∼50 % or smaller than the Twomey effect, decrease to negligible values with higher spatial resolution data. As Earth system models move toward higher spatial resolutions in the future, evaluating the LWP and CF adjustment contributions to the radiative forcing budget at these finer resolutions will be essential for evaluation and model development.

Aerosol-Cloud Interactions↗

Commuting embeddings for parallel strategies in non-local games

Non-local games provide a versatile framework for probing quantum correlations and for benchmarking the power of entanglement. In finite dimensions, the standard method for playing several games in parallel requires a tensor product of the local Hilbert spaces, which scales additively in the number of qubits. In this work, we show that this additive cost can be reduced by exploiting algebraic embeddings. We introduce two forms of compressions. First, when a referee selects one game from a finite collection of games at random, the game quantum strategy can be implemented using a maximally entangled state of dimension equal to the largest individual game, thereby eliminating the need for repeated state preparations. Second, we establish conditions under which several games can be played simultaneously in parallel on fewer qubits than the tensor product baseline. These conditions are expressed in terms of commuting embeddings of the game algebras. Moreover, we provide a constructive framework for building such embeddings. Using tools from Lie theory, we show that aligning the various game algebras into a common Cartan decomposition enables such a qubit reduction. Beyond the theoretical contribution, our framework casts NLGs as algebraic primitives for distributed and resource-constrained quantum computations and suggested NLGs as a comparable device-independent dimension witness.

Commuting embeddings↗

The Role of Data Filtering in Open Source Software Ranking and Selection

Faced with more than 100M open source projects, a more manageable small subset is needed for most empirical investigations. More than half of the research papers in leading venues investigated filtering projects by some measure of popularity with explicit or implicit arguments that unpopular projects are not of interest, may not even represent "real" software projects, or that less popular projects are not worthy of study. However, such filtering may have enormous effects on the results of the studies if and precisely because the sought-out response or prediction is in any way related to the filtering criteria.This paper exemplifies the impact of this common practice on research outcomes, specifically how filtering of software projects on GitHub based on inherent characteristics affects the assessment of their popularity. Using a dataset of over 100,000 repositories, we used multiple regression to model the number of stars -a commonly used proxy for popularity- based on factors such as the number of commits, the duration of the project, the number of authors and the number of core developers. Our control model included the entire dataset, while a second filtered model considered only projects with ten or more authors. The results indicated that while certain characteristics of the repository consistently predict popularity, the filtering process significantly alters the relationships between these characteristics and the response. We found that the number of commits exhibited a positive correlation with popularity in the control sample but showed a negative correlation in the filtered sample. These findings highlight the potential biases introduced by data filtering and emphasize the need for careful sample selection in empirical research of mining software repositories. We recommend that empirical work should either analyze complete datasets such as World of Code, or employ stratified random sampling from a complete dataset to ensure that filtering is not biasing the results.

Malviya Thakur, Addi↗

Sub-pilot-scale Production of High-Value Products from U.S. Coals

Investigators from the University of Utah, University of Wyoming and Marshall University pursued a program to study the conversion of raw coal to high-value products of carbon fiber and silicon carbide. Team members also developed an initial framework for a data portal that can incorporate laboratory data on coal processing and product quality, and also work with tools for machine learning for data analysis, data visualization and economic assessment. Experimental R&D efforts focused on the conversion of raw coal to coal tar and other byproducts, and the resulting tar intermediates were upgraded to form anisotropic and isotropic pitch materials. These pitch materials were produced from coal using both thermal (pyrolysis) and chemical (mild solvolysis liquefaction) decomposition of raw coal. Four different coals were studied: Utah bituminous coal (Sufco), Wyoming PRB coal (Black Thunder), Illinois bituminous coal (Illinois #6), and West Virginia bituminous coal (Flying Eagle). Both metallurgical-grade coking coals and lower-grade steam coals were investigated, and controlled secondary gas-phase reactions were used during a two-stage pyrolysis process to induce cracking and condensation reactions among the pyrolytic tar species. This approach successfully improved the performance of the lower grade coals for yielding pitch materials, with properties more consistent with a commercial-grade pitch that had previously demonstrated success for quality carbon fiber production. The use of waste plastic materials was also studied, to help improve physical and chemical characteristics of the intermediate tars and final pitch product; in particular, for lowering the pitch softening point to an acceptable level for melt spinning carbon fiber. Mild solvolysis liquefaction was also used as a method for producing pitch for carbon fiber production. As expected, significantly higher pitch yields were obtained using this approach, and waste plastic materials were also successfully used to reduce pitch softening point to an acceptable level. The plastic materials were also utilized to create a solvent for the mild solvolysis process, and this plastic-derived solvent was shown to provide results consistent with more expensive commercial chemical solvents, and could thus avoid the need for costly recovery and recycle of a liquefaction solvent. Additional experimental R&D focused on the production of silicon carbide (β-SiC) from the residual char byproduct from pitch production, and also on the production of carbon fiber from the anisotropic pitch. SiC was successfully synthesized using a mixture of residual char and sandstone at a ratio of 1:1. Reaction temperature and residence time were optimized and yielded a product purity of 81%. For carbon fiber production, the most successful pitch samples were obtained from the mild solvolysis liquefaction approach, combined with the use of a plastic (HDPE)-derived solvent. Fiber properties improved over time as laboratory fiber production methodologies improved, and final yields of carbon fiber were obtained with a diameter of 12.14 ± 1.10 um, Modulus of 173.73 ± 15.25 GPa, and Tensile Strength of 1.04 ± 0.10 GPa. A proof-of-concept Modern Community Research Data Portal (MCRDP) was developed and deployed for coal and coal-derived pitch characterization, with the full support of (i) remote web-based access, (ii) distributed analysis, (iii) interactive visualization and exploration, (iv) shared and long-term data access, (v) advanced query capabilities and (vi) real-time collaboration. The Coal to Products Data Portal “coaltoproducts.org” provides researchers with space to store and share data within a project, tools for analyzing and understanding data for scientific investigation, and the ability to publish data to the broader community for reproducibility. The portal leverages the Material Commons 2.0 (MC) platform developed by the Center for PRedictive Integrated Structural Materials Science (PRISMS) of the University of Michigan, to achieve long-term longevity of data collections and, more importantly, collaborative science. A number of data visualization tools were also assessed and implemented for interrogating the experimental and modeling data. The machine learning portion of this project analyzed datasets from two different coal conversion processes performed on a diverse set of coal samples from both the coal pyrolysis experiments and the solvent liquefaction experiments. The work was initiated by exploring standard regression models on the pyrolysis data, aiming to understand the impact of sample characteristics and processing conditions on key product metrics. Over the course of the project, the focus expanded to include a variety of machine learning tools, delving into both supervised and unsupervised learning methods. Models tested on the pyrolysis data included linear, ridge, lasso, elastic-net, Gaussian process, random forest regression, and AutoSklearn, and the approach was continually refined to enhance predictive accuracy and model interpretability. Similar techniques were applied to the liquefaction data with an additional focus on feature engineering. Along with mesophase content, additional outputs of interest were the pitch yield, softening point, and QI content. Insights derived from these analyses are crucial in determining the factors influencing the quality and yield of coal-derived products. As the work progressed, the research evolved from foundational model comparisons to analyses of random forests, decision paths, and feature importance scores. A thorough market analysis was performed to examine the prospects of coal-based carbon fibers. The best opportunities for coal come from its lower and more stable price relative to petroleum, particularly for subbituminous coals, which is the primary advantage that a coal refinery may have over a petroleum refinery. Before a commercial CTP production facility can be modeled, however, several things need to be understood regarding the nature of the would-be coal refinery. These include the technology to be deployed, the size of facility, the volume(s) of co-product(s), and the waste and emissions profile of the plant. The volume of co-products and waste may be substantial and will require separate market analysis to ensure viability. In the near-term, the importance of coal tar pitch, in the form of carbon pitch, to the aluminum and steel industries is likely to overshadow the alternative use of this material as an input for carbon fiber. The importance of steel and aluminum in building materials, and the need for carbon materials in their manufacturing, will ensure that demand for these products remains for the long run. In addition, carbon fiber may also be the best substitute for steel and aluminum well into the future. While society will eventually be able to shift production of much of its electricity needs to renewables, it will not be able to shift away from fossil fuels for production of high-strength construction and vehicular materials. Demand for carbon fiber is expected to increase quickly, but the volume of carbon fiber and the amount of coal that would be needed to produce even a sizeable share of this market may still be relatively small compared to current coal production. Thus, other coal-based products like graphene, graphite, carbon foams, resins, and carbon-based building products will play important roles in sustaining coal production as coal-fired power generation continues to decline.

01 COAL, LIGNITE, AND PEAT↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

On the connection between least squares, regularization, and classical shadows

Classical shadows (CS) offer a resource-efficient means to estimate quantum observables, circumventing the need for exhaustive state tomography. Here, we clarify and explore the connection between CS techniques and least squares (LS) and regularized least squares (RLS) methods commonly used in machine learning and data analysis. By formal identification of LS and RLS ``shadows'' completely analogous to those in CS---namely, point estimators calculated from the empirical frequencies of single measurements---we show that both RLS and CS can be viewed as regularizers for the underdetermined regime, replacing the pseudoinverse with invertible alternatives. Through numerical simulations, we evaluate RLS and CS from three distinct angles: the tradeoff in bias and variance, mismatch between the expected and actual measurement distributions, and the interplay between the number of measurements and number of shots per measurement. Compared to CS, RLS attains lower variance at the expense of bias, is robust to distribution mismatch, and is more sensitive to the number of shots for a fixed number of state copies---differences that can be understood from the distinct approaches taken to regularization. Conceptually, our integration of LS, RLS, and CS under a unifying ``shadow'' umbrella aids in advancing the overall picture of CS techniques, while practically our results highlight the tradeoffs intrinsic to these measurement approaches, illuminating the circumstances under which either RLS or CS would be preferred, such as unverified randomness for the former or unbiased estimation for the latter.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluation of RANS vs. LES simulation of fluid flow through 3 × 3 rod bundle with a simple spacer grid as a precursor to coupled fluid–structure interaction simulations

The research literature on Computational Fluid Dynamics (CFD) of coolant flow through rod bundles with spacer-grids and mixing vanes is replete, ranging from high fidelity Large Eddy Simulation (LES)/Direct Numerical Simulation (DNS) simulations to Reynolds-Averaged Navier–Stokes (RANS) modeled studies. The mixing of flow between subchannels and the pressure drop through the bundle are fundamental quantities useful for comparing and evaluating CFD methods. Less commonly observed and compared are the forces exerted onto the structure by the fluid. The present study seeks to evaluate the use of RANS simulations for predicting the structural response to fluid flow. Wall resolved RANS simulations are benchmarked against LES simulations of fluid flow at a Reynolds number of 15,000 through a 3 × 3 fuel rod bundle with a simple spacer grid. Velocity line-plots are compared showing good agreement between RANS and LES results, ascertaining that the former is capable of capturing the essential time-averaged velocity profile. Additionally, the distribution of forces on the spacer grid and fuel rods are collected as a function of time and space. The RANS methods are evaluated using the frequency and magnitude of the fluctuating forces on various portions of the structure as compared to LES. In conclusion, the power spectral density evaluation of the models reveal underprediction of force amplitude on the rod walls by RANS and also discrepancy in the prediction of high frequency spectra, especially in the immediate vicinity of spacer-grid structure, which may be attributed to the lack of random turbulence fluctuation or insufficient modeling of small-scale eddies in RANS simulation.

FIV↗

Lineage frequency time series reveal elevated levels of genetic drift in SARS-CoV-2 transmission in England

Genetic drift in infectious disease transmission results from randomness of transmission and host recovery or death. The strength of genetic drift for SARS-CoV-2 transmission is expected to be high due to high levels of superspreading, and this is expected to substantially impact disease epidemiology and evolution. However, we don’t yet have an understanding of how genetic drift changes over time or across locations. Furthermore, noise that results from data collection can potentially confound estimates of genetic drift. To address this challenge, we develop and validate a method to jointly infer genetic drift and measurement noise from time-series lineage frequency data. Our method is highly scalable to increasingly large genomic datasets, which overcomes a limitation in commonly used phylogenetic methods. We apply this method to over 490,000 SARS-CoV-2 genomic sequences from England collected between March 2020 and December 2021 by the COVID-19 Genomics UK (COG-UK) consortium and separately infer the strength of genetic drift for pre-B.1.177, B.1.177, Alpha, and Delta. We find that even after correcting for measurement noise, the strength of genetic drift is consistently, throughout time, higher than that expected from the observed number of COVID-19 positive individuals in England by 1 to 3 orders of magnitude, which cannot be explained by literature values of superspreading. Our estimates of genetic drift suggest low and time-varying establishment probabilities for new mutations, inform the parametrization of SARS-CoV-2 evolutionary models, and motivate future studies of the potential mechanisms for increased stochasticity in this system.

60 APPLIED LIFE SCIENCES↗

General trends of superconducting pairing and magnetic correlations in the Ruddlesden-Popper nickelate 𝑚-layered superconductors La 𝑚+1 ⁢Ni 𝑚 ⁢O 3⁢𝑚+1

Here, we report a comprehensive theoretical analysis of the Ruddlesden-Popper layered nickelates La 𝑚+1 ⁢Ni 𝑚 ⁢O 3⁢𝑚+1 (𝑚 = 1 to 6) under pressure. These materials have recently received significant attention due to the discovery of superconductivity in some nickelates under pressure. Our results suggest that, while these Ruddlesden-Popper layered nickelates display many similarities, they also show noticeable differences. One of the common features of La 𝑚+1 ⁢Ni 𝑚 ⁢O 3⁢𝑚+1 is that the electronic states near the Fermi level are mainly contributed by Ni 3⁢𝑑 orbitals, slightly hybridized with O 2⁢𝑝 orbitals. The Ni 𝑑 3⁢𝑧 2 −𝑟 2 orbitals display bonding-antibonding, or bonding-antibonding-nonbonding, characteristic splittings, depending on the even or odd number of stacking layers 𝑚. In addition, the ratio of the in-plane interorbital hopping between 𝑑 3⁢𝑧 2 −𝑟 2 and 𝑑 𝑥 2 −𝑦 2 orbitals and in-plane intraorbital hopping between 𝑑 𝑥 2 −𝑦 2 orbitals was found to be large in La 𝑚+1 ⁢Ni 𝑚 ⁢O 3⁢𝑚+1 (𝑚 = 1 to 6), and this ratio increases from 𝑚 = 1 to 𝑚 = 6, suggesting that the in-plane hybridization will increase as the layer number 𝑚 increases. In contrast to the dominant 𝑠 ± -wave state driven by spin fluctuations in the bilayer La 3 ⁢Ni 2⁢ O 7 and trilayer La 4 ⁢Ni 3 ⁢O 10 , two nearly degenerate 𝑑 𝑥 2 −𝑦 2 -wave and 𝑠 ± -wave leading states were obtained in the four-layer stacking La 5⁢ Ni 4 ⁢O 13 and five-layer stacking La 6 ⁢Ni 5 ⁢O 16 . The leading 𝑠 ± -wave state was recovered in the six-layer material La 7 ⁢Ni 6 ⁢O 19 with slightly higher calculated pairing strength 𝜆 than that of the 𝑑 𝑥 2 −𝑦 2 -wave state. All this evidence suggests that both 𝑠 ± -wave and 𝑑 𝑥 2 −𝑦 2 -wave channels are strongly competing in the high-order niceklates based on our random-phase approximation calculations. In general, at the level of the random-phase approximation treatment, the superconducting transition temperature 𝑇 𝑐 decreases in stoichiometric bulk systems from the bilayer La 3 ⁢Ni 2 ⁢O 7 to the six-layer La 7 ⁢Ni 6 ⁢O 19 , despite the 𝑚-dependent dominant pairing. Both in-plane and out-of-plane magnetic correlations are found to be quite complex. Within the in-plane direction, we obtained the peak of the magnetic susceptibility at 𝐪 = (0.6⁢𝜋, 0.6⁢𝜋) for La 5 ⁢Ni 4 ⁢O 13 (𝑚 = 4) and La 7 ⁢Ni 6 ⁢O 19 (𝑚 = 6) and at 𝐪 = (0.7⁢𝜋, 0.7⁢𝜋) for La 6⁢ Ni 5⁢ O 16 (𝑚 = 5). Along the out-of-plane direction, four layers are coupled as ↓−↑−↑−↓ in La 5 ⁢Ni 4⁢ O 13 , five layers are coupled as ↑−↑−↓−↑−↑ in La 6 ⁢Ni 5 ⁢O 16 , and six layers are coupled as ↑−↓−↓−↑−↑−↓ in La 7 ⁢Ni 6 ⁢O 19 .

Zhang, Yang [Univ. of Tennessee, Knoxville, TN (Un↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗