Search NASASearch

Engineering topics

Ke, Ruian

Publications and source records attributed to Ke, Ruian.

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES

The kinetics of SARS-CoV-2 infection based on a human challenge study

Studying the early events that occur after viral infection in humans is difficult unless one intentionally infects volunteers in a human challenge study. Here, we use data about severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) in such a study in combination with mathematical modeling to gain insights into the relationship between the amount of virus in the upper respiratory tract and the immune response it generates. We propose a set of dynamic models of increasing complexity to dissect the roles of target cell limitation, innate immunity, and adaptive immunity in determining the observed viral kinetics. We introduce an approach for modeling the effect of humoral immunity that describes a decline in infectious virus after immune activation. We fit our models to viral load and infectious titer data from all the untreated infected participants in the study simultaneously. We found that a power-law with a power h < 1 describes the relationship between infectious virus and viral load. Viral replication at the early stage of infection is rapid, with a doubling time of ~2 h for viral RNA and ~3 h for infectious virus. We estimate that adaptive immunity is initiated ~7 to 10 d postinfection and appears to contribute to a multiphasic viral decline experienced by some participants; the viral rebound experienced by other participants is consistent with a decline in the interferon response. Altogether, we quantified the kinetics of SARS-CoV-2 infection, shedding light on the early dynamics of the virus and the potential role of innate and adaptive immunity in promoting viral decline during infection.

59 BASIC BIOLOGICAL SCIENCES

CoVTransformer

The code is a transformer model to forecast SARS-CoV-2 lineage frequencies in the future.

Feng, Yinan