Search NASA⌕ Search

SEARCH · Search NASA

Results for “HIV epidemiology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology↗

Molecular Epidemiology of HIV-1 in Ghana: Subtype Distribution, Drug Resistance and Coreceptor Usage

The greatest HIV-1 genetic diversity is found in West/Central Africa due to the pandemic’s origins in this region, but this diversity remains understudied. We characterized HIV-1 subtype diversity (from both sub-genomic and full-genome viral sequences), drug resistance and coreceptor usage in 103 predominantly (90%) antiretroviral-naive individuals living with HIV-1 in Ghana. Full-genome HIV-1 subtyping confirmed the circulating recombinant form CRF02_AG as the dominant (53.9%) subtype in the region, with the complex recombinant 06_cpx (4%) present as well. Unique recombinants, most of which were mosaics containing CRF02_AG and/or 06_cpx, made up 37% of sequences, while “pure” subtypes were rare (<6%). Pretreatment resistance to at least one drug class was observed in 17% of the cohort, with NNRTI resistance being the most common (12%) and INSTI resistance being relatively rare (2%). CXCR4-using HIV-1 sequences were identified in 23% of participants. Overall, our findings advance our understanding of HIV-1 molecular epidemiology in Ghana. Extensive HIV-1 genetic diversity in the region appears to be fueling the ongoing creation of novel recombinants, the majority CRF02_AG-containing, in the region. The relatively high prevalence of pretreatment NNRTI resistance but low prevalence of INSTI resistance supports the use of INSTI-based first-line regimens in Ghana.

59 BASIC BIOLOGICAL SCIENCES↗

Combining biomarker and virus phylogenetic models improves HIV-1 epidemiological source identification

To identify and stop active HIV transmission chains new epidemiological techniques are needed. Here, we describe the development of a multi-biomarker augmentation to phylogenetic inference of the underlying transmission history in a local population. HIV biomarkers are measurable biological quantities that have some relationship to the amount of time someone has been infected with HIV. To train our model, we used five biomarkers based on real data from serological assays, HIV sequence data, and target cell counts in longitudinally followed, untreated patients with known infection times. The biomarkers were modeled with a mixed effects framework to allow for patient specific variation and general trends, and fit to patient data using Markov Chain Monte Carlo (MCMC) methods. Subsequently, the density of the unobserved infection time conditional on observed biomarkers were obtained by integrating out the random effects from the model fit. This probabilistic information about infection times was incorporated into the likelihood function for the transmission history and phylogenetic tree reconstruction, informed by the HIV sequence data. To critically test our methodology, we developed a coalescent-based simulation framework that generates phylogenies and biomarkers given a specific or general transmission history. Testing on many epidemiological scenarios showed that biomarker augmented phylogenetics can reach 90% accuracy under idealized situations. Under realistic within-host HIV-1 evolution, involving substantial within-host diversification and frequent transmission of multiple lineages, the average accuracy was at about 50% in transmission clusters involving 5–50 hosts. Realistic biomarker data added on average 16 percentage points over using the phylogeny alone. Using more biomarkers improved the performance. Shorter temporal spacing between transmission events and increased transmission heterogeneity reduced reconstruction accuracy, but larger clusters were not harder to get right. More sequence data per infected host also improved accuracy. We show that the method is robust to incomplete sampling and that adding biomarkers improves reconstructions of real HIV-1 transmission histories. The technology presented here could allow for better prevention programs by providing data for locally informed and tailored strategies.

60 APPLIED LIFE SCIENCES↗

Characterization of a novel HIV-1 circulating recombinant form, CRF91_cpx, comprising CRF02_AG, G, J, and U, mostly among men who have sex with men

Prospective molecular studies of HIV-1 pol region (2253–5250 in HXB2 genome) sequences from sequenced samples of 269 HIV-1-infected patients in Cyprus (2017–2021) revealed a transmission cluster of 14 unknown HIV-1 recombinants that were not classified as previously established CRFs. The earliest recombinant was collected in September 2017, and the transmission cluster continued to grow until November 2020. Near full-length HIV-1 genome sequences of the 11 of the 14 recombinants were successfully obtained (790–8795 in HXB2 genome) and aligned against a reference dataset of HIV-1 subtypes and CRFs. We employed MEGAX for maximum-likelihood tree construction (GTR model, 1000 bootstrap replicates), Cluster-Picker for phylogenetic clustering analysis (genetic distance ≤0.045, bootstrap support value ≥70%), and REGA-3.0 for subtype determination. Bootscan and similarity plot analyses (sliding window of 400 nucleotides overlapped by 40 nucleotides) were conducted using SimPlot-v3.5.1, and subregion confirmatory neighbour-joining tree analyses were conducted using MEGAX (Kimura two-parameter model, 1000 bootstrap replicates, ≥70% bootstrap-support value). Exclusive clustering of the HIV-1 recombinants revealed their uniqueness. The recombination analyses illustrated the same unique mosaic pattern with six putative intersubtype recombination breakpoints, seven fragments of subtypes CRF02_AG, G, J and an unclassified fragment. We conclusively characterized the mosaic structure of the novel HIV-1 CRF, named CRF91_cpx, by the Los Alamos HIV Sequence Database. Additionally, we identified a URF of CRF91_cpx with two additional recombination sites, generated by a recombination event between subtype B and CRF91_cpx. Since the identification of CRF91_cpx, two additional patient samples have been entered into the CRF91_cpx transmission cluster, demonstrating active growth.

60 APPLIED LIFE SCIENCES↗

Comprehensive Genetic Characterization of Four Novel HIV-1 Circulating Recombinant Forms (CRF129_56G, CRF130_A1B, CRF131_A1B, and CRF138_cpx): Insights from Molecular Epidemiology in Cyprus

Molecular investigations of the HIV-1 pol region (2253–5250 in the HXB2 genome) were conducted on sequences obtained from 331 individuals infected with HIV-1 in Cyprus between 2017 and 2021. This study unveiled four distinct HIV-1 putative transmission clusters, encompassing 19 previously unidentified HIV-1 recombinants. These recombinants, each comprising eight, three, four, and four sequences, respectively, did not align with previously established Circulating Recombinant Forms (CRFs). To characterize these novel HIV-1 recombinants, near-full-length genome sequences were successfully obtained for 16 of the 19 recombinants (790–8795 in the HXB2 genome) using an in-house-developed RT-PCR assay. Phylogenetic analyses, employing MEGAX and Cluster-Picker, along with confirmatory neighbor-joining tree analyses of subregions, were conducted to identify distinct clusters and determine subtypes. The uniqueness of the HIV-1 recombinants was evident in their exclusive clustering within generated maximum likelihood trees. Recombination analyses highlighted the distinct chimeric nature of these recombinants, with consistent mosaic patterns observed across all sequences within each of the four putative transmission clusters. Conclusive genetic characterization identified four novel HIV-1 CRFs: CRF129_56G, CRF130_A1B, CRF131_A1B, and CRF138_cpx. CRF129_56G exhibited two recombination breakpoints and three fragments of subtypes CRF56_cpx and G. Both CRF130_A1B and CRF131_A1B featured seven recombination breakpoints and eight fragments of subtypes A1 and B. CRF138_cpx displayed five recombination breakpoints and six fragments of subtypes CRF22_01A1 and F2, along with an unclassified fragment. Additional BLAST analyses identified a Unique Recombinant Form (URF) of CRF138_cpx with three additional recombination sites, involving subtype F2, a fragment of unknown subtype origin, and CRF138_cpx. Post-identification, all putative transmission clusters remained active, with CRF130_A1B, CRF131_A1B, and CRF138_cpx clusters exhibiting further growth. Furthermore, international connections were identified through BLAST analyses, linking one sequence from the USA to the CRF130_A1B strain, and three sequences from Belgium and Cameroon to the CRF138_cpx strain. This study contributes valuable insights into the dynamic landscape of HIV-1 diversity and transmission patterns, emphasizing the need for ongoing molecular surveillance and global collaboration in tracking emerging viral variants.

60 APPLIED LIFE SCIENCES↗

A deep learning approach to real-time HIV outbreak detection using genetic data

Pathogen genomic sequence data are increasingly made available for epidemiological monitoring. A main interest is to identify and assess the potential of infectious disease outbreaks. While popular methods to analyze sequence data often involve phylogenetic tree inference, they are vulnerable to errors from recombination and impose a high computational cost, making it difficult to obtain real-time results when the number of sequences is in or above the thousands. Here, we propose an alternative strategy to outbreak detection using genomic data based on deep learning methods developed for image classification. The key idea is to use a pairwise genetic distance matrix calculated from viral sequences as an image, and develop convolutional neutral network (CNN) models to classify areas of the images that show signatures of active outbreak, leading to identification of subsets of sequences taken from an active outbreak. We showed that our method is efficient in finding HIV-1 outbreaks with R0 ≥ 2.5, and overall a specificity exceeding 98% and sensitivity better than 92%. We validated our approach using data from HIV-1 CRF01 in Europe, containing both endemic sequences and a well-known dual outbreak in intravenous drug users. Our model accurately identified known outbreak sequences in the background of slower spreading HIV. Importantly, we detected both outbreaks early on, before they were over, implying that had this method been applied in real-time as data became available, one would have been able to intervene and possibly prevent the extent of these outbreaks. This approach is scalable to processing hundreds of thousands of sequences, making it useful for current and future real-time epidemiological investigations, including public health monitoring using large databases and especially for rapid outbreak identification.

59 BASIC BIOLOGICAL SCIENCES↗

Large Evolutionary Rate Heterogeneity among and within HIV-1 Subtypes and CRFs

HIV-1 is a fast-evolving, genetically diverse virus presently classified into several groups and subtypes. The virus evolves rapidly because of an error-prone polymerase, high rates of recombination, and selection in response to the host immune system and clinical management of the infection. The rate of evolution is also influenced by the rate of virus spread in a population and nature of the outbreak, among other factors. HIV-1 evolution is thus driven by a range of complex genetic, social, and epidemiological factors that complicates disease management and prevention. Here, we quantify the evolutionary (substitution) rate heterogeneity among major HIV-1 subtypes and recombinants by analyzing the largest collection of HIV-1 genetic data spanning the widest possible geographical (100 countries) and temporal (1981–2019) spread. We show that HIV-1 substitution rates vary substantially, sometimes by several folds, both across the virus genome and between major subtypes and recombinants, but also within a subtype. Across subtypes, rates ranged 3.5-fold from 1.34 × 10−3 to 4.72 × 10−3 in env and 2.3-fold from 0.95 × 10−3 to 2.18 × 10−3 substitutions site−1 year−1 in pol. Within the subtype, 3-fold rate variation was observed in env in different human populations. It is possible that HIV-1 lineages in different parts of the world are operating under different selection pressures leading to substantial rate heterogeneity within and between subtypes. We further highlight how such rate heterogeneity can complicate HIV-1 phylodynamic studies, specifically, inferences on epidemiological linkage of transmission clusters based on genetic distance or phylogenetic data, and can mislead estimates about the timing of HIV-1 lineages.

59 BASIC BIOLOGICAL SCIENCES↗

Intra- and inter-subtype HIV diversity between 1994 and 2018 in southern Uganda: a longitudinal population-based study

There is limited data on human immunodeficiency virus (HIV) evolutionary trends in African populations. We evaluated changes in HIV viral diversity and genetic divergence in southern Uganda over a 24-year period spanning the introduction and scale-up of HIV prevention and treatment programs using HIV sequence and survey data from the Rakai Community Cohort Study, an open longitudinal population-based HIV surveillance cohort. Gag (p24) and env (gp41) HIV data were generated from people living with HIV (PLHIV) in 31 inland semi-urban trading and agrarian communities (1994–2018) and four hyperendemic Lake Victoria fishing communities (2011–2018) under continuous surveillance. HIV subtype was assigned using the Recombination Identification Program with phylogenetic confirmation. Inter-subtype diversity was evaluated using the Shannon diversity index, and intra-subtype diversity with the nucleotide diversity and pairwise TN93 genetic distance. Genetic divergence was measured using root-to-tip distance and pairwise TN93 genetic distance analyses. Demographic history of HIV was inferred using a coalescent-based Bayesian Skygrid model. Evolutionary dynamics were assessed among demographic and behavioral population subgroups, including by migration status. 9931 HIV sequences were available from 4999 PLHIV, including 3060 and 1939 persons residing in inland and fishing communities, respectively. In inland communities, subtype A1 viruses proportionately increased from 14.3% in 1995 to 25.9% in 2017 (P < .001), while those of subtype D declined from 73.2% in 1995 to 28.2% in 2017 (P < .001). The proportion of viruses classified as recombinants significantly increased by nearly four-fold from 12.2% in 1995 to 44.8% in 2017. Inter-subtype HIV diversity has generally increased. While intra-subtype p24 genetic diversity and divergence leveled off after 2014, intra-subtype gp41 diversity, effective population size, and divergence increased through 2017. Intra- and inter-subtype viral diversity increased across all demographic and behavioral population subgroups, including among individuals with no recent migration history or extra-community sexual partners. This study provides insights into population-level HIV evolutionary dynamics following the scale-up of HIV prevention and treatment programs. Continued molecular surveillance may provide a better understanding of the dynamics driving population HIV evolution and yield important insights for epidemic control and vaccine development.

60 APPLIED LIFE SCIENCES↗

Identification of a new circulating recombinant form of human immunodeficiency virus type 1, CRF124_cpx involving subtypes A, G, H, and CRF27_cpx in Angola

Angola, located in Central Africa, has around 320,000 (270,000–380,000) people living with human immunodeficiency virus (HIV)/AIDS, equivalent to 1% of the country’s population at the end of 2021. A previous study conducted in 2012, using Angolan samples collected between 2008 and 2010 revealed a high prevalence of HIV-1 recombinants, around 42% of sequences, with 21% showing the same UH profile in partial pol region which were grouped into a monophyletic cluster with high bootstrap support. Thus, the objective of the present work was to obtain complete genomes of those sequences and characterize them, aiming at a description of a new circulating recombinant form (CRF). Whole blood from nine HIV-1 UH pol -infected individuals had their genomic DNA extracted, and nested PCR was used to amplify seven overlapping fragments targeting the full-length HIV-1 genome. The final classification was based on maximum likelihood trees, and recombination analyses were performed using a bootscan from the Simplot program. BLAST and Los Alamos Database inspections were used to search other similar H-like pol sequences. Complete genome amplification was possible for three samples, partial genomes were obtained for the other three, and only pol was available for the remaining three sequences. Bootscan analysis of the two whole-genome and three partial genome sequences retrieved from people living with HIV/AIDS (PLHIVA) without epidemiological linkage showed the same complex recombination profile involving HIV-1 subtypes A/G/H/CRF27_cpx, with a total of six recombinant breakpoints, aiming to classify a new HIV-1 CRF124_cpx. We found no other full-length HIV-1 genomes with the same mosaic profile; however, we identified 33 partial pol sequences, mainly sampled from Angola between 2001 to 2019, with the same H-like profile. Bayesian analysis of H and H-like pol sequences indicates that CRF124_cpx probably originated in Angola at mid-1970s, indicating that this CRF has been circulating in the country for a long time. In summary, our study describes a new CRF circulating principally in Angola and highlights the importance of continuing molecular surveillance studies, especially in countries with high molecular diversity of HIV.

60 APPLIED LIFE SCIENCES↗

Updated HIV-1 Consensus Sequences Change but Stay Within Similar Distance From Worldwide Samples

HIV consensus sequences are used in various bioinformatic, evolutionary, and vaccine related research. Since the previous HIV-1 subtype and CRF consensus sequences were constructed in 2002, the number of publicly available HIV-1 sequences have grown exponentially, especially from non-EU and US countries. Here, we reconstruct 90 new HIV-1 subtype and CRF consensus sequences from 3,470 high-quality, representative, full genome sequences in the LANL HIV database. While subtypes and CRFs are unevenly spread across the world, in total 89 countries were represented. For consensus sequences that were based on at least 20 genomes, we found that on average 2.3% (range 0.8–10%) of the consensus genome site states changed from 2002 to 2021, of which about half were nucleotide state differences and the rest insertions and deletions. Interestingly, the 2021 consensus sequences were shorter than in 2002, and compared to 4,674 HIV-1 worldwide genome sequences, the 2021 consensuses were somewhat closer to the worldwide genome sequences, i.e., showing on average fewer nucleotide state differences. Some subtypes/CRFs have had limited geographical spread, and thus sampling of subtypes/CRFs is uneven, at least in part, due to the epidemiological dynamics. Thus, taken as a whole, the 2021 consensus sequences likely are good representations of the typical subtype/CRF genome nucleotide states. The new consensus sequences are available at the LANL HIV database.

60 APPLIED LIFE SCIENCES↗

How robust are estimates of key parameters in standard viral dynamic models?

Mathematical models of viral infection have been developed, fitted to data, and provide insight into disease pathogenesis for multiple agents that cause chronic infection, including HIV, hepatitis C, and B virus. However, for agents that cause acute infections or during the acute stage of agents that cause chronic infections, viral load data are often collected after symptoms develop, usually around or after the peak viral load. Consequently, we frequently lack data in the initial phase of viral growth, i.e., when pre-symptomatic transmission events occur. Missing data may make estimating the time of infection, the infectious period, and parameters in viral dynamic models, such as the cell infection rate, difficult. However, having extra information, such as the average time to peak viral load, may improve the robustness of the estimation. Here, we evaluated the robustness of estimates of key model parameters when viral load data prior to the viral load peak is missing, when we know the values of some parameters and/or the time from infection to peak viral load. Although estimates of the time of infection are sensitive to the quality and amount of available data, particularly pre-peak, other parameters important in understanding disease pathogenesis, such as the loss rate of infected cells, are less sensitive. Viral infectivity and the viral production rate are key parameters affecting the robustness of data fits. Fixing their values to literature values can help estimate the remaining model parameters when pre-peak data is missing or limited. We find a lack of data in the pre-peak growth phase underestimates the time to peak viral load by several days, leading to a shorter predicted growth phase. On the other hand, knowing the time of infection (e.g., from epidemiological data) and fixing it results in good estimates of dynamical parameters even in the absence of early data. While we provide ways to approximate model parameters in the absence of early viral load data, our results also suggest that these data, when available, are needed to estimate model parameters more precisely.

59 BASIC BIOLOGICAL SCIENCES↗