Search NASA⌕ Search

SEARCH · Search NASA

Results for “viral genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CheckV assesses the quality and completeness of metagenome-assembled viral genomes

Abstract Millions of new viral sequences have been identified from metagenomes, but the quality and completeness of these sequences vary considerably. Here we present CheckV, an automated pipeline for identifying closed viral genomes, estimating the completeness of genome fragments and removing flanking host regions from integrated proviruses. CheckV estimates completeness by comparing sequences with a large database of complete viral genomes, including 76,262 identified from a systematic search of publicly available metagenomes, metatranscriptomes and metaviromes. After validation on mock datasets and comparison to existing methods, we applied CheckV to large and diverse collections of metagenome-assembled viral sequences, including IMG/VR and the Global Ocean Virome. This revealed 44,652 high-quality viral genomes (that is, >90% complete), although the vast majority of sequences were small fragments, which highlights the challenge of assembling viral genomes from short-read metagenomes. Additionally, we found that removal of host contamination substantially improved the accurate identification of auxiliary metabolic genes and interpretation of viral-encoded functions.

59 BASIC BIOLOGICAL SCIENCES↗

Tunturi virus isolates and metagenome-assembled viral genomes provide insights into the virome of Acidobacteriota in Arctic tundra soils

Arctic soils are climate-critical areas, where microorganisms play crucial roles in nutrient cycling processes. Acidobacteriota are phylogenetically and physiologically diverse bacteria that are abundant and active in Arctic tundra soils. Still, surprisingly little is known about acidobacterial viruses in general and those residing in the Arctic in particular. Here, we applied both culture-dependent and -independent methods to study the virome of Acidobacteriota in Arctic soils. Five virus isolates, Tunturi 1–5, were obtained from Arctic tundra soils, Kilpisjärvi, Finland (69°N), using Tunturiibacter spp. strains originating from the same area as hosts. The new virus isolates have tailed particles with podo- (Tunturi 1, 2, 3), sipho- (Tunturi 4), or myovirus-like (Tunturi 5) morphologies. The dsDNA genomes of the viral isolates are 63–98 kbp long, except Tunturi 5, which is a jumbo phage with a 309-kbp genome. Tunturi 1 and Tunturi 2 share 88% overall nucleotide identity, while the other three are not related to one another. For over half of the open reading frames in Tunturi genomes, no functions could be predicted. To further assess the Acidobacteriota-associated viral diversity in Kilpisjärvi soils, bulk metagenomes from the same soils were explored and a total of 1881 viral operational taxonomic units (vOTUs) were bioinformatically predicted. Almost all vOTUs (98%) were assigned to the class Caudoviricetes. For 125 vOTUs, including five (near-)complete ones, Acidobacteriota hosts were predicted. Acidobacteriota-linked vOTUs were abundant across sites, especially in fens. Terriglobia-associated proviruses were observed in Kilpisjärvi soils, being related to proviruses from distant soils and other biomes. Approximately genus- or higher-level similarities were found between the Tunturi viruses, Kilpisjärvi vOTUs, and other soil vOTUs, suggesting some shared groups of Acidobacteriota viruses across soils. This study provides acidobacterial virus isolates as laboratory models for future research and adds insights into the diversity of viral communities associated with Acidobacteriota in tundra soils. Predicted virus-host links and viral gene functions suggest various interactions between viruses and their host microorganisms. Largely unknown sequences in the isolates and metagenome-assembled viral genomes highlight a need for more extensive sampling of Arctic soils to better understand viral functions and contributions to ecosystem-wide cycling processes in the Arctic.

54 ENVIRONMENTAL SCIENCES↗

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology↗

Visualizing a viral genome with contrast variation small angle X-ray scattering

Despite the threat to human health posed by some single-stranded RNA viruses, little is understood about their assembly. The goal of this work is to introduce a new tool for watching an RNA genome direct its own packaging and encapsidation by proteins. Contrast variation small-angle X-ray scattering (CV-SAXS) is a powerful tool with the potential to monitor the changing structure of a viral RNA through this assembly process. The proteins, though present, do not contribute to the measured signal. As a first step in assessing the feasibility of viral genome studies, the structure of encapsidated MS2 RNA was exclusively detected with CV-SAXS and compared with a structure derived from asymmetric cryo-EM reconstructions. Additional comparisons with free RNA highlight the significant structural rearrangements induced by capsid proteins and invite the application of time-resolved CV-SAXS to reveal interactions that result in efficient viral assembly.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding standards in viromics: in silico evaluation of dsDNA viral genome identification, classification, and auxiliary metabolic gene curation

Viruses influence global patterns of microbial diversity and nutrient cycles. Though viral metagenomics (viromics), specifically targeting dsDNA viruses, has been critical for revealing viral roles across diverse ecosystems, its analyses differ in many ways from those used for microbes. To date, viromics benchmarking has covered read pre-processing, assembly, relative abundance, read mapping thresholds and diversity estimation, but other steps would benefit from benchmarking and standardization. Here we use in silico-generated datasets and an extensive literature survey to evaluate and highlight how dataset composition (i.e., viromes vs bulk metagenomes) and assembly fragmentation impact (i) viral contig identification tool, (ii) virus taxonomic classification, and (iii) identification and curation of auxiliary metabolic genes (AMGs). The in silico benchmarking of five commonly used virus identification tools show that gene-content-based tools consistently performed well for long (≥3 kbp) contigs, while k -mer- and blast-based tools were uniquely able to detect viruses from short (≤3 kbp) contigs. Notably, however, the performance increase of k -mer- and blast-based tools for short contigs was obtained at the cost of increased false positives (sometimes up to ~5% for virome and ~75% bulk samples), particularly when eukaryotic or mobile genetic element sequences were included in the test datasets. Furthermore, for viral classification, variously sized genome fragments were assessed using gene-sharing network analytics to quantify drop-offs in taxonomic assignments, which revealed correct assignations ranging from ~95% (whole genomes) down to ~80% (3 kbp sized genome fragments). A similar trend was also observed for other viral classification tools such as VPF-class, ViPTree and VIRIDIC, suggesting that caution is warranted when classifying short genome fragments and not full genomes. Finally, we highlight how fragmented assemblies can lead to erroneous identification of AMGs and outline a best-practices workflow to curate candidate AMGs in viral genomes assembled from metagenomes. Together, these benchmarking experiments and annotation guidelines should aid researchers seeking to best detect, classify, and characterize the myriad viruses ‘hidden’ in diverse sequence datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis for cloverleaf RNA-initiated viral genome replication

The genomes of positive-strand RNA viruses serve as a template for both protein translation and genome replication. In enteroviruses, a cloverleaf RNA structure at the 5' end of the genome functions as a switch to transition from viral translation to replication by interacting with host poly(C)-binding protein 2 (PCBP2) and the viral 3CD pro protein. We determined the structures of cloverleaf RNA from coxsackievirus and poliovirus. Cloverleaf RNA folds into an H-type four-way junction and is stabilized by a unique adenosine-cytidine-uridine (A•C-U) base triple involving the conserved pyrimidine mismatch region. The two PCBP2 binding sites are spatially proximal and are located on the opposite end from the 3CD pro binding site on cloverleaf. We determined that the A•C-U base triple restricts the flexibility of the cloverleaf stem–loops resulting in partial occlusion of the PCBP2 binding site, and elimination of the A•C-U base triple increases the binding affinity of PCBP2 to the cloverleaf RNA. Based on the cloverleaf structures and biophysical assays, we propose a new mechanistic model by which enteroviruses use the cloverleaf structure as a molecular switch to transition from viral protein translation to genome replication.

59 BASIC BIOLOGICAL SCIENCES↗

$\mathrm{COBRA}$ improves the completeness and contiguity of viral genomes assembled from metagenomes

Viruses are often studied using metagenome-assembled sequences, but genome incompleteness hampers comprehensive and accurate analyses. Contig Overlap Based Re-Assembly (COBRA) resolves assembly breakpoints based on the de Bruijn graph and joins contigs. Here we benchmarked COBRA using ocean and soil viral datasets. COBRA accurately joined the assembled sequences and achieved notably higher genome accuracy than binning tools. From 231 published freshwater metagenomes, we obtained 7,334 bacteriophage clusters, ~83% of which represent new phage species. Notably, ~70% of these were circular, compared with 34% before COBRA analyses. We expanded sampling of huge phages (≥200 kbp), the largest of which was curated to completion (717 kbp). Improved phage genomes from Rotsee Lake provided context for metatranscriptomic data and indicated the in situ activity of huge phages, whiB-encoding phages and cysC- and cysH-encoding phages. COBRA improves viral genome assembly contiguity and completeness, thus the accuracy and reliability of analyses of gene content, diversity and evolution.

54 ENVIRONMENTAL SCIENCES↗

Propagation of viral genomes by replicating ammonia-oxidising archaea during soil nitrification

Ammonia-oxidising archaea (AOA) are a ubiquitous component of microbial communities and dominate the first stage of nitrification in some soils. While we are beginning to understand soil virus dynamics, we have no knowledge of the composition or activity of those infecting nitrifiers or their potential to influence processes. This study aimed to characterise viruses having infected autotrophic AOA in two nitrifying soils of contrasting pH by following transfer of assimilated CO 2 -derived 13 C from host to virus via DNA stable-isotope probing and metagenomic analysis. Incorporation of 13 C into low GC mol% AOA and virus genomes increased DNA buoyant density in CsCl gradients but resulted in co-migration with dominant non-enriched high GC mol% genomes, reducing sequencing depth and contig assembly. We therefore developed a hybrid approach where AOA and virus genomes were assembled from low buoyant density DNA with subsequent mapping of 13 C isotopically enriched high buoyant density DNA reads to identify activity of AOA. Metagenome-assembled genomes were different between the two soils and represented a broad diversity of active populations. Sixty-four AOA-infecting viral operational taxonomic units (vOTUs) were identified with no clear relatedness to previously characterised prokaryote viruses. These vOTUs were also distinct between soils, with 42% enriched in 13 C derived from hosts. The majority were predicted as capable of lysogeny and auxiliary metabolic genes included an AOA-specific multicopper oxidase suggesting infection may augment copper uptake essential for central metabolic functioning. These findings indicate virus infection of AOA may be a frequent process during nitrification with potential to influence host physiology and activity.

59 BASIC BIOLOGICAL SCIENCES↗

The Number and Pattern of Viral Genomic Reassortments are not Necessarily Identifiable from Segment Trees

Reassortment is an evolutionary process common in viruses with segmented genomes. These viruses can swap whole genomic segments during cellular co-infection, giving rise to novel progeny formed from the mixture of parental segments. Since large-scale genome rearrangements have the potential to generate new phenotypes, reassortment is important to both evolutionary biology and public health research. However, statistical inference of the pattern of reassortment events from phylogenetic data is exceptionally difficult, potentially involving inference of general graphs in which individual segment trees are embedded. In this paper, we argue that, in general, the number and pattern of reassortment events are not identifiable from segment trees alone, even with theoretically ideal data. We call this fact the fundamental problem of reassortment, which we illustrate using the concept of the “first-infection tree,” a potentially counterfactual genealogy that would have been observed in the segment trees had no reassortment occurred. Further, we illustrate four additional problems that can arise logically in the inference of reassortment events and show, using simulated data, that these problems are not rare and can potentially distort our observation of reassortment even in small data sets. Finally, we discuss how existing methods can be augmented or adapted to account for not only the fundamental problem of reassortment, but also the four additional situations that can complicate the inference of reassortment.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling the Influenza A NP-vRNA-Polymerase Complex in Atomic Detail

Seasonal flu is an acute respiratory disease that exacts a massive toll on human populations, healthcare systems and economies. The disease is caused by an enveloped Influenza virus containing eight ribonucleoprotein (RNP) complexes. Each RNP incorporates multiple copies of nucleoprotein (NP), a fragment of the viral genome (vRNA), and a viral RNA-dependent RNA polymerase (POL), and is responsible for packaging the viral genome and performing critical functions including replication and transcription. A complete model of an Influenza RNP in atomic detail can elucidate the structural basis for viral genome functions, and identify potential targets for viral therapeutics. In this work we construct a model of a complete Influenza A RNP complex in atomic detail using multiple sources of structural and sequence information and a series of homology-modeling techniques, including a motif-matching fragment assembly method. Our final model provides a rationale for experimentally-observed changes to viral polymerase activity in numerous mutational assays. Further, our model reveals specific interactions between the three primary structural components of the RNP, including potential targets for blocking POL-binding to the NP-vRNA complex. The methods developed in this work open the possibility of elucidating other functionally-relevant atomic-scale interactions in additional RNP structures and other biomolecular complexes.

59 BASIC BIOLOGICAL SCIENCES↗

FutureTense

Protective vaccines and reliable diagnostics are essential tools for controlling viral diseases. However, the efficacy of these tools can be diminished by mutations in viral genomes. The delay between the emergence of new viral strains and the redesign of vaccines and diagnostics allows for continued viral transmission. Is it possible to address this challenge by computationally predicting viral genome sequence evolution? Can we “future-proof” vaccines and diagnostics by targeting both current and anticipated future sequence variants? While predicting viral evolution is still an unsolved, “grand challenge” problem in biology, the large, and rapidly growing, number of SARS-CoV-2 genome sequences provide an opportunity to quantify the ability of machine learning to predict viral genome sequence evolution. Towards this end, we have developed a simple computational model for predicting viral evolution at the level of individual nucleotides. The key metric for quantifying the per-base, prediction accuracy for viral evolution is the Mann-Whitney U statistic (or, equivalently, the area under the receiver operator curve). Since the Mann-Whitney U statistic is not a differentiable function, existing deep leaning packages (like Pytorch and Keras/TensorFlow) are not useful, as they require that the accuracy metric/objective function be analytically differentiable with respect to the model parameters. To overcome this challenge, we have implemented custom software, “FutureTense”, that can train a machine learning model by maximizing the non-differentiable Mann-Whitney U statistic. This software trains a machine learning model by exploring along the direction of the discrete gradient of the Mann-Whitney U statistic in the model parameter space. Parallel computing and genome sequence-specific optimizations are used to accelerate model training. The resulting machine learning model learns the observed high C->U mutation rates in the SARS-CoV-2 genome (which are potentially induced by host defenses) and provides prediction accuracies that are significantly better than one would expect from random chance. While predicting viral evolution is still quite far from a solved problem, the surprising performance of this simple model gives hope that the accuracy of predicting viral genome evolution can be further increased by more sophisticated approaches.

Gans, Jason↗

Recovering new viruses from New Mexico soils

Here, we utilized metagenomic and size-filtered virome sequencing to recover 4,157 medium, high, or complete quality viral genomes from soils taken from three high elevation sites in New Mexico, USA. Among recovered viral genomes, 90% were from size-filtered samples, indicating the importance of this enrichment in assessments of complex viromes.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid assembly of SARS-CoV-2 genomes reveals attenuation of the Omicron BA.1 variant through NSP6

Although the SARS-CoV-2 Omicron variant (BA.1) spread rapidly across the world and effectively evaded immune responses, its viral fitness in cell and animal models was reduced. The precise nature of this attenuation remains unknown as generating replication-competent viral genomes is challenging because of the length of the viral genome (~30 kb). Here, we present a plasmid-based viral genome assembly and rescue strategy (pGLUE) that constructs complete infectious viruses or noninfectious subgenomic replicons in a single ligation reaction with >80% efficiency. Fully sequenced replicons and infectious viral stocks can be generated in 1 and 3 weeks, respectively. By testing a series of naturally occurring viruses as well as Delta-Omicron chimeric replicons, we show that Omicron nonstructural protein 6 harbors critical attenuating mutations, which dampen viral RNA replication and reduce lipid droplet consumption. Thus, pGLUE overcomes remaining barriers to broadly study SARS-CoV-2 replication and reveals deficits in nonstructural protein function underlying Omicron attenuation.

60 APPLIED LIFE SCIENCES↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Full scale structural, mechanical and dynamical properties of HIV-1 liposomes

Enveloped viruses are enclosed by a lipid membrane inside of which are all of the components necessary for the virus life cycle; viral proteins, the viral genome and metabolites. Viral envelopes are lipid bilayers that adopt morphologies ranging from spheres to tubes. The envelope is derived from the host cell during viral replication. Thus, the composition of the bilayer depends on the complex constitution of lipids from the host-cell’s organelle(s) where assembly and/or budding of the viral particle occurs. Here, molecular dynamics (MD) simulations of authentic, asymmetric HIV-1 liposomes are used to derive a unique level of resolution of its full-scale structure, mechanics and dynamics. Analysis of the structural properties reveal the distribution of thicknesses of the bilayers over the entire liposome as well as its global fluctuations. Moreover, full-scale mechanical analyses are employed to derive the global bending rigidity of HIV-1 liposomes. Finally, dynamical properties of the lipid molecules reveal important relationships between their 3D diffusion, the location of lipid-rafts and the asymmetrical composition of the envelope. Overall, our simulations reveal complex relationships between the rich lipid composition of the HIV-1 liposome and its structural, mechanical and dynamical properties with critical consequences to different stages of HIV-1’s life cycle.

59 BASIC BIOLOGICAL SCIENCES↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

Multi-choice Viromics Pipeline (MVP) v1

MVP stands for Multi-choice Viromics Pipeline. It is a pipeline that utilizes a suite of state-of-art tools: geNomad to identify viruses, proviruses, and plasmids in sequencing data, CheckV to assess the quality, and completeness of identified viral genomes, including identification of host contamination for integrated proviruses, A custom code for a rapid genome clustering based on pairwise ANI, Bowtie2, Samtools, and CoverM to calculate coverage of individual viral genomes by read mapping, A custom code to create a vOTU table of abundance, MMseqs2 to compare viral proteins to multiple databases. It provides a quick, and intuitive pipeline to get viral sequences and corresponding properties that can be used for downstream analyses.

Roux, Simon↗