Search NASA⌕ Search

SEARCH · Search NASA

Results for “structural genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An orthologous gene coevolution network provides insight into eukaryotic cellular and genomic structure and function

The evolutionary rates of functionally related genes often covary. We present a gene coevolution network inferred from examining nearly 3 million orthologous gene pairs from 332 budding yeast species spanning ~400 million years of evolution. Network modules provide insight into cellular and genomic structure and function. Examination of the phenotypic impact of network perturbation using deletion mutant data from the baker’s yeast Saccharomyces cerevisiae, which were obtained from previously published studies, suggests that fitness in diverse environments is affected by orthologous gene neighborhood and connectivity. Mapping the network onto the chromosomes of S. cerevisiae and Candida albicans revealed that coevolving orthologous genes are not physically clustered in either species; rather, they are often located on different chromosomes or far apart on the same chromosome. The coevolution network captures the hierarchy of cellular structure and function, provides a roadmap for genotype-to-phenotype discovery, and portrays the genome as a linked ensemble of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Structural genomics of bacterial drug targets: Application of a high-throughput pipeline to solve 58 protein structures from pathogenic and related bacteria

Antibiotic resistance remains a leading cause of severe infections worldwide. Small changes in protein sequence can impact antibiotic efficacy. Here, we report deposition of 58 X-ray crystal structures of bacterial proteins that are known targets for antibiotics, which expands knowledge of structural variation to support future antibiotic discovery or modifications.

PDB↗

Chromatin structures from integrated AI and polymer physics model

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure fromindirectmeasures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Biochemistry & Molecular Biology↗

Genomes to Structure and Function Workshop Report 2022

The goal of the U.S. Department of Energy (DOE) Biological and Environmental Research (BER) Program is to achieve a predictive understanding of complex biological, earth, and environmental systems with the aim of advancing the nation’s energy and infrastructure security. (https://www.energy.gov/science/ ber/biological-and-environmental-research). To pursue this goal, collaborations among experts in diverse research areas that lead to multidisciplinary projects are indispensable. The roles of DOE’s User Facilities, which offer unique and powerful resources for such research projects, are evolving, and expectations for the facilities are increasing. To respond to Users’ needs, the Joint Genome Institute (JGI) and Environmental Molecular Sciences Laboratory (EMSL) initiated the Facilities Integrating Collaborations for User Science (FICUS) program in 2014. This collaboration has grown into a popular and successful program, advancing more than 100 multidisciplinary projects to date. Similarly, the new interFacility collaborations among the JGI, EMSL, and User resources for BER structural biology and imaging at the Basic Energy Science (BES) Program’s synchrotron and neutron facilities are becoming essential for cutting-edge transdisciplinary science. To further explore the need for the BER research community to combine genomic, functional, and structural approaches to advance their research, an organizing committee was formed to develop and jointly host a 3-part workshop. The committee’s members represented seven DOE National Laboratory User Facilities (Appendix 1 lists the members). The “Genomes to Structure and Function” virtual workshop (see Appendices 2–5) was composed of three sessions. The first session, titled “Molecular Structures” (October 27– 28, 2021), highlighted diverse integrative experimental and computational approaches correlating structural data with sequencing and functional information, as well as predicting protein structures to model complex biological systems. The second session, “Intracellular Organization, and Material Synthesis and Decomposition” (December 15–16, 2021), covered imaging methods for observing, quantifying, and manipulating biosystems. The third session, “Imaging the Rhizosphere and Cellular Organization” (January 26–27, 2022) emphasized advanced and non-invasive imaging techniques applied to plant root-microbe-soil interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Giant Starship Elements Mobilize Accessory Genes in Fungal Genomes

Accessory genes are variably present among members of a species and are a reservoir of adaptive functions. In bacteria, differences in gene distributions among individuals largely result from mobile elements that acquire and disperse accessory genes as cargo. In contrast, the impact of cargo-carrying elements on eukaryotic evolution remains largely unknown. Here, we show that variation in genome content within multiple fungal species is facilitated by Starships, a newly discovered group of massive mobile elements that are 110 kb long on average, share conserved components, and carry diverse arrays of accessory genes. We identified hundreds of Starship-like regions across every major class of filamentous Ascomycetes, including 28 distinct Starships that range from 27 to 393 kb and last shared a common ancestor ca. 400 Ma. Using new long-read assemblies of the plant pathogen Macrophomina phaseolina, we characterize four additional Starships whose activities contribute to standing variation in genome structure and content. One of these elements, Voyager, inserts into 5S rDNA and contains a candidate virulence factor whose increasing copy number has contrasting associations with pathogenic and saprophytic growth, suggesting Voyager’s activity underlies an ecological trade-off. We propose that Starships are eukaryotic analogs of bacterial integrative and conjugative elements based on parallels between their conserved components and may therefore represent the first dedicated agents of active gene transfer in eukaryotes. Our results suggest that Starships have shaped the content and structure of fungal genomes for millions of years and reveal a new concerted route for evolution throughout an entire eukaryotic phylum.

59 BASIC BIOLOGICAL SCIENCES↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Genomic patterns of structural variation among diverse genotypes of Sorghum bicolor and a potential role for deletions in local adaptation

Genomic structural mutations, especially deletions, are an important source of variation in many species and can play key roles in phenotypic diversification and evolution. Previous work in many plant species has identified multiple instances of structural variations (SVs) occurring in or near genes related to stress response and disease resistance, suggesting a possible role for SVs in local adaptation. Sorghum [Sorghum bicolor (L.) Moench] is one of the most widely grown cereal crops in the world. It has been adapted to an array of different climates as well as bred for multiple purposes, resulting in a striking phenotypic diversity. In this study, we identified genome-wide SVs in the Biomass Association Panel, a collection of 347 diverse sorghum genotypes collected from multiple countries and continents. Using Illumina-based, short-read whole-genome resequencing data from every genotype, we found a total of 24,648 SVs, including 22,359 deletions. The global site frequency spectrum of deletions and other types of SVs fit a model of neutral evolution, suggesting that the majority of these mutations were not under any types of selection. Clustering results based on single nucleotide polymorphisms separated the genotypes into eight clusters which largely corresponded with geographic origins, with many of the large deletions we uncovered being unique to a single cluster. Even though most deletions appeared to be neutral, a handful of cluster-specific deletions were found in genes related to biotic and abiotic stress responses, supporting the possibility that at least some of these deletions contribute to local adaptation in sorghum.

59 BASIC BIOLOGICAL SCIENCES↗

SvAnna: efficient and accurate pathogenicity prediction of coding and regulatory structural variants in long-read genome sequencing

Structural variants (SVs) are implicated in the etiology of Mendelian diseases but have been systematically underascertained owing to sequencing technology limitations. Long-read sequencing enables comprehensive detection of SVs, but approaches for prioritization of candidate SVs are needed. Structural variant Annotation and analysis (SvAnna) assesses all classes of SVs and their intersection with transcripts and regulatory sequences, relating predicted effects on gene function with clinical phenotype data. SvAnna places 87% of deleterious SVs in the top ten ranks. The interpretable prioritizations offered by SvAnna will facilitate the widespread adoption of long-read sequencing in diagnostic genomics. SvAnna is available at https://github.com/TheJacksonLaboratory/SvAnna.

59 BASIC BIOLOGICAL SCIENCES↗

Extreme dimensions — how big (or small) can tailed phages be?

Bacteriophages (phages), that is, viruses that infect bacteria, represent an extremely diverse, yet under-characterized, group of viruses. Although most known phages harbour genomes that are shorter than 200 kb packaged into capsids with a diameter under 100 Å, more and more ‘extremely large’ phages are being discovered. Early reports of phages with genomes larger than 200 kb defined them as ‘jumbo phages’1. Phylogenetic analysis of these jumbo phages revealed that they form distinct and cohesive groups with multiple independent origins. This suggested that jumbo phages are not merely mistakes or aberrations of smaller phages that got too big, but that large genome size is generally a stable trait. In fact, a larger genome can be advantageous; almost all jumbo phages encode their own specific transcription factors, at least some parts of the replication machinery, and many also have their own tRNA genes, which enables increased independence from their host. Here, large genomes also provide more flexibility in terms of transcription strategy; smaller phage genomes tend to have a strict modular genome structure to increase efficiency, whereas jumbo phage genomes may be much less organized in the absence of a strong selection pressure to maintain a compact genome.

59 BASIC BIOLOGICAL SCIENCES↗

A haplotype-resolved, chromosome-scale genome assembly for the southern live oak, Quercus virginiana

Hybridization is a major force driving diversification, migration, and adaptation in Quercus species. While population genetics and phylogenetics have traditionally been used for studying these processes, advances in sequencing technology now enable us to incorporate comparative and pan-genomic approaches as well. Here, we present a highly contiguous, chromosome-scale and haplotype-resolved genome assembly for the southern live oak, Quercus virginiana, the first reference genome for section Virentes, as part of the American Campus Tree Genomes program. Originating from a clone of Auburn University's historic “Toomer's Oak,” this assembly contributes to the pool of genomic resources for investigating recombination, haplotype variation, and structural genomic changes influencing hybridization potential in this clade and across Quercus. It also provides insights into the architecture of the putative centromeric regions within the genus. Alongside other oak references, the Q. virginiana genome will support research into the evolution and adaptation of the Quercus genus.

Quercus virginiana↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Multiple Wheat Genomes Reveal Novel Gli-2 Sublocus Location and Variation of Celiac Disease Epitopes in Duplicated α-Gliadin Genes

The seed protein α-gliadin is a major component of wheat flour and causes gluten-related diseases. However, due to the complexity of this multigene family with a genome structure composed of dozens of copies derived from tandem and genome duplications, little was known about the variation between accessions, and thus little effort has been made to explicitly target α-gliadin for bread wheat breeding. Here, we analyzed genomic variation in α-gliadins across 11 recently published chromosome-scale assemblies of hexaploid wheat, with validation using long-read data. We unexpectedly found that the Gli-B2 locus is not a single contiguous locus but is composed of two subloci, suggesting the possibility of recombination between the two during breeding. We confirmed that the number of immunogenic epitopes among 11 accessions varied. The D subgenome of a European spelt line also contained epitopes, in agreement with its hybridization history. Evolutionary analysis identified amino acid sites under diversifying selection, suggesting their functional importance. The analysis opens the way for improved grain quality and safety through wheat breeding.

59 BASIC BIOLOGICAL SCIENCES↗

Crystal structure of N-terminally hexahistidine-tagged Onchocerca volvulus macrophage migration inhibitory factor-1

Onchocerca volvulus causes blindness, onchocerciasis, skin infections and devastating neurological diseases such as nodding syndrome. New treatments are needed because the currently used drug, ivermectin, is contraindicated in pregnant women and those co-infected with Loa loa . The Seattle Structural Genomics Center for Infectious Disease (SSGCID) produced, crystallized and determined the apo structure of N-terminally hexahistidine-tagged O. volvulus macrophage migration inhibitory factor-1 (His- Ov MIF-1). Ov MIF-1 is a possible drug target. His- Ov MIF-1 has a unique jellyfish-like structure with a prototypical macrophage migration inhibitory factor (MIF) trimer as the `head' and a unique C-terminal `tail'. Deleting the N-terminal tag reveals an Ov MIF-1 structure with a larger cavity than that observed in human MIF that can be targeted for drug repurposing and discovery. Removal of the tag will be necessary to determine the actual biological oligomer of Ov MIF-1 because size-exclusion chomatographic analysis of His- Ov MIF-1 suggests a monomer, while PISA analysis suggests a hexamer stabilized by the unique C-terminal tails.

Kimble, Amber D. (ORCID:0000000167851596)↗

Genetic and behavioral adaptation of Candida parapsilosis to the microbiome of hospitalized infants revealed by in situ genomics, transcriptomics, and proteomics

Background Candida parapsilosis is a common cause of invasive candidiasis, especially in newborn infants, and infections have been increasing over the past two decades. C. parapsilosis has been primarily studied in pure culture, leaving gaps in understanding of its function in a microbiome context. Results. Here, we compare five unique C. parapsilosis genomes assembled from premature infant fecal samples, three of which are newly reconstructed, and analyze their genome structure, population diversity, and in situ activity relative to reference strains in pure culture. All five genomes contain hotspots of single nucleotide variants, some of which are shared by strains from multiple hospitals. A subset of environmental and hospital-derived genomes share variants within these hotspots suggesting derivation of that region from a common ancestor. Four of the newly reconstructed C. parapsilosis genomes have 4 to 16 copies of the gene RTA3, which encodes a lipid translocase and is implicated in antifungal resistance, potentially indicating adaptation to hospital antifungal use. Time course metatranscriptomics and metaproteomics on fecal samples from a premature infant with a C. parapsilosis blood infection revealed highly variable in situ expression patterns that are distinct from those of similar strains in pure cultures. For example, biofilm formation genes were relatively less expressed in situ, whereas genes linked to oxygen utilization were more highly expressed, indicative of growth in a relatively aerobic environment. In gut microbiome samples, C. parapsilosis co-existed with Enterococcus faecalis that shifted in relative abundance over time, accompanied by changes in bacterial and fungal gene expression and proteome composition. Conclusions The results reveal potentially medically relevant differences in Candida function in gut vs. laboratory environments, and constrain evolutionary processes that could contribute to hospital strain persistence and transfer into premature infant microbiomes.

59 BASIC BIOLOGICAL SCIENCES↗