Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

An RNA ligase partner for the prokaryotic protein-only RNase P: insights into the functional diversity of RNase P from genome mining

RNase P can use either an RNA- or a protein-based active site to catalyze 5'-maturation of transfer RNAs (tRNAs). This distinctive attribute in the biocatalytic repertoire raises questions about the underlying evolutionary driving forces, especially if each variant somehow affords a selective advantage under certain conditions. Upon mining all publicly available prokaryotic genomes and examining gene co-occurrence, we discovered that an RNA ligase with circularization activity was significantly overrepresented in genomes that contain the protein form of RNase P. This unexpected linkage inspires testable ideas to understand the bases for scenarios that might favor RNase P variants of different architectures/make-up.

HARP↗

Functional characterization of glycosyltransferases in duckweed to enable predictive biology

Glycosyltransferases (GTs) catalyze the formation of glycosidic linkages to produce almost all complex carbohydrates. This project used a multi-disciplinary, high-throughput (HTP) biochemical and computational biology approach focused on duckweed as a model energy crop, to study carbohydrate metabolic processes. To achieve this, developed and carried out out high-throughput (HTP) functional characterization of plant glycosyltransferases (GTs) role of enzymatic microenvironments be assessed through a combined proteomic and computational biology approach, and the combined data was used to populate deep-learning frameworks to predict plant GT function. Functional validation achieved through this research is being used to assign gene function and study plant processes at the systems level to efficiently link the genome sequence with gene function. Together, the combined approaches used within this study provide a foundation for how computational prediction, in combination with high-throughput functional validation, can be used to study plant processes at the systems level and translate knowledge gained to efficiently link genome sequence with gene function in a species agnostic manner.

09 BIOMASS FUELS↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

Identifying microbial functional guilds performing cryptic organotrophic and lithotrophic redox cycles in anaerobic granular biofilms

Granular biofilms used in anaerobic digester systems contain diverse microbial populations that interact to hydrolyze organic matter and produce methane within controlled environments. Prior research investigated the feasibility of utilizing granular biofilms obtained from an anaerobic digester to remove nitrate without the addition of exogenous electron donors. These granules possessed a unique structure of alternating light and dark iron sulfide and pyrite rich layers that potentially served as both an electron source and sink, linking carbon, nitrogen, sulfur, and iron cycles. To characterize the functional roles of diverse microbial populations enriched within these layered biofilms, we analyzed metagenomes obtained from three different granules. Comparisons between the functional gene content of forty metagenome assembled genomes (MAGs) identified phylogenetically cohesive functional guilds. Each of these functional MAG clusters was assigned to specific steps in anaerobic digestion (hydrolysis, acidogenesis, acetogenesis, and methanogenesis) and anaerobic respiration (denitrification and sulfate reduction). Comparisons with metagenomes derived from a variety of natural and engineered ecosystems confirmed that the enriched denitrifying bacteria were similar to populations typically found in wetlands and biological nitrogen removal systems. Analysis of read alignments to individual genes within the forty MAGs identified conserved genomic features that were representative of the functions that distinguished functional guilds. Overall, this research illustrates the utility of functional based classification of microorganisms for characterizing ecosystem functions and highlights the potential application of engineered ecosystems to serve as experimental models for complex natural ecosystems.

Ecosystem engineering↗

CRISPR-GRIT: Guide RNAs with Integrated Repair Templates Enable Precise Multiplexed Genome Editing in the Diploid Fungal Pathogen Candida albicans

Candida albicans, an opportunistic fungal pathogen, causes severe infections in immunocompromised individuals. Limited classes and overuse of current antifungals have led to the rapid emergence of antifungal resistance. Thus, there is an urgent need to understand fungal pathogen genetics to develop new antifungal strategies. Genetic manipulation of C. albicans is encumbered by its diploid chromosomes requiring editing both alleles to elucidate gene function. Although the recent development of CRISPR-Cas systems has facilitated genome editing in C. albicans, large-scale and multiplexed functional genomic studies are still hindered by the necessity of cotransforming repair templates for homozygous knockouts. Here, we present CRISPR-GRIT (Guide RNAs with Integrated Repair Templates), a repair template-integrated guide RNA design for expedited gene knockouts and multiplexed gene editing in C. albicans. Here, we envision that this method can be used for high-throughput library screens and identification of synthetic lethal pairs in both C. albicans and other diploid organisms with strong homologous recombination machinery.

60 APPLIED LIFE SCIENCES↗

Telomere-to-telomere assemblies of chromosome 10 reveal complex adaptive variation of 3-ketoacyl-CoA-synthases in Populus trichocarpa likely driven by Helitrons

The model woody plant Populus trichocarpa displays an atypical alkene-diverse wax cuticle likely driven by copy number variation (CNV) of 3-ketoacyl-CoA synthases ( KCS ), which has been difficult to confirm with short-read assemblies. Long-read sequencing enables the development of telomere-to-telomere resources to detect cryptic variation, including CNVs, which are currently missed. Integrating this information can improve genomic prediction for breeding and provide insights into the evolutionary basis of important traits. Our analysis of 78 long-read haplotypes from chromosome 10 identified more than twice as many KCS genes as previously reported, and numerous intragenic non-synonymous substitutions. Random Forest predictive models highlighted the importance of Potri.010G079500 in producing very long chain alkenes; however, its absence did not predict previously reported alkene-deficient phenotypes. Instead, alkene levels are best predicted by the combinations of KCS copies. Additionally, amino acid substitutions clustered around ligand and donor binding pockets, suggesting they contribute to differing wax cuticle composition. Finally, each KCS gene and copy was linked to a Helitron transposon. A phylogenetic analysis suggests Helitrons are the evolutionary mechanism for generating KCS tandem arrays. Long-read generated telomere-to-telomere assemblies of P. trichocarpa chromosome 10 revealed large-effect loci critical to genetic studies that are unattainable from short-reads. This new resource produced novel insights into genome structure and function, and a novel mechanism for generating tandem gene duplication. Our results highlight that, given current challenges in annotation and assembly, detailed and focused long-read sequences are key to interpreting complex genomic regions that contain tandem copy number variants.

09 BIOMASS FUELS↗

Structural basis of differential gene expression at eQTLs loci from high-resolution ensemble models of 3D single-cell chromatin conformations

Abstract Motivation Techniques such as high-throughput chromosome conformation capture (Hi-C) have provided a wealth of information on nucleus organization and genome important for understanding gene expression regulation. Genome-Wide Association Studies have identified numerous loci associated with complex traits. Expression quantitative trait loci (eQTL) studies have further linked the genetic variants to alteration in expression levels of associated target genes across individuals. However, the functional roles of many eQTLs in noncoding regions remain unclear. Current joint analyses of Hi-C and eQTLs data lack advanced computational tools, limiting what can be learned from these data. Results We developed a computational method for simultaneous analysis of Hi-C and eQTL data, capable of identifying a small set of nonrandom interactions from all Hi-C interactions. Using these nonrandom interactions, we reconstructed large ensembles (×105) of high-resolution single-cell 3D chromatin conformations with thorough sampling, accurately replicating Hi-C measurements. Our results revealed many-body interactions in chromatin conformation at the single-cell level within eQTL loci, providing a detailed view of how 3D chromatin structures form the physical foundation for gene regulation, including how genetic variants of eQTLs affect the expression of associated eGenes. Furthermore, our method can deconvolve chromatin heterogeneity and investigate the spatial associations of eQTLs and eGenes at subpopulation level, revealing their regulatory impacts on gene expression. Together, ensemble modeling of thoroughly sampled single-cell chromatin conformations combined with eQTL data, helps decipher how 3D chromatin structures provide the physical basis for gene regulation, expression control, and aid in understanding the overall structure-function relationships of genome organization. Availability and implementation It is available at https://github.com/uic-liang-lab/3DChromFolding-eQTL-Loci.

Du, Lin (ORCID:0009000289869812)↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

Multi-omics analysis reveals the dynamic interplay between Vero host chromatin structure and function during vaccinia virus infection

The genome folds into complex configurations and structures thought to profoundly impact its function. The intricacies of this dynamic structure-function relationship are not well understood particularly in the context of viral infection. To unravel this interplay, here we provide a comprehensive investigation of simultaneous host chromatin structural (via Hi-C and ATAC-seq) and functional changes (via RNA-seq) in response to vaccinia virus infection. Over time, infection significantly impacts global and local chromatin structure by increasing long-range intra-chromosomal interactions and B compartmentalization and by decreasing chromatin accessibility and inter-chromosomal interactions. Local accessibility changes are independent of broad-scale chromatin compartment exchange (~12% of the genome), underscoring potential independent mechanisms for global and local chromatin reorganization. While infection structurally condenses the host genome, there is nearly equal bidirectional differential gene expression. Despite global weakening of intra-TAD interactions, functional changes including downregulated immunity genes are associated with alterations in local accessibility and loop domain restructuring. Therefore, chromatin accessibility and local structure profiling provide impactful predictions for host responses and may improve development of efficacious anti-viral counter measures including the optimization of vaccine design.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of key steps in the evolution of anaerobic methanotrophy in Candidatus Methanovorans (ANME-3) archaea

Despite their large environmental impact and multiple independent emergences, the processes leading to the evolution of anaerobic methanotrophic archaea (ANME) remain unclear. This work uses comparative metagenomics of a recently evolved but understudied ANME group, “Candidatus Methanovorans” (ANME-3), to identify evolutionary processes and innovations at work in ANME, which may be obscured in earlier evolved lineages. We identified horizontal transfer of hdrA homologs and convergent evolution in carbon and energy metabolic genes as potential early steps in Methanovorans evolution. We also identified the erosion of genes required for methylotrophic methanogenesis along with horizontal acquisition of multiheme cytochromes and other loci uniquely associated with ANME. The assembly and comparative analysis of multiple Methanovorans genomes offers important functional context for understanding the niche-defining metabolic differences between methane-oxidizing ANME and their methanogen relatives. Furthermore, this work illustrates the multiple evolutionary modes at play in the transition to a globally important metabolic niche.

59 BASIC BIOLOGICAL SCIENCES↗

Phylogenomic discovery and engineering of nitrogen fixation into the bioenergy woody crop poplar

Biological nitrogen fixation (BNF) is a key process enabling plants in specific lineages to convert atmospheric dinitrogen (N₂) into bioavailable ammonia through symbioses with diazotrophic microbes. Expanding this capability beyond native nitrogen-fixing clades into non-nodulating crops would reduce synthetic fertilizer use, lowering energy inputs and environmental impacts in agriculture. Supported by DOE Funding Award DE-SC0018247, the NitFix project advanced foundational knowledge required to engineer root-nodule symbioses in new host species. The team generated the most comprehensive phylogenomic analysis to date of all known nodulating lineages, resolving the evolutionary history of nitrogen-fixing symbiosis and identifying core gene suites retained across nodulating taxa. Through multimodal genomics, transcriptomics, and functional analyses in Medicago truncatula and related species, the project mapped regulatory networks underlying nodule organogenesis, bacterial infection, and nitrogen-fixation efficiency. Key discoveries include the identification of conserved signaling modules for rhizobial recognition, transcription factors controlling nodule differentiation, and metabolic pathways integrating fixed nitrogen into plant growth. The project also developed enabling tools—including optimized transformation pipelines, gene-editing workflows, and imaging-based phenotyping—to accelerate engineering efforts in emerging models. Together, these results refine the mechanistic framework of symbiotic nitrogen fixation and highlight transferable components essential for rewiring these traits into non-nodulating crops.

59 BASIC BIOLOGICAL SCIENCES↗

Conjugation-based genome engineering enables rapid prototyping and bioproduction in non-model bacteria

Abstract Non-model bacteria offer unique metabolic capabilities for sustainable bioproduction, yet their limited genetic accessibility hinders systematic strain development. Here we present conjugation-based serine recombinase-assisted genome engineering (cSAGE), a broad-host-range platform that enables predictable, iterative genomic integration in transformation-resistant bacteria. cSAGE combines conjugative DNA delivery, standardized low-copy vectors, orthogonal recombinases, and modular genetic parts to support rapid pathway assembly and cross-host benchmarking. Using purple nonsulfur bacteria as a testbed, we integrate promoter engineering, multi-payload genome modification, and genome-scale metabolic modeling to empirically evaluate host-dependent pathway performance. Applying this workflow, we identify strain-specific differences in photosynthetic conversion of lignin-derived p -coumarate to the thermoplastic precursor p -vinylphenol. By enabling genome engineering and functional comparison across diverse bacteria using a single plasmid system, cSAGE provides a general framework for non-model strain prototyping and biotransformation discovery.

Guzman, Michael S. [Department of Chemical Enginee↗

Cross-family and phage-specific gene requirements for Klebsiella infection revealed by scalable RB-TnSeq genetic screens.

Bacteriophages are being cataloged at an accelerating pace and are recognized as key players in nutrient and energy cycling across ecosystems. Yet the bacterial genetic determinants that govern phage-host specificity and infection success remain poorly understood, particularly in clinically and ecologically important genera such as Klebsiella where prior receptor characterization has been almost entirely limited to capsulated strains. Here we used a randomly barcoded, genome-wide, loss-of-function transposon mutant library (RB-TnSeq) of Klebsiella sp. M5al, a naturally acapsular, nitrogen-fixing rhizobacterium, to generate the first systematic, cross-family map of phage receptor gene dependencies in Klebsiella. Challenging the library against 25 double-stranded DNA phages spanning five families in 213 parallel assays, we identified 42 bacterial genes associated with phage infection, of which 15 had no prior association with phage infection in any bacterial system. Disruption of surface receptor biosynthesis genes conferred cross-resistance across multiple phage families, while intracellular gene disruptions had predominantly phage-specific effects. Clonal validation of eight genes confirmed LPS outer core biosynthesis genes as primary receptor determinants alongside additional host factors spanning outer membrane transport, cofactor biosynthesis, and two-component signaling. Comparative analysis across all 25 phages revealed that phage genus rather than family is the stronger predictor of host gene dependency profiles, a finding with direct implications for the functional annotation of uncharacterized phage isolates and rational phage cocktail design. Together, these findings provide a community resource for linking phage genomic diversity to functional host interaction space in this ecologically and clinically important genus.

Gittrich, Marissa R↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Eco-evolutionary strategies for relieving carbon limitation under salt stress differ across microbial clades

With the continuous expansion of saline soils under climate change, understanding the eco-evolutionary tradeoff between the microbial mitigation of carbon limitation and the maintenance of functional traits in saline soils represents a significant knowledge gap in predicting future soil health and ecological function. Through shotgun metagenomic sequencing of coastal soils along a salinity gradient, we show contrasting eco-evolutionary directions of soil bacteria and archaea that manifest in changes to genome size and the functional potential of the soil microbiome. In salt environments with high carbon requirements, bacteria exhibit reduced genome sizes associated with a depletion of metabolic genes, while archaea display larger genomes and enrichment of salt-resistance, metabolic, and carbon-acquisition genes. This suggests that bacteria conserve energy through genome streamlining when facing salt stress, while archaea invest in carbon-acquisition pathways to broaden their resource usage. These findings suggest divergent directions in eco-evolutionary adaptations to soil saline stress amongst microbial clades and serve as a foundation for understanding the response of soil microbiomes to escalating climate change.

54 ENVIRONMENTAL SCIENCES↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete↗