Search NASASearch

SEARCH · Search NASA

Results for “functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota

Multi-omics analysis reveals the dynamic interplay between Vero host chromatin structure and function during vaccinia virus infection

The genome folds into complex configurations and structures thought to profoundly impact its function. The intricacies of this dynamic structure-function relationship are not well understood particularly in the context of viral infection. To unravel this interplay, here we provide a comprehensive investigation of simultaneous host chromatin structural (via Hi-C and ATAC-seq) and functional changes (via RNA-seq) in response to vaccinia virus infection. Over time, infection significantly impacts global and local chromatin structure by increasing long-range intra-chromosomal interactions and B compartmentalization and by decreasing chromatin accessibility and inter-chromosomal interactions. Local accessibility changes are independent of broad-scale chromatin compartment exchange (~12% of the genome), underscoring potential independent mechanisms for global and local chromatin reorganization. While infection structurally condenses the host genome, there is nearly equal bidirectional differential gene expression. Despite global weakening of intra-TAD interactions, functional changes including downregulated immunity genes are associated with alterations in local accessibility and loop domain restructuring. Therefore, chromatin accessibility and local structure profiling provide impactful predictions for host responses and may improve development of efficacious anti-viral counter measures including the optimization of vaccine design.

59 BASIC BIOLOGICAL SCIENCES

Identification of key steps in the evolution of anaerobic methanotrophy in Candidatus Methanovorans (ANME-3) archaea

Despite their large environmental impact and multiple independent emergences, the processes leading to the evolution of anaerobic methanotrophic archaea (ANME) remain unclear. This work uses comparative metagenomics of a recently evolved but understudied ANME group, “Candidatus Methanovorans” (ANME-3), to identify evolutionary processes and innovations at work in ANME, which may be obscured in earlier evolved lineages. We identified horizontal transfer of hdrA homologs and convergent evolution in carbon and energy metabolic genes as potential early steps in Methanovorans evolution. We also identified the erosion of genes required for methylotrophic methanogenesis along with horizontal acquisition of multiheme cytochromes and other loci uniquely associated with ANME. The assembly and comparative analysis of multiple Methanovorans genomes offers important functional context for understanding the niche-defining metabolic differences between methane-oxidizing ANME and their methanogen relatives. Furthermore, this work illustrates the multiple evolutionary modes at play in the transition to a globally important metabolic niche.

59 BASIC BIOLOGICAL SCIENCES

Phylogenomic discovery and engineering of nitrogen fixation into the bioenergy woody crop poplar

Biological nitrogen fixation (BNF) is a key process enabling plants in specific lineages to convert atmospheric dinitrogen (N₂) into bioavailable ammonia through symbioses with diazotrophic microbes. Expanding this capability beyond native nitrogen-fixing clades into non-nodulating crops would reduce synthetic fertilizer use, lowering energy inputs and environmental impacts in agriculture. Supported by DOE Funding Award DE-SC0018247, the NitFix project advanced foundational knowledge required to engineer root-nodule symbioses in new host species. The team generated the most comprehensive phylogenomic analysis to date of all known nodulating lineages, resolving the evolutionary history of nitrogen-fixing symbiosis and identifying core gene suites retained across nodulating taxa. Through multimodal genomics, transcriptomics, and functional analyses in Medicago truncatula and related species, the project mapped regulatory networks underlying nodule organogenesis, bacterial infection, and nitrogen-fixation efficiency. Key discoveries include the identification of conserved signaling modules for rhizobial recognition, transcription factors controlling nodule differentiation, and metabolic pathways integrating fixed nitrogen into plant growth. The project also developed enabling tools—including optimized transformation pipelines, gene-editing workflows, and imaging-based phenotyping—to accelerate engineering efforts in emerging models. Together, these results refine the mechanistic framework of symbiotic nitrogen fixation and highlight transferable components essential for rewiring these traits into non-nodulating crops.

59 BASIC BIOLOGICAL SCIENCES

Conjugation-based genome engineering enables rapid prototyping and bioproduction in non-model bacteria

Abstract Non-model bacteria offer unique metabolic capabilities for sustainable bioproduction, yet their limited genetic accessibility hinders systematic strain development. Here we present conjugation-based serine recombinase-assisted genome engineering (cSAGE), a broad-host-range platform that enables predictable, iterative genomic integration in transformation-resistant bacteria. cSAGE combines conjugative DNA delivery, standardized low-copy vectors, orthogonal recombinases, and modular genetic parts to support rapid pathway assembly and cross-host benchmarking. Using purple nonsulfur bacteria as a testbed, we integrate promoter engineering, multi-payload genome modification, and genome-scale metabolic modeling to empirically evaluate host-dependent pathway performance. Applying this workflow, we identify strain-specific differences in photosynthetic conversion of lignin-derived p -coumarate to the thermoplastic precursor p -vinylphenol. By enabling genome engineering and functional comparison across diverse bacteria using a single plasmid system, cSAGE provides a general framework for non-model strain prototyping and biotransformation discovery.

Guzman, Michael S. [Department of Chemical Enginee

Resistance of virus to extinction on bottleneck passages: study of a decaying and fluctuating pattern of fitness loss

RNA viruses display high mutation rates and their populations replicate as dynamic and complex mutant distributions, termed viral quasispecies. Repeated genetic bottlenecks, which experimentally are carried out through serial plaque-to-plaque transfers of the virus, lead to fitness decrease (measured here as diminished capacity to produce infectious progeny). Here we report an analysis of fitness evolution of several low fitness foot-and-mouth disease virus clones subjected to 50 plaque-to-plaque transfers. Unexpectedly, fitness decrease, rather than being continuous and monotonic, displayed a fluctuating pattern, which was influenced by both the virus and the state of the host cell as shown by effects of recent cell passage history. The amplitude of the fluctuations increased as fitness decreased, resulting in a remarkable resistance of virus to extinction. Whereas the frequency distribution of fitness in control (independent) experiments follows a log-normal distribution, the probability of fitness values in the evolving bottlenecked populations fitted a Weibull distribution. We suggest that multiple functions of viral genomic RNA and its encoded proteins, subjected to high mutational pressure, interact with cellular components to produce this nontrivial, fluctuating pattern.

Serial Passage

Research from the NASA Twins Study and Omics in Support of Mars Missions

The NASA Twins Study, NASA's first foray into integrated omic studies in humans, illustrates how an integrated omics approach can be brought to bear on the challenges to human health and performance on a Mars mission. The NASA Twins Study involves US Astronaut Scott Kelly and his identical twin brother, Mark Kelly, a retired US Astronaut. No other opportunity to study a twin pair for a prolonged period with one subject in space and one on the ground is available for the foreseeable future. A team of 10 principal investigators are conducting the Twins Study, examining a very broad range of biological functions including the genome, epigenome, transcriptome, proteome, metabolome, gut microbiome, immunological response to vaccinations, indicators of atherosclerosis, physiological fluid shifts, and cognition. A novel aspect of the study is the integrated study of molecular, physiological, cognitive, and microbiological properties. Major sample and data collection from both subjects for this study began approximately six months before Scott Kelly's one year mission on the ISS, continue while Scott Kelly is in flight and will conclude approximately six months after his return to Earth. Mark Kelly will remain on Earth during this study, in a lifestyle unconstrained by this study, thereby providing a measure of normal variation in the properties being studied. An overview of initial results and the future plans will be described as well as the technological and ethical issues raised for spaceflight studies involving omics.

Kundrot, C.

Novel Approach to Quantification of Telomere Length with Direct Nanopore Sequencing and PCR Amplification

The ends of human chromosomes contain telomeres, or tandem arrays of repeating DNA sequences capped by multiple associated proteins that protect chromosomal ends from degradation. Telomeres function to preserve genomic stability by preventing natural chromosomal ends from being recognized as broken DNA double-strand breaks and triggering inappropriate DNA damage responses. Mounting evidence shows telomere length is an inherited trait that decreases with cellular division and normal aging. In addition, telomere length also appears to be influenced by other factors such as cellular oxidative stress, radiation and mechanical unloading of tissues as in microgravity. To measure these potential effects of the space environment on telomere lengths and cellular aging and regenerative potential we developed a novel telomere measurement approach based on nanopore sequencing of PCR amplified bar-coded chromosome termini. Specifically, telomeres can be directly enriched using barcode sequences ligated to the end of a free end- repaired telomere using the WetLab-2 facility SmartCycler on ISS. Prior to the ligation and amplification protocol a proteinase K digestion of capping proteins followed by a single 95-degree C heat denaturation of the protease is included. After digestion and bar-code ligation, PCR amplification will initiate with the ligated barcoded sequence, suppressing amplification of intra-genomic fragments and resulting in long read barcoded telomere amplicons including the nanopore motor protein sequences. Purified PCR amplicons are then used for nanopore sequencing library generation by simple addition of motor proteins and sequencing library is loaded into the MinION nanopore DNA-sequencer. Amplicon sequence reads from the nanopore device can be base-called quickly on ISS due to barcoding ligation and subsequent PCR amplification enhancing the telomere sequence resolution. If successfully implemented on ISS this technique will provide a novel means of measuring regenerative ability of somatic stem cells in astronauts, and of determining whether spaceflight in microgravity alters their telomere lengths and causes premature cellular aging.

Ma, Kristin R.

Cross-family and phage-specific gene requirements for Klebsiella infection revealed by scalable RB-TnSeq genetic screens.

Bacteriophages are being cataloged at an accelerating pace and are recognized as key players in nutrient and energy cycling across ecosystems. Yet the bacterial genetic determinants that govern phage-host specificity and infection success remain poorly understood, particularly in clinically and ecologically important genera such as Klebsiella where prior receptor characterization has been almost entirely limited to capsulated strains. Here we used a randomly barcoded, genome-wide, loss-of-function transposon mutant library (RB-TnSeq) of Klebsiella sp. M5al, a naturally acapsular, nitrogen-fixing rhizobacterium, to generate the first systematic, cross-family map of phage receptor gene dependencies in Klebsiella. Challenging the library against 25 double-stranded DNA phages spanning five families in 213 parallel assays, we identified 42 bacterial genes associated with phage infection, of which 15 had no prior association with phage infection in any bacterial system. Disruption of surface receptor biosynthesis genes conferred cross-resistance across multiple phage families, while intracellular gene disruptions had predominantly phage-specific effects. Clonal validation of eight genes confirmed LPS outer core biosynthesis genes as primary receptor determinants alongside additional host factors spanning outer membrane transport, cofactor biosynthesis, and two-component signaling. Comparative analysis across all 25 phages revealed that phage genus rather than family is the stronger predictor of host gene dependency profiles, a finding with direct implications for the functional annotation of uncharacterized phage isolates and rational phage cocktail design. Together, these findings provide a community resource for linking phage genomic diversity to functional host interaction space in this ecologically and clinically important genus.

Gittrich, Marissa R

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant

Eco-evolutionary strategies for relieving carbon limitation under salt stress differ across microbial clades

With the continuous expansion of saline soils under climate change, understanding the eco-evolutionary tradeoff between the microbial mitigation of carbon limitation and the maintenance of functional traits in saline soils represents a significant knowledge gap in predicting future soil health and ecological function. Through shotgun metagenomic sequencing of coastal soils along a salinity gradient, we show contrasting eco-evolutionary directions of soil bacteria and archaea that manifest in changes to genome size and the functional potential of the soil microbiome. In salt environments with high carbon requirements, bacteria exhibit reduced genome sizes associated with a depletion of metabolic genes, while archaea display larger genomes and enrichment of salt-resistance, metabolic, and carbon-acquisition genes. This suggests that bacteria conserve energy through genome streamlining when facing salt stress, while archaea invest in carbon-acquisition pathways to broaden their resource usage. These findings suggest divergent directions in eco-evolutionary adaptations to soil saline stress amongst microbial clades and serve as a foundation for understanding the response of soil microbiomes to escalating climate change.

54 ENVIRONMENTAL SCIENCES

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete

Development of male-sterile lines of Setaria viridis to accelerate C 4 model plant genetics

Setaria viridis is a diploid C 4 grass in the Poaceae family, notable for its rapid life cycle of 6–8 weeks from sowing to seed—much shorter than the 4–5 months required by crops such as Zea mays and Sorghum bicolor . This fast growth makes S. viridis a valuable model for C 4 crop research. Genetic crosses are essential for studying gene function, but manual crossing is labor-intensive and time-consuming. Here, to address this, we developed a male-sterile line by targeting the S. viridis ortholog of Setaria italica NO POLLEN 1 ( SiNP1 ), which encodes a glucose–methanol–choline oxidoreductase required for pollen exine formation. Using Cas9 and TREX2 -mediated genome editing, we generated SiNP1 knockouts in both the S. viridis ME034V and A10.1 backgrounds that were fully male-sterile. Backcrossing T 0 male-sterile plants to ME034V wild-type followed by selfing yielded a stable BC 1 F 2 line homozygous for a 59 bp deletion in the S. viridis NO POLLEN 1 gene, easily genotyped by PCR and maintained by heterozygous siblings. Using this line, we developed a simple and efficient crossing protocol that eliminates the need for emasculation. This method enables a single person to perform up to 100 crosses per day—compared to 15 using traditional methods—and yields 20–32 F 1 hybrid seeds per panicle with 100% genetic purity. We also quantified pollen flow and outcrossing frequencies under greenhouse conditions to develop optimal bagging strategies and prevent unintended pollination. This resource accelerates genetic research in S. viridis , enhancing its utility as a premier C 4 model for mapping and functional genomics.

C4 research

PERCEPTIVE: an R shiny $\underline{p}$ipelin$\underline{e}$ for the p$\underline{r}$edi$\underline{c}$tion of $\underline{ep}$igenetic modula$\underline{t}$ors $\underline{i}$n no$\underline{v}$el sp$\underline{e}$cies

Epigenetic processes are central to regulating gene expression, genome stability, and metabolic function across the tree of life; yet, their roles remain underexplored in microalgae, especially as new species continue to be identified and characterized. This is likely due to the cumbersome nature and species-dependent attributes of epigenetic wet-lab methodologies, which preclude the rapid identification of epigenetic modifications and modulators. However, there is high conservation of epigenetic processes from budding yeast to humans; in many cases, one may infer how behavior and function are epigenetically regulated in novel species by identifying epigenetic modulators, or the proteins responsible for conferring epigenetic modifications. Here, to this end, we have developed a graphical software package, titled PERCEPTIVE (pipeline for the prediction of epigenetic modulators in novel species). This platform solely uses the genomic sequence of an algal species, and preexisting information from other model organisms, to predict the epigenetic modulators and associated modifications in algae. Predictions are presented to the user in a graphical interface, which provides literature-based interpretation of results, enabling users to quickly understand potential epigenetic processes in their algal species of interest and plan follow-up experiments. To test PERCEPTIVE, we predicted epigenetic modulators in several feedstock candidate algae species. To validate these predictions, wet-lab studies were performed, including mass spectrometry; these results underscore the high accuracy of PERCEPTIVE predictions. Overall, PERCEPTIVE represents a powerful in silico tool for the research and manipulation of algal species, which does not require a priori knowledge of epigenetics and is accessible to a broad set of investigators.

59 BASIC BIOLOGICAL SCIENCES

A genomic view of Earth’s biomes

Microorganisms are essential to all life on Earth through critical roles in key biological processes and diverse interactions with other organisms that shape ecosystems, drive biogeochemical cycles and influence both human health and environmental health. High-throughput sequencing from environmental samples has revolutionized the understanding of microbial diversity and functions. With vast amounts of genomes now available across Earth’s biomes, these data provide a blueprint of microbial life that can be harnessed for a more holistic understanding of microbiome structure and function across the various ecosystems on Earth. Here we review the application of genome-centric approaches, including recent advances in single-cell sequencing and functional profiling, to survey microbial and viral diversity. Furthermore, we highlight some of the most impactful evolutionary and functional discoveries, explore the spatial diversity and temporal dynamics of microorganisms across diverse environments, and discuss genome-enabled insights into host-associated microorganisms.

Ecology

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)