Search NASA⌕ Search

SEARCH · Search NASA

Results for “genome analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CYP76BK1 orthologs catalyze furan and lactone ring formation in clerodane diterpenoids across the mint family

The Lamiaceae (mint family) is the largest known source of furanoclerodanes, a subset of clerodane diterpenoids with broad bioactivities including insect antifeedant properties. The Ajugoideae subfamily, in particular, accumulates significant numbers of structurally related furanoclerodanes. The biosynthetic capacity for formation of these diterpenoids is retained across most Lamiaceae subfamilies, including the early-diverging Callicarpoideae which forms a sister clade to the rest of Lamiaceae. VacCYP76BK1, a cytochrome P450 monooxygenase from Vitex agnus-castus, was previously found to catalyze the formation of the proposed precursor to furan and lactone-containing labdane diterpenoids. Through transcriptome-guided pathway exploration, we identified orthologs of VacCYP76BK1 in Ajuga reptans and Callicarpa americana. Functional characterization demonstrated that both could catalyze the oxidative cyclization of clerodane backbones to yield a furan ring. Subsequent investigation revealed a total of 10 CYP76BK1 orthologs across six Lamiaceae subfamilies. Through analysis of available chromosome-scale genomes, we identified four CYP76BK1 members as syntelogs within a conserved syntenic block across divergent subfamilies. This suggests an evolutionary lineage that predates the speciation of the Lamiaceae. Functional characterization of the CYP76BK1 orthologs affirmed conservation of function, as all catalyzed furan ring formation. Additionally, some orthologs yielded two novel lactone ring moieties. The presence of the CYP76BK1 orthologs across Lamiaceae subfamilies closely overlaps with the distribution of reported furanoclerodanes. Together, the activities and distribution of the CYP76BK1 orthologs identified here support their central role in furanoclerodane biosynthesis within the Lamiaceae family. Our findings lay the groundwork for biotechnological applications to harness the economic potential of this promising class of compounds.

59 BASIC BIOLOGICAL SCIENCES↗

A bloom of a single bacterium shapes the microbiome during outdoor diatom cultivation collapse

Algae-dominated ecosystems are fundamentally influenced by their microbiome. We lack information on the identity and function of bacteria that specialize in consuming algal-derived dissolved organic matter in high algal density ecosystems such as outdoor algal ponds used for biofuel production. Here, we describe the metagenomic and metaproteomic signatures of a single bacterial strain that bloomed during a population-wide crash of the diatom, Phaeodactylum tricornutum, grown in outdoor ponds. 16S rRNA gene data indicated that a single Kordia sp. strain (family Flavobacteriaceae) contributed up to 93% of the bacterial community during P. tricornutum demise. Kordia sp. expressed proteins linked to microbial antagonism and biopolymer breakdown, which likely contributed to its dominance over other microbial taxa during diatom demise. Analysis of accompanying downstream microbiota (primarily of the Rhodobacteraceae family) provided evidence that cross-feeding may be a pathway supporting microbial diversity during diatom demise. In situ and laboratory data with a different strain suggested that Kordia was a primary degrader of biopolymers during algal demise, and co-occurring Rhodobacteraceae exploited degradation molecules for carbon. An analysis of 30 Rhodobacteraceae metagenome assembled genomes suggested that algal pond Rhodobacteraceae commonly harbored pathways to use diverse carbon and energy sources, including carbon monoxide, which may have contributed to the prevalence of this taxonomic group within the ponds. These observations further constrain the roles of functionally distinct heterotrophic bacteria in algal microbiomes, demonstrating how a single dominant bacterium, specialized in processing senescing or dead algal biomass, shapes the microbial community of outdoor algal biofuel ponds.

Kordia↗

MONet/1000 Soils metagenome pathway modelling narrative w/ auto batch import

This narrative performs metabolic modeling and flux balance analysis (FBA) using metagenome-assembled genomes (MAGs) from the 1000 Soils samples, as described by Song et al. (2026, accepted). The set of MAGs (in FASTA format) is converted into a set of assembly objects compatible with functional annotation via RASTtk, yielding a set of genome objects that undergo metabolic modeling via OMEGGA. These genome objects are then used to conduct FBA, generating tables of metabolite uptake rates across the MAGs under investigation.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular Basis of the Increase in Invertase Activity Elicited by Gravistimulation of Oat-Shoot Pulvini

An asymmetric (top vs. bottom) increase in invertase activity is elicited by gravistimulation in oatshoot pulvini starting within 3h after treatment. In order to analyze the regulation of invertase gene expression in this system, we examined the effect of gravistimulation on invertase mRNA induction. Total RNA and poly(A)(+)RNA, isolated from oat pulvini, and two oligonucleotide primers, corresponding to two conserved amino-acid sequences (NDPNG and WECPD) found in invertase from other species, were used for the Polymerase Chain Reaction (PCR). A partial-length cDNA (550 base pairs) was obtained and characterized. There was a 52 % deduced amino-acid sequence homology to that of carrot beta-fructosi- dase and a 48 % homology to that of tomato invertase. Northern blot analysis showed that there was an obvious transient accumulation of invertase mRNA elicited by gravistimulation of oat pulvini. The mRNA was rapidly induced to a maximum level at 1h following gravistimulation treatment and gradually decreased afterwards. The mRNA level in the bottom half of the oat pulvinus was significantly higher (five-fold) than that in the top half of the pulvinus tissue. The induction of invertase mRNA was consistent with the transient enhancement of invertase activity during the graviresponse of the pulvinus. These data indicate that the expression of the invertase gene(s) could be regulated by gravistimulation at the transcriptional and/or translational levels. Southern blot analysis showed that there were four genomic DNA fragments hybridized to the invertase cDNA. This suggests that an invertase gene family may exist in oat plants.

Wu, Liu-Lai↗

Murine Host-gut Microbiota Interactions are Modulated During Spaceflight

The rodent habitat on the International Space Station has provided critical insight into the impact of spaceflight on mammalian physiology. These effects include dysfunction of carbohydrate, steroid and lipid metabolism, and immune response, as well as induction of symptoms characteristic of liver disease, insulin resistance, osteopenia and myopathy, which are anticipated to intensify over long-duration spaceflight. Although these physiological responses can involve the microbiome, the host-microorganism interactions during spaceflight are still largely unknown. NASA GeneLab curates a wide range of space research data and the current work harnesses GeneLab multi’omic data from recent Rodent Research studies to explore changes to gut microbiota during spaceflight and their associations with host physiology when compared to ground controls. Using a hybrid analysis of DNA barcoding and whole genome shotgun data, an array of bacteria, fungi and nematodes could be identified at species level, and significant differences in relative abundances associated with spaceflight. Functional prediction based on differential abundance of species and metagenome gene inventories as well as metatranscriptomic gene expression at the host-gut microbiome interface implicate microbiota interactions could contribute to spaceflight pathology. Harnessing carefully curated publicly available data, such as from Genelab, to generate multi‘omic space science discoveries can help decipher the complex host-microbiome interactions that influence both health on Earth and the feasibility of long-duration spaceflight.

Microbiome↗

Discovery of FoTO1 and Taxol genes enables biosynthesis of baccatin III

Abstract Plants make complex and potent therapeutic molecules 1,2 , but sourcing these molecules from natural producers or through chemical synthesis is difficult, which limits their use in the clinic. A prominent example is the anti-cancer therapeutic paclitaxel (sold under the brand name Taxol), which is derived from yew trees (Taxusspecies) 3 . Identifying the full paclitaxel biosynthetic pathway would enable heterologous production of the drug, but this has yet to be achieved despite half a century of research 4 . WithinTaxus’ large, enzyme-rich genome 5 , we suspected that the paclitaxel pathway would be difficult to resolve using conventional RNA-sequencing and co-expression analyses. Here, to improve the resolution of transcriptional analysis for pathway identification, we developed a strategy we term multiplexed perturbation × single nuclei (mpXsn) to transcriptionally profile cell states spanning tissues, cell types, developmental stages and elicitation conditions. Our data show that paclitaxel biosynthetic genes segregate into distinct expression modules that suggest consecutive subpathways. These modules resolved seven new genes, allowing a de novo 17-gene biosynthesis and isolation of baccatin III, the industrial precursor to Taxol, inNicotiana benthamianaleaves, at levels comparable with the natural abundance inTaxusneedles. Notably, we found that a nuclear transport factor 2 (NTF2)-like protein, FoTO1, is crucial for promoting the formation of the desired product during the first oxidation, resolving a long-standing bottleneck in paclitaxel pathway reconstitution. Together with a new β-phenylalanine-CoA ligase, the eight genes discovered here enable the de novo biosynthesis of 3’-N-debenzoyl-2’-deoxypaclitaxel. More broadly, we establish a generalizable approach to efficiently scale the power of co-expression analysis to match the complexity of large, uncharacterized genomes, facilitating the discovery of high-value gene sets.

Science & Technology - Other Topics↗

Steps Toward Improved Integration, Search, and Analysis of Heterogeneous Data in the Astrobiology Habitable Environments Database

The Astrobiology Habitable Environments Database (AHED) is a new data system being developed as a long-term, open-access repository for astrobiology data. AHED is intended to store user-contributed results from NASA or externally-funded research in astrobiology, and to encourage sharing and synergy within the astrobiology community. However, the interdisciplinary nature of astrobiology presents some specific challenges to data management, integration, and analysis within AHED. In some disciplines (e.g., genomics), open databases thrive because the contributed products are fairly uniform and standardized (e.g., sequence data). In astrobiology, each investigation produces a unique set of data products; this makes it difficult to search across different datasets to find similar data, or to combine results from separate investigations. With AHED, we are taking steps to ensure there is adequate metadata - both at the dataset and record levels - to facilitate search, integration, and analysis. At the dataset level, we are developing a new metadata standard for describing astrobiology datasets, with detailed information about content, funding source, and scientific relevance, along with a set of topical keywords for characterizing datasets. At the record level, we are encouraging users to provide more structured content and finer-grained metadata. In many user-contributed science data repositories, few restrictions are placed on the uploaded data format, and minimal or no record-level metadata is required; thus users are unburdened when it comes to data preparation. The tradeoff is that deep integration and search across datasets is almost impossible without standardized structures and metadata. Although AHED users are free to upload minimally-described datasets, they will be encouraged to use database authoring tools (supplied by the underlying platform - Open Data Repository's Data Publisher) plus a set of customizable astrobiology-specific templates to help structure their data and provide standardized metadata. In reward for their extra effort, AHED will be able to deliver enhanced search, discovery, and analysis capabilities.

astrobiology↗

YeastWGD2025

Supplementary data for Discovery of additional ancient genome duplications in yeasts wgd_syn / - directory containing wgd syn output for all contiguous genomes [dataset] Tree - phylogeny [dataset]Duplications - duplication table from OrthoFinder output KOannotations - KEGG annotations used for enrichment analysis IPRannotations - InterPro annotations used for enrichment analysis DipodascalesOrthogroups - formatted orthogroup assignments for Dipodascales genes.fa and .gff3 files for each new genome assembly are also provided, those these are not required to replicate the analysis

Genomics↗

A Rhodopseudomonas strain with a substantially smaller genome retains the core metabolic versatility of its genus

ABSTRACT Rhodopseudomonas are a group of phototrophic microbes with a marked metabolic versatility and flexibility that underpins their potential use in the production of value-added products, bioremediation, and plant growth promotion. Members of this group have an average genome size of about 5.5 Mb, but two closely related strains have genome sizes of about 4.0 Mb. To identify the types of genes missing in a reduced genome strain, we compared strain DSM127 with other Rhodopseudomonas isolates at the genomic and phenotypic levels. We found that DSM127 can grow as well as other members of the Rhodopseudomonas genus and retains most of their metabolic versatility, but it has many fewer genes associated with high-affinity transport of nutrients, iron uptake, nitrogen metabolism, and biodegradation of aromatic compounds. This analysis indicates genes that can be deleted in genome reduction campaigns and suggests that DSM127 could be a favorable choice for biotechnology applications using Rhodopseudomonas or as a strain that can be engineered further to reside in a specialized natural environment. IMPORTANCE Rhodopseudomonas are a cohort of phototrophic bacteria with broad metabolic versatility. Members of this group are present in diverse soil and water environments, and some strains are found associated with plants and have plant growth-promoting activity. Motivated by the idea that it may be possible to design bacteria with reduced genomes that can survive well only in a specific environment or that may be more metabolically efficient, we compared Rhodopseudomonas strains with typical genome sizes of about 5.5 Mb to a strain with a reduced genome size of 4.0 Mb. From this, we concluded that metabolic versatility is part of the identity of the Rhodopseudomonas group, but high-affinity transport genes and genes of apparent redundant function can be dispensed with.

59 BASIC BIOLOGICAL SCIENCES↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

ldrd_virus_work

This is a Python code base that takes openly-available genetic information on known viruses and performs supervised machine learning and feature importance analysis on the relationship of the viral genomes to the competence to infect humans or bind to a specific host cell receptor.

Reddy, Tyler [LANL]↗

Quantitative analysis of bristle number in Drosophila mutants identifies genes involved in neural development

BACKGROUND: The identification of the function of all genes that contribute to specific biological processes and complex traits is one of the major challenges in the postgenomic era. One approach is to employ forward genetic screens in genetically tractable model organisms. In Drosophila melanogaster, P element-mediated insertional mutagenesis is a versatile tool for the dissection of molecular pathways, and there is an ongoing effort to tag every gene with a P element insertion. However, the vast majority of P element insertion lines are viable and fertile as homozygotes and do not exhibit obvious phenotypic defects, perhaps because of the tendency for P elements to insert 5' of transcription units. Quantitative genetic analysis of subtle effects of P element mutations that have been induced in an isogenic background may be a highly efficient method for functional genome annotation. RESULTS: Here, we have tested the efficacy of this strategy by assessing the extent to which screening for quantitative effects of P elements on sensory bristle number can identify genes affecting neural development. We find that such quantitative screens uncover an unusually large number of genes that are known to function in neural development, as well as genes with yet uncharacterized effects on neural development, and novel loci. CONCLUSIONS: Our findings establish the use of quantitative trait analysis for functional genome annotation through forward genetics. Similar analyses of quantitative effects of P element insertions will facilitate our understanding of the genes affecting many other complex traits in Drosophila.

Non-NASA Center↗

Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses

ABSTRACT Although long-read sequencing has enabled obtaining high-quality and complete genomes from metagenomes, many challenges still remain to completely decompose a metagenome into its constituent prokaryotic and viral genomes. This study focuses on decomposing an estuarine metagenome to obtain a more accurate estimate of microbial diversity. To achieve this, we developed a new bead-based DNA extraction method, a novel bin refinement method, and obtained 150 Gbp of Nanopore sequencing. We estimate that there are ~500 bacterial and archaeal species in our sample and obtained 68 high-quality bins (>90% complete, <5% contamination, ≤5 contigs, contig length of >100 kbp, and all ribosomal and tRNA genes). We also obtained many contigs of picoeukaryotes, environmental DNA of larger eukaryotes such as mammals, and complete mitochondrial and chloroplast genomes and detected ~40,000 viral populations. Our analysis indicates that there are only a few strains that comprise most of the species abundances. IMPORTANCE Ocean and estuarine microbiomes play critical roles in global element cycling and ecosystem function. Despite the importance of these microbial communities, many species still have not been cultured in the lab. Environmental sequencing is the primary way the function and population dynamics of these communities can be studied. Long-read sequencing provides an avenue to overcome limitations of short-read technologies to obtain complete microbial genomes but comes with its own technical challenges, such as needed sequencing depth and obtaining high-quality DNA. We present here new sampling and bioinformatics methods to attempt decomposing an estuarine microbiome into its constituent genomes. Our results suggest there are only a few strains that comprise most of the species abundances from viruses to picoeukaryotes, and to fully decompose a metagenome of this diversity requires 1 Tbp of long-read sequencing. We anticipate that as long-read sequencing technologies continue to improve, less sequencing will be needed.

Lui, Lauren M.↗

Multiscale Simulation of Deployable Composite Structures

In this paper, a multiscale simulation method for analyzing deployable composite structures is presented. Effective shell properties of the composites are obtained based on Mechanics of Structure Genome (MSG) homogenization, and then implemented into a user-subroutine UGENS for structural simulation with shell elements in Abaqus. The column bending test(CBT) of a flat thin flexure and lenticular composite boom in a simplified deployer structure are studied for demonstration. A viscoelastic material model with direct integration is adopted in this paper. The CBT simulation shows good agreement with experiments during relaxation, while errors are observed when comparing residual deformation. It is shown that this CBT model can be calibrated to CBT test results. For the lenticular boom analysis, the complete process of flattening, coiling, stowage, deployment and recovery is simulated with the viscoelastic shell model.

Finite element analysis↗

Deuterostome Evolution: Large Data Set Analysis

This award allowed us to develop novel hardware for phylogenetics, collect genomic data and produce several phylogenies of deuterostome organisms, communicate the results publicly, release software into the public domain, publish textbooks and papers, and prepare for the next research projects. There are no resulting subject inventions to report. We review these activities in three sections: 1) Hardware and software and development; 2) Evolutionary biology research; 3) Our proposed future direction, predictive analysis of pathogens in support of the NASA mission.

Janies, Daniel↗

Biomass yields, reproductive fertility, compositional analysis, and genetic diversity of newly developed triploid giant miscanthus hybrids

Abstract Miscanthus × giganteus (giant miscanthus), first found as a naturally occurring hybrid, has shown promise as a bioenergy/biomass crop throughout much of the temperate world. This allotriploid (2 n = 3 x = 57) hybrid resulted from a cross between tetraploid Miscanthus sacchariflorus (2 n = 4 x = 76) and diploid Miscanthus sinensis (2 n = 2 x = 38) and is particularly desirable due to its low fertility that minimizes reseeding and potential invasiveness. However, there is limited genetic diversity in commonly grown cultivars of triploid M. × giganteus and breeding and development efforts to improve and domesticate this crop have been minimal. Here, we report on newly developed M. × giganteus hybrids compared with the industry standard M. × giganteus '1993‐1780'. Dry biomass yields of new hybrids ranged from 19.5 to 32.4 Mg/ha/year for the fourth growing season, compared with 21.0 Mg/ha/year for M. × giganteus '1993‐1780'. Plant reproductive fertility remained low for all accessions with overall fertility [(seed set × seed germination)/100] ranging from 0.3% to 4.5% for new hybrids compared to 0.4% for M. × giganteus '1993‐1780'. Culm density and height varied among accessions and were positively correlated with increased biomass. Based on compositional analyses, theoretical ethanol yields ranged from 9, 740 to 16,278 L/ha/year for new hybrids compared to 10,406 L/ha/year for M. × giganteus '1993‐1780'. Relative feed value indices were low overall and ranged between 66.0 and 72.8 for new hybrids compared to M. × giganteus '1993‐1780' with 71.3. The genetic diversity of new hybrids, compared with existing cultivars, was characterized using whole genome sequences. Based on pair‐wise distances, cluster analysis clearly showed increased diversity of new hybrids compared with earlier selections. These results document new triploid hybrids of M. × giganteus with enhanced biomass and theoretical ethanol yields in combination with broader genetic diversity and lowreproductive fertility.

Touchell, Darren H.↗

GeneLab

GeneLab collects and enables analysis of spaceflight and ground-based spaceflight simulation genomic data, RNA and protein expression, and metabolic profiles. It interfaces with other existing databases containing spaceflight omic data. The 2011 National Research Council (NRC) Decadal Survey on NASA Life and Physical Sciences called for increased opportunities for multi-investigator spaceflight opportunities and greater use of genomic approaches to meet the needs of NASA researchers. To address these recommendations of the NRC Decadal Survey, the Space Life and Physical Sciences Research and Applications Division of NASA's Human Exploration and Operations Mission Directorate has initiated a transition to an Open Science architecture to increase research opportunities, and has developed the GeneLab Platform based on highly leveraged and integrated bioinformatics analytics. GeneLab is an interactive, open-access resource where scientists can upload, download, store, search, share, transfer, and analyze omics data from spaceflight and corresponding analogue experiments. Users can explore GeneLab datasets in the Data Repository, analyze data using the Analysis Platform, visualize high-order data and create collaborative projects using the Collaborative Workspace. Our primary goal is to maximize the utilization of the valuable biological research conducted aboard the International Space Station (ISS) by collecting genomic, transcriptomic, proteomic, and metabolomics data known as “omics”. By providing a portal linking processed data to flight parameters, GeneLab enables exploration of the molecular network responses of terrestrial biology to the space environment. This allows researchers to understand the complex responses of biological systems to the space environment. This technology development activity was transferred from the Human Exploration and Operations Mission Directorate to the Science Mission Directorate Division of Biological and Physical Sciences (BPS) in October 2020.

GeneLab↗