Search NASA⌕ Search

SEARCH · Search NASA

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii↗

Complex expression patterns of lymphocyte-specific genes during the development of cartilaginous fish implicate unique lymphoid tissues in generating an immune repertoire

Cartilaginous fish express canonical B and T cell recognition genes, but their lymphoid organs and lymphocyte development have been poorly defined. Here, the expression of Ig, TCR, recombination-activating gene (Rag)-1 and terminal deoxynucleosidase (TdT) genes has been used to identify roles of various lymphoid tissues throughout development in the cartilaginous fish, Raja eglanteria (clearnose skate). In embryogenesis, Ig and TCR genes are sharply up-regulated at 8 weeks of development. At this stage TCR and TdT expression is limited to the thymus; later, TCR gene expression appears in peripheral sites in hatchlings and adults, suggesting that the thymus is a source of T cells as in mammals. B cell gene expression indicates more complex roles for the spleen and two special organs of cartilaginous fish-the Leydig and epigonal (gonad-associated) organs. In the adult, the Leydig organ is the site of the highest IgM and IgX expression. However, the spleen is the first site of IgM expression, while IgX is expressed first in gonad, liver, Leydig and even thymus. Distinctive spatiotemporal patterns of Ig light chain gene expression also are seen. A subset of Ig genes is pre-rearranged in the germline of the cartilaginous fish, making expression possible without rearrangement. To assess whether this allows differential developmental regulation, IgM and IgX heavy chain cDNA sequences from specific tissues and developmental stages have been compared with known germline-joined genomic sequences. Both non-productively rearranged genes and germline-joined genes are transcribed in the embryo and hatchling, but not in the adult.

Non-NASA Center↗

Leiomodins: larger members of the tropomodulin (Tmod) gene family

The 64-kDa autoantigen D1 or 1D, first identified as a potential autoantigen in Graves' disease, is similar to the tropomodulin (Tmod) family of actin filament pointed end-capping proteins. A novel gene with significant similarity to the 64-kDa human autoantigen D1 has been cloned from both humans and mice, and the genomic sequences of both genes have been identified. These genes form a subfamily closely related to the Tmods and are here named the Leiomodins (Lmods). Both Lmod genes display a conserved intron-exon structure, as do three Tmod genes, but the intron-exon structure of the Lmods and the Tmods is divergent. mRNA expression analysis indicates that the gene formerly known as the 64-kDa autoantigen D1 is most highly expressed in a variety of human tissues that contain smooth muscle, earning it the name smooth muscle Leiomodin (SM-Lmod; HGMW-approved symbol LMOD1). Transcripts encoding the novel Lmod gene are present exclusively in fetal and adult heart and adult skeletal muscle, and it is here named cardiac Leiomodin (C-Lmod; HGMW-approved symbol LMOD2). Human C-Lmod is located near the hypertrophic cardiomyopathy locus CMH6 on human chromosome 7q3, potentially implicating it in this disease. Our data demonstrate that the Lmods are evolutionarily related and display tissue-specific patterns of expression distinct from, but overlapping with, the expression of Tmod isoforms. Copyright 2001 Academic Press.

Carrier Proteins/biosynthesis/genetics↗

Benchmarking Computational Tools for Calling SNPs and Indels in Complex Microbial Populations

The NASA BioNutrients missions seek to understand the suitability of microorganisms for bioproduction during space flight. One topic of interest is the stability of microbial genomes during long-term ambient storage and subsequent rehydration and growth. To address these questions, samples from 8 species were flown to ISS for 5 years of desiccated storage at ambient temperature (Stasis Packs) and 2 species were packaged along with powdered media inside a bioreactor system to allow hydration and growth in microgravity (Production Packs). For both systems, Whole Genome Sequencing (WGS) of the DNA extracted from the returned samples and paired ground controls will be conducted to identify changes in genome stability due to time, storage conditions and growth in space. Across the technical replicates, ground controls, 10 timepoints, and multiple experimental conditions, ~300 samples have been selected for initial analysis with WGS sequencing to 100x coverage. A flexible and resource efficient mutation calling pipeline is needed to process this large dataset and allow for comparisons between species. Many bioinformatics tools for calling Indels and Single Nucleotide Variants (SNVs) are designed for use with pure isolates, where true variations from the reference genome are expected to dominate the reads aligning to the location of mutation. In contrast, DNA from the Stasis Pack (SP) samples was collected directly after recovery from desiccated storage and the Production Pack (PP) samples were collected after fermentation. In this context, reads with mutations are expected to be less frequent than reads that align with the reference genome, as each sample will include multiple lines of cells. Thus, BioNutrients samples are expected to be similar to samples from cancer cell or “pooled” sequencing approaches. In preparation for the analysis of the BioNutrients samples, we have tested three mutation calling tools (GATK for Microbes, BreSeq and DiscoSNP) designed for complex samples. A challenge of validating mutation identification pipelines is a lack of “Ground Truth” datasets, especially for complex samples. To compare these three tools, we sought to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO as part of the Space Algae-1 mission. Here we present a summary of these tools against the analysis originally conducted using the CRISP tool. Critical metrics are compared such as runtime, the number of SNPs, the number and size of Indels, and patterns of transversion and transitions identified by each tool are reported. By sharing these benchmarking results collected in support of the BioNutrients mission, we aim to guide others seeking to identify SNVs in similarly complex microbial samples.

Biology↗

Genomic Analysis of Aspergillus Section Terrei Reveals a High Potential in Secondary Metabolite Production and Plant Biomass Degradation

Aspergillus terreus has attracted interest due to its application in industrial biotechnology, particularly for the production of itaconic acid and bioactive secondary metabolites. As related species also seem to possess a prosperous secondary metabolism, they are of high interest for genome mining and exploitation. Here, we present draft genome sequences for six species from Aspergillus section Terrei and one species from Aspergillus section Nidulantes. Whole-genome phylogeny confirmed that section Terrei is monophyletic. Genome analyses identified between 70 and 108 key secondary metabolism genes in each of the genomes of section Terrei, the highest rate found in the genus Aspergillus so far. The respective enzymes fall into 167 distinct families with most of them corresponding to potentially unique compounds or compound families. Moreover, 53% of the families were only found in a single species, which supports the suitability of species from section Terrei for further genome mining. Intriguingly, this analysis, combined with heterologous gene expression and metabolite identification, suggested that species from section Terrei use a strategy for UV protection different to other species from the genus Aspergillus. Section Terrei contains a complete plant polysaccharide degrading potential and an even higher cellulolytic potential than other Aspergilli, possibly facilitating additional applications for these species in biotechnology.

60 APPLIED LIFE SCIENCES↗

An archaeal genomic signature

Comparisons of complete genome sequences allow the most objective and comprehensive descriptions possible of a lineage's evolution. This communication uses the completed genomes from four major euryarchaeal taxa to define a genomic signature for the Euryarchaeota and, by extension, the Archaea as a whole. The signature is defined in terms of the set of protein-encoding genes found in at least two diverse members of the euryarchaeal taxa that function uniquely within the Archaea; most signature proteins have no recognizable bacterial or eukaryal homologs. By this definition, 351 clusters of signature proteins have been identified. Functions of most proteins in this signature set are currently unknown. At least 70% of the clusters that contain proteins from all the euryarchaeal genomes also have crenarchaeal homologs. This conservative set, which appears refractory to horizontal gene transfer to the Bacteria or the Eukarya, would seem to reflect the significant innovations that were unique and fundamental to the archaeal "design fabric." Genomic protein signature analysis methods may be extended to characterize the evolution of any phylogenetically defined lineage. The complete set of protein clusters for the archaeal genomic signature is presented as supplementary material (see the PNAS web site, www.pnas.org).

Non-NASA Center↗

Genome-resolved analysis of Serratia marcescens strain SMTT infers niche specialization as a hydrocarbon-degrader

Abstract Bacteria that are chronically exposed to high levels of pollutants demonstrate genomic and corresponding metabolic diversity that complement their strategies for adaptation to hydrocarbon-rich environments. Whole genome sequencing was carried out to infer functional traits of Serratia marcescens strain SMTT recovered from soil contaminated with crude oil. The genome size (Mb) was 5,013,981 with a total gene count of 4,842. Comparative analyses with carefully selected S. marcescens strains, 2 of which are associated with contaminated soil, show conservation of central metabolic pathways in addition to intra-specific genetic diversity and metabolic flexibility. Genome comparisons also indicated an enrichment of genes associated with multidrug resistance and efflux pumps for SMTT. The SMTT genome contained genes that enable the catabolism of aromatic compounds via the protocatechuate para-degradation pathway, in addition to meta-cleavage of catechol (meta-cleavage pathway II); gene enrichment for aromatic compound degradation was markedly higher for SMTT compared to the other S. marcescens strains analysed. Our data presents a valuable genetic inventory for future studies on strains of S. marcescens and provides insights into those genomic features of SMTT with industrial potential.

Genetics & Heredity↗

Conserved gene clusters in bacterial genomes provide further support for the primacy of RNA

Five complete bacterial genome sequences have been released to the scientific community. These include four (eu)Bacteria, Haemophilus influenzae, Mycoplasma genitalium, M. pneumoniae, and Synechocystis PCC 6803, as well as one Archaeon, Methanococcus jannaschii. Features of organization shared by these genomes are likely to have arisen very early in the history of the bacteria and thus can be expected to provide further insight into the nature of early ancestors. Results of a genome comparison of these five organisms confirm earlier observations that gene order is remarkably unpreserved. There are, nevertheless, at least 16 clusters of two or more genes whose order remains the same among the four (eu)Bacteria and these are presumed to reflect conserved elements of coordinated gene expression that require gene proximity. Eight of these gene orders are essentially conserved in the Archaea as well. Many of these clusters are known to be regulated by RNA-level mechanisms in Escherichia coli, which supports the earlier suggestion that this type of regulation of gene expression may have arisen very early. We conclude that although the last common ancestor may have had a DNA genome, it likely was preceded by progenotes with an RNA genome.

Non-NASA Center↗

Methanococcus jannaschii genome: revisited

Analysis of genomic sequences is necessarily an ongoing process. Initial gene assignments tend (wisely) to be on the conservative side (Venter, 1996). The analysis of the genome then grows in an iterative fashion as additional data and more sophisticated algorithms are brought to bear on the data. The present report is an emendation of the original gene list of Methanococcus jannaschii (Bult et al., 1996). By using a somewhat more updated database and more relaxed (and operator-intensive) pattern matching methods, we were able to add significantly to, and in a few cases amend, the gene identification table originally published by Bult et al. (1996).

Non-NASA Center↗

Comparative genomics provides insights into the cold adaptation of endophytic fungi associated with Deschampsia antarctica

Endophytic fungi from Deschampsia antarctica , the southernmost flowering plant, provide insights into the cold adaptation mechanisms of plant-associated fungi in extreme environments. This study presents the genome sequences and comparative analysis of eight fungal isolates from D. antarctica leaves. These Antarctic fungal isolates were analyzed alongside 121 plant-associated fungal genomes to uncover signatures of adaptation and endophytic specialization. Antarctic endophytes show striking patterns, including reduced genome size (∼26.3 Mb on average), streamlined gene content (∼8844 genes), and notably small secretomes (∼288 proteins). Despite this reduced gene repertoire, they maintain a robust set of genes encoding carbohydrate-active enzymes (CAZymes) but lack those for lignin and bacterial cell wall degradation, indicating a symbiotic lifestyle that avoids host damage and predation. One isolate, Alternaria sp. UNIPAMPA017 stood out, with 26% of its genome occupied by transposable elements. Lifestyle, rather than phylogeny, was the main driver of CAZyme and secretome profiles, underscoring ecological convergence. Compared to endophytes from Arabidopsis and Populus, D. antarctica endophytes harbor fewer pectin-degrading enzymes, reflecting their adaptation to the cell wall structure of their monocot host. Together, these fungi reveal a pattern of genomic reduction and functional fine-tuning, hallmarks of life adapted to persist in cold, nutrient-scarce niches.

Ascomycota↗

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R 2 of 0.604 [0.593–0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156–5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

COVID19↗

Pangenomes suggest ecological-evolutionary responses to experimental soil warming

ABSTRACT Below-ground carbon transformations that contribute to healthy soils represent a natural climate change mitigation, but newly acquired traits adaptive to climate stress may alter microbial feedback mechanisms. To better define microbial evolutionary responses to long-term climate warming, we study microorganisms from an ongoing in situ soil warming experiment where, for over three decades, temperate forest soils are continuously heated at 5°C above ambient. We hypothesize that across generations of chronic warming, genomic signatures within diverse bacterial lineages reflect adaptations related to growth and carbon utilization. From our bacterial culture collection isolated from experimental heated and control plots, we sequenced genomes representing dominant taxa sensitive to warming, including lineages of Actinobacteria, Alphaproteobacteria, and Betaproteobacteria. We investigated genomic attributes and functional gene content to identify signatures of adaptation. Comparative pangenomics revealed accessory gene clusters related to central metabolism, competition, and carbon substrate degradation, with few functional annotations explicitly associated with long-term warming. Trends in functional gene patterns suggest genomes from heated plots were relatively enriched in central carbohydrate and nitrogen metabolism pathways, while genomes from control plots were relatively enriched in amino acid and fatty acid metabolism pathways. We observed that genomes from heated plots had less codon bias, suggesting potential adaptive traits related to growth or growth efficiency. Codon usage bias varied for organisms with similar 16S rrn operon copy number, suggesting that these organisms experience different selective pressures on growth efficiency. Our work suggests the emergence of lineage-specific trends as well as common ecological-evolutionary microbial responses to climate change. IMPORTANCE Anthropogenic climate change threatens soil ecosystem health in part by altering below-ground carbon cycling carried out by microbes. Microbial evolutionary responses are often overshadowed by community-level ecological responses, but adaptive responses represent potential changes in traits and functional potential that may alter ecosystem function. We predict that microbes are adapting to climate change stressors like soil warming. To test this, we analyzed the genomes of bacteria from a soil warming experiment where soil plots have been experimentally heated 5°C above ambient for over 30 years. While genomic attributes were unchanged by long-term warming, we observed trends in functional gene content related to carbon and nitrogen usage and genomic indicators of growth efficiency. These responses may represent new parameters in how soil ecosystems feedback to the climate system.

Choudoir, Mallory J. (ORCID:0000000291175150)↗

Host population dynamics influence Leptospira spp. transmission patterns among Rattus norvegicus in Boston, Massachusetts, US

Leptospirosis (caused by pathogenic bacteria in the genus Leptospira ) is prevalent worldwide but more common in tropical and subtropical regions. Transmission can occur following direct exposure to infected urine from reservoir hosts, or a urine-contaminated environment, which then can serve as an infection source for additional rats and other mammals, including humans. The brown rat, Rattus norvegicus , is an important reservoir of Leptospira spp. in urban settings. We investigated the presence of Leptospira spp. among brown rats in Boston, Massachusetts and hypothesized that rat population dynamics in this urban setting influence the transportation, persistence, and diversity of Leptospira spp. We analyzed DNA from 328 rat kidney samples collected from 17 sites in Boston over a seven-year period (2016–2022); 59 rats representing 12 of 17 sites were positive for Leptospira spp. We used 21 neutral microsatellite loci to genotype 311 rats and utilized the resulting data to investigate genetic connectivity among sampling sites. We generated whole genome sequences for 28 Leptospira spp. isolates obtained from frozen and fresh tissue from some of the 59 positive rat kidneys. When isolates were not obtained, we attempted genomic DNA capture and enrichment, which yielded 14 additional Leptospira spp. genomes from rats. We also generated an enriched Leptospira spp. genome from a 2018 human case in Boston. We found evidence of high genetic structure among rat populations that is likely influenced by major roads and/or other dispersal barriers, resulting in distinct rat population groups within the city; at certain sites these groups persisted for multiple years. We identified multiple distinct phylogenetic clades of L. interrogans among rats that were tightly linked to distinct rat populations. This pattern suggests L. interrogans persists in local rat populations and its transportation is influenced by rat population dynamics. Finally, our genomic analyses of the Leptospira spp. detected in the 2018 human leptospirosis case in Boston suggests a link to rats as the source. These findings will be useful for guiding rat control and human leptospirosis mitigation efforts in this and other similar urban settings.

Stone, Nathan E.↗

Rapid Detection and Quick Characterization of African Swine Fever Virus Using the VolTRAX Automated Library Preparation Platform

African swine fever virus (ASFV) is the causative agent of a severe and highly contagious viral disease affecting domestic and wild swine. The current ASFV pandemic strain has a high mortality rate, severely impacting pig production and, for countries suffering outbreaks, preventing the export of their pig products for international trade. Early detection and diagnosis of ASFV is necessary to control new outbreaks before the disease spreads rapidly. One of the rate-limiting steps to identify ASFV by next-generation sequencing platforms is library preparation. Here, we investigated the capability of the Oxford Nanopore Technologies’ VolTRAX platform for automated DNA library preparation with downstream sequencing on Nanopore sequencing platforms as a proof-of-concept study to rapidly identify the strain of ASFV. Within minutes, DNA libraries prepared using VolTRAX generated near-full genome sequences of ASFV. Thus, our data highlight the use of the VolTRAX as a platform for automated library preparation, coupled with sequencing on the MinION Mk1C for field sequencing or GridION within a laboratory setting. These results suggest a proof-of-concept study that VolTRAX is an effective tool for library preparation that can be used for the rapid and real-time detection of ASFV.

60 APPLIED LIFE SCIENCES↗

Nanoscale Bio-engineering Solutions for Space Exploration: The Nanopore Sequencer

Characterization of biological systems at the molecular level and extraction of essential information for nano-engineering design to guide the nano-fabrication of solid-state sensors and molecular identification devices is a computational challenge. The alpha hemolysin protein ion channel is used as a model system for structural analysis of nucleic acids like DNA. Applied voltage draws a DNA strand and surrounding ionic solution through the biological nanopore. The subunits in the DNA strand block ion flow by differing amounts. Atomistic scale simulations are employed using NASA supercomputers to study DNA translocation, with the aim to enhance single DNA subunit identification. Compared to protein channels, solid-state nanopores offer a better temporal control of the translocation of DNA and the possibility to easily tune its chemistry to increase the signal resolution. Potential applications for NASA missions, besides real-time genome sequencing include astronaut health, life detection and decoding of various genomes.

Stolc, Viktor↗

Nanoscale Bioengineering Solutions for Space Exploration the Nanopore Sequencer

Characterization of biological systems at the molecular level and extraction of essential information for nano-engineering design to guide the nano-fabrication of solid-state sensors and molecular identification devices is a computational challenge. The alpha hemolysin protein ion channel is used as a model system for structural analysis of nucleic acids like DNA. Applied voltage draws a DNA strand and surrounding ionic solution through the biological nanopore. The subunits in the DNA strand block ion flow by differing amounts. Atomistic scale simulations are employed using NASA supercomputers to study DNA translocation. with the aim to enhance single DNA subunit identification. Compared to protein channels, solid-state nanopores offer a better temporal control of the translocation of DNA and the possibility to easily tune its chemistry to increase the signal resolution. Potential applications for NASA missions, besides real-time genome sequencing include astronaut health, life detection and decoding of various genomes. http://phenomrph.arc.nasa.gov/index.php

Ioana, Cozmuta↗