Search NASA⌕ Search

SEARCH · Search NASA

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Evolution, language and analogy in functional genomics

Almost a century ago, Wittgenstein pointed out that theory in science is intricately connected to language. This connection is not a frequent topic in the genomics literature. But a case can be made that functional genomics is today hindered by the paradoxes that Wittgenstein identified. If this is true, until these paradoxes are recognized and addressed, functional genomics will continue to be limited in its ability to extrapolate information from genomic sequences.

Evolution, Molecular↗

Monitoring Astronaut Health with DNA Sequencing

In recent years microbe a plethora of microbe populations have been identified onboard the ISS (International Space Station). Approaches for real-time tracking of microbes for routine housekeeping and food/water safety monitoring will be critical for mission safety and crew health on future longer duration missions to the Moon or Mars. This work is a proof-of-concept study demonstrating an end-to-end phylogenetic identification and full genome sequencing effort of multiple microbial populations. Our methodology utilized the ISS flight-certified WetLab-2 molecular toolbox and the Biomolecule Sequencer projects for real-time end-to-end on-orbit microbial biological samples processing and molecular analysis with real time results generated utilizing only field "offline" analytic software. For this experiment we colony-cultured several ISS isolated microorganisms before generation of the pre-sequencing library via the automated VolTRAX device which enabled high library turnover with little wet-bench activity or potential future costly astronaut time. The pre-sequencing library is diluted in loading buffer and injected into the MinION sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown. 16S and full genome alignment, nucleotide matching, gene identification, and phylogenetic sorting was accomplished utilizing the Epi2me software and the offline NCBI Blast viral, microbiome, and human somatic databases. In short, the methodologies developed herein replace the myriad of specific, often highly targeted microbiological tests used in the clinical laboratory, which would be difficult if not impossible to currently implement aboard the ISS or in deep space, with a single metagenomics test.

genomics↗

Aeromonas in South Asia: genomic insights into an environmental pathogen and reservoir of antimicrobial resistance

Aeromonads are an ecologically versatile group of bacteria that cause infections in aquatic animals and are recognised as emerging human pathogens. Despite this, our understanding of Aeromonas diversity, especially the relationship between clinical and environmental strains, remains limited. Here, we present a genomic analysis of the Aeromonas genus, comprising 1853 genomes, and a detailed comparison of clinical and environmental strains from South Asia, including 996 newly sequenced genomes from Bangladesh and India. Phylogenetic analyses revealed that Aeromonas is a highly diverse genus, with no distinct clade separating clinical and environmental isolates. We identified 28 Aeromonas species and 905 novel sequence types, comprising 72.5% of the genomes. Notably, we show a high incidence of antimicrobial resistance (AMR) genes across all isolates, including against front and last-line antibiotics. Finally, we highlight frequent misidentification of Aeromonas as Vibrio cholerae, which is relevant to cholera-endemic regions where both genera co-exist and are associated with diarrhoeal disease. Our study underscores Aeromonas as an important environmental AMR reservoir and emerging multi-species pathogen capable of spilling over into human populations.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient mutagenesis and genotyping of maize inbreds using biolistics, multiplex CRISPR/Cas9 editing, and Indel-Selective PCR

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Efficient Mutagenesis and Genotyping of Maize Inbreds Using Biolistics, Multiplex CRISPR/Cas9 Editing, and Indel-Selective PCR"

CRISPR/Cas9 based genome editing has advanced our understanding of a myriad of important biological phenomena. Important challenges to multiplex genome editing in maize include assembly of large complex DNA constructs, few genotypes with efficient transformation systems, and costly/labor-intensive genotyping methods. Here we present an approach for multiplex CRISPR/Cas9 genome editing system that delivers a single compact DNA construct via biolistics to Type I embryogenic calli, followed by a novel efficient genotyping assay to identify desirable editing outcomes. We first demonstrate the creation of heritable mutations at multiple target sites within the same gene. Next, we successfully created individual and stacked mutations for multiple members of a gene family. Genome sequencing found off-target mutations are rare. Multiplex genome editing was achieved for both the highly transformable inbred line H99 and Illinois Low Protein1 (ILP1), a genotype where transformation has not previously been reported. In addition to screening transformation events for deletion alleles by PCR, we also designed PCR assays that selectively amplify deletion or insertion of a single nucleotide, the most common outcome from DNA repair of CRISPR/Cas9 breaks by non-homologous end-joining. The Indel-Selective PCR (IS-PCR) method enabled rapid tracking of multiple edited alleles in progeny populations. The ‘end to end’ pipeline presented here for multiplexed CRISPR/Cas9 mutagenesis can be applied to accelerate maize functional genomics in a broader diversity of genetic backgrounds.

gene editing↗

Cas3-Mediated Genome Reduction: Demonstration in Cupriavidus Necator H16 Improves Growth on Heterotrophic and Autotrophic Carbon Sources

Genome reduction is widely used to improve microbial bioprocessing hosts by reducing the burden of inessential physiology. Rationally identifying genomic regions that are dispensable or even detrimental to bioprocessing is challenged by our inability to map genome sequence to function across complex regulation and physiology. Thus, there is a need for tools that rapidly generate reduced genome strains with improved performance in process-relevant conditions. Here, we report a Cascade-Cas3-enabled method called TRIM3 that generates large deletions by targeting a randomly integrated transposon, enabling facile generation of a genome-reduced mutant library. Mutants with improved performance were isolated following growth-coupled selection and analyzed by long-read DNA sequencing to identify deletions in their genomes. We deploy this system iteratively in the industrial host Cupriavidus necator H16 on fructose and on formate. After two rounds of TRIM3, we isolate a strain containing a total reduction of 1.4 Mb (18.4% of the genome) that grows 25% faster in a bioreactor on fructose and a strain with a total reduction of 0.5 Mb (7.3% of the genome) that grows 14% faster on formate. This work demonstrates a method for random, iterative, growth-selectable genome reduction that represents a new avenue for large-scale genome modifications and the development of improved bioprocessing hosts.

09 BIOMASS FUELS↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

genomeocean: a pretrained microbial genome foundational model (genomeoceanLLM) v1.0

We present Genomeocean, a foundational genome language model that represents the microbial genome sequences from complex environmental samples. By training on a large, diverse metagenomic dataset, Genomeocean learns species-specific sequence composition and can generate long, realistic open reading frames (ORFs). Our model employs a Byte-pair-encoding (BPE) tokenization strategy, allowing it to efficiently process large genomic datasets and generate long sequences up to 50kb. We demonstrate that fine-tuning Genomeocean can generate novel gene clusters encoding biosynthetic pathways, showcasing its ability to model both fundamental and complex biological processes. Our work establishes Genomeocean as a powerful tool for understanding microbial genome biology and paves the way for its application in a range of fields, from synthetic biology to microbiome research.

Wang, Zhong [Lawrence Berkeley National Laboratory↗

The Ciona intestinalis genome: when the constraints are off

The recent genome sequencing of a non-vertebrate deuterostome, the ascidian tunicate Ciona intestinalis, makes a substantial contribution to the fields of evolutionary and developmental biology.1 Tunicates have some of the smallest bilaterian genomes, embryos with relatively few cells, fixed lineages and early determination of cell fates. Initial analyses of the C. intestinalis genome indicate that it has been evolving rapidly. Comparisons with other bilaterians show that C. intestinalis has lost a number of genes, and that many genes linked together in most other bilaterians have become uncoupled. In addition, a number of independent, lineage-specific gene duplications have been detected. These new results, although interesting in themselves, will take on a deeper significance once the genomes of additional invertebrate deuterostomes (e.g. echinoderms, hemichordates and amphioxus) have been sequenced. With such a broadened database, comparative genomics can begin to ask pointed questions about the relationship between the evolution of genomes and the evolution of body plans. Copyright 2003 Wiley Periodicals, Inc.

Review, Tutorial↗

Data Mining of Groundwater to Identify MAGs with Methane, Propane and Toluene Monooxygenases

Whole genome sequencing datasets, involving more than 600 groundwater samples, from nine countries, were analyzed to identify metagenome assembled genomes (MAGs) containing full operons for propane monooxygenase, soluble methane monooxygease, toluene monooxygenase and particulate ammonia/methane monooxygenase. The enzymes encoded by these genes are a focus of interest because of their ability to degrade common groundwater contaminants. Due to the large amount of data, sequence analyses involved more than 80 individual KBase narratives. The approach followed the KBase tutorial called "Metagenome-Assembled Genome Extraction from a Compost Microbiome Enrichment" The generated MAGs were exported from each individual narrative into separate summary KBase narratives for each monooxygenase. Three KBase narratives were generated for particulate ammonia/methane monooxygenase, due to the large number of MAGs identified.

59 BASIC BIOLOGICAL SCIENCES↗

Vanderwaltozyma urihicola sp. nov., a yeast species isolated from rotting wood and beetles in a Brazilian Amazonian rainforest biome

Five yeast isolates belonging to a candidate for novel species were obtained from rotting wood and the gut of a passalid beetle larva in a site of Amazonian rainforest biome in Brazil. Sequence analysis of the Internal Transcribed Spacer (ITS)-5.8S region and the D1/D2 domains of the large subunit rRNA gene showed that the isolates represent a novel species of the genus Vanderwaltozyma. The closest relative of the novel species is Vanderwaltozyma huisunica. These species differs due to 44 nt substitutions and 21 indels in the sequences of the ITS region, as well as by 15 substitutions and four indels in the sequences of the D1/D2 domains. A phylogenomic analysis of the Vanderwaltozyma species with genomes sequenced showed that this novel species is an outgroup to the other species of this genus. We propose the name Vanderwaltozyma urihicola sp. nov. (CBS 18107T, MycoBank MB 856975) to accommodate these isolates. Furthermore, the species is homothallic, producing one to two ascospores per ascus. The habitat of V. urihicola is rotting wood in the Brazilian Amazonian rainforest biome.

Amazonian Forest↗

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

High-quality Acinetobacter genomes recovered from combat wounds via metagenomic sequencing resemble cultured isolate genomes

The ability to accurately characterize wound pathogens is critical to informing clinical decisions for wound infections with complex treatment requirements. Acinetobacter baumannii is an impactful nosocomial pathogen in combat wounds and civilian hospital-acquired infections. An informed understanding of the phylogenetics and epidemiology of A. baumannii infections in military and civilian environments could guide approaches that improve antibiotic treatment regimens for both military and civilian patients. Whole-genome data for bacterial strains can be difficult to obtain due to challenges in culturing isolates from preserved military specimens. Metagenomic sequencing and assembly create opportunities for genomic analysis of pathogens directly from clinical specimens. The ability to perform comparative analyses between metagenome-derived genomes and culture-derived genomes would support a range of comparative bacterial genomic studies. Wound tissue biopsy and effluent samples from combat injuries were subjected to metagenomic sequencing and assembly. In total, 42 microbial metagenome-assembled genomes (MAGs) were obtained directly from metagenomic sequence data, 36 of which were designated “high” quality. Thirty of these genomes corresponded to Acinetobacter, with 29 mapping specifically to A. baumannii. Other observed genera included Bordetella, Citrobacter, Escherichia, and Pseudomonas. Single-copy and multi-copy orthologs were identified across Acinetobacter MAGs and publicly available isolate genomes derived from military and civilian sources. Both MAG and military isolate genomes were annotated with antimicrobial resistance data, and MAG genomes were statistically comparable to genomes obtained from isolates. Our results highlight the potential of de novo metagenome assembly for enabling high-resolution characterization directly from clinical specimens, thereby improving diagnostic precision, guiding antimicrobial stewardship, and enhancing understanding of pathogen evolution across diverse healthcare and battlefield environments.

Acinetobacter baumannii↗

Complex expression patterns of lymphocyte-specific genes during the development of cartilaginous fish implicate unique lymphoid tissues in generating an immune repertoire

Cartilaginous fish express canonical B and T cell recognition genes, but their lymphoid organs and lymphocyte development have been poorly defined. Here, the expression of Ig, TCR, recombination-activating gene (Rag)-1 and terminal deoxynucleosidase (TdT) genes has been used to identify roles of various lymphoid tissues throughout development in the cartilaginous fish, Raja eglanteria (clearnose skate). In embryogenesis, Ig and TCR genes are sharply up-regulated at 8 weeks of development. At this stage TCR and TdT expression is limited to the thymus; later, TCR gene expression appears in peripheral sites in hatchlings and adults, suggesting that the thymus is a source of T cells as in mammals. B cell gene expression indicates more complex roles for the spleen and two special organs of cartilaginous fish-the Leydig and epigonal (gonad-associated) organs. In the adult, the Leydig organ is the site of the highest IgM and IgX expression. However, the spleen is the first site of IgM expression, while IgX is expressed first in gonad, liver, Leydig and even thymus. Distinctive spatiotemporal patterns of Ig light chain gene expression also are seen. A subset of Ig genes is pre-rearranged in the germline of the cartilaginous fish, making expression possible without rearrangement. To assess whether this allows differential developmental regulation, IgM and IgX heavy chain cDNA sequences from specific tissues and developmental stages have been compared with known germline-joined genomic sequences. Both non-productively rearranged genes and germline-joined genes are transcribed in the embryo and hatchling, but not in the adult.

Non-NASA Center↗

Leiomodins: larger members of the tropomodulin (Tmod) gene family

The 64-kDa autoantigen D1 or 1D, first identified as a potential autoantigen in Graves' disease, is similar to the tropomodulin (Tmod) family of actin filament pointed end-capping proteins. A novel gene with significant similarity to the 64-kDa human autoantigen D1 has been cloned from both humans and mice, and the genomic sequences of both genes have been identified. These genes form a subfamily closely related to the Tmods and are here named the Leiomodins (Lmods). Both Lmod genes display a conserved intron-exon structure, as do three Tmod genes, but the intron-exon structure of the Lmods and the Tmods is divergent. mRNA expression analysis indicates that the gene formerly known as the 64-kDa autoantigen D1 is most highly expressed in a variety of human tissues that contain smooth muscle, earning it the name smooth muscle Leiomodin (SM-Lmod; HGMW-approved symbol LMOD1). Transcripts encoding the novel Lmod gene are present exclusively in fetal and adult heart and adult skeletal muscle, and it is here named cardiac Leiomodin (C-Lmod; HGMW-approved symbol LMOD2). Human C-Lmod is located near the hypertrophic cardiomyopathy locus CMH6 on human chromosome 7q3, potentially implicating it in this disease. Our data demonstrate that the Lmods are evolutionarily related and display tissue-specific patterns of expression distinct from, but overlapping with, the expression of Tmod isoforms. Copyright 2001 Academic Press.

Carrier Proteins/biosynthesis/genetics↗