Search NASASearch

SEARCH · Search NASA

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

6051R & 6051S Assembly and Annotation

We report the draft genomes of two morphologically distinct variants of Bacillus subtilis ATCC 6051 [NCBI3610]. The two isolates exhibit differences in not only morphology but also their genetics, despite identical 16S rRNA sequences. Investigating the genetic differences of colony morphology variation in this model organism can provide valuable insights.

59 BASIC BIOLOGICAL SCIENCES

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B

Signatures of Selection for Resistance/Tolerance to Perkinsus olseni in Grooved Carpet Shell Clam ( Ruditapes decussatus ) Using a Population Genomics Approach

ABSTRACT The grooved carpet shell clam ( Ruditapes decussatus ) is a bivalve of high commercial value distributed throughout the European coast. Its production has suffered a decline caused by different factors, especially by the parasite Perkinsus olsenii . Improving production of R . decussatus requires genomic resources to ascertain the genetic factors underlying resistance/tolerance to P. olseni i . In this study, the first reference genome of R . decussatus was assembled through long‐ and short‐read sequencing (1677 contigs; 1.386 Mb) and further scaffolded at chromosome level with Hi‐C (19 superscaffolds; 95.4% of assembly). Repetitive elements were identified (32%) and masked for annotation of 38,276 coding‐ and 13,056 non‐coding genes. This genome was used as a reference to develop a 2bRAD‐Seq 13,438 SNP panel for a genomic screening on six shellfish beds distributed across the Atlantic Ocean and Mediterranean Sea. Beds were selected by perkinsosis prevalence and the infection level was individually evaluated in all the samples. Genetic diversity was significantly higher in the Mediterranean than in the Atlantic region. The main genetic breakage was detected between those regions (F ST = 0.224), being the Mediterranean more heterogeneous than the Atlantic. Several loci under divergent selection (394 outliers; 261 genomic windows) were detected across shellfish beds. Samples were also inspected to detect signals of selection for resistance/tolerance to P. olseni i by using infection‐level and population‐genomics approaches, and 90 common divergent outliers for resistance/tolerance to perkinsosis were identified and used for gene mining. Candidate genes and markers identified provide invaluable information for controlling perkinsosis and for improving production of the grooved carpet shell clam.

Sambade, Inés M. [Department of Zoology, Genetics

High-quality draft genome sequence of Thermobifida halotolerans DSM 44931

Here, we report the genome sequence of Thermobifida halotolerans DSM 44931, a bacterium that was originally isolated from a salt mine in the Yunnan Province of China. This genome was sequenced using Pacific Biosciences sequencing technology and was assembled into 2 contigs in 2 scaffolds. It has a total length of 5,506,851 bp and a GC content of 71.16%. Functional annotation of this genome provides further metabolic insight into this species.

actinomycete

Conservation of Fold and Topology of Functional Elements in Thiamin Pyrophosphate Enzymes

Thiamin pyrophosphate (TPP)-dependent enzymes are a highly divergent family of proteins binding both TPP and metal ions. They perform decarboxylation-hydroxyaldehydes. Prior -ketoacids and of a common - (O=)C-C(OH)- fragment of to knowledge of three-dimensional structures of these enzmes, the GDGY25-30NN sequence was used to identify these enzymes. Subsequently, a number of structural studies on those enzymes revealed multi-subunit organization and the features of the two duplicate cofactor binding sites. Analyzing the structures of 44 structurally known enzymes, we found that the common structure of these enzymes is reduced to 180-220 amino acid long fragments of two PP and two PYR domains that form the [PP:PYR]2 binding center of two cofactor molecules. The structures of PP and PYR are arranged in a similar fold-sheet with triplets of helices on both sides.Dconsisting of a six-stranded Residues surrounding the cofactors are not strictly conserved, but they provide the same interatomic contacts required for the catalytic functions that these enzymes perform while maintaining interactive structural integrity. These structural and functional amino acids are topological counterparts located in the same positions of the conserved fold of sets of PP and PYR domains. Additional parallels include short fragments of sequences that link these amino acids to the fold and function. This report on the structural commonalities amongst TPP dependent enzymes is thought to contribute new approaches to annotation that may assist in advancing the functional proteomics of TPP dependent enzymes, and trace their complexity within evolutionary context.

Dominiak, P.

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES

[Columbia Sensor Diagrams]

A two dimensional graphical event sequence of the time history of relevant sensor information located in the left wing and wheel well areas of the Space Shuttle Columbia Orbiter is presented. Information contained in this graphical event sequence include: 1) Sensor location on orbiter and its associated wire bindle in X-Y plane; 2) Wire bundle routing; 3) Description of each anomalous sensor event; 4) Time annotation by (a) GMT, (b) time relative to LOS, (c) time history bar, and (d) ground track; and 5) Graphical display of temperature rise (based on delta temperature from point it is determined to be anomalous).

Source record

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS

The small protein SbtC is a functional component of the CO 2 concentrating mechanism in Synechocystis sp. PCC 6803

Oxygenic phototrophs fix CO 2 via the enzyme ribulose-1,5-bisphosphate carboxylase/oxygenase (RubisCO), which shows relatively low CO 2 affinity and specificity. To circumvent low and fluctuating CO 2 concentrations in aquatic systems, cyanobacteria and algae have evolved sophisticated inorganic carbon (Ci) concentrating mechanisms (CCMs). Bicarbonate transporters such as SbtA play a crucial role in the cyanobacterial CCM and hence display multiple layers of tight regulation. Control of sbtA gene expression and corresponding transporter activity involves the PII-like protein SbtB, whose gene is frequently co-transcribed with sbtA. A previously non-annotated gene located upstream of the sbtAB operon in the model Synechocystis sp. PCC 6803 encodes the small protein SbtC, composed of 80 amino acids. Presence of SbtC was confirmed by immunoblotting of the sbtC-coding sequence fused to a Flag-tag. Similar to sbtAB , transcription of the sbtC locus is induced by low CO 2 availability; however, it is controlled independently. Mutation of the sbtC locus in a wild-type background produced only a mild phenotype, even under low CO 2 , but impaired diurnal growth resembled that of the mutant ΔsbtB . Biochemical analysis indicated a trimeric SbtABC complex in the membrane. Bicarbonate leakage from cells was strongly elevated when either sbtB or sbtC was deleted from recombinant Synechocystis strains harboring only SbtA as single Ci uptake system. Here, our results provide evidence that SbtC contributes to the formation of the SbtAB complex, thereby regulating bicarbonate exchange at the cytoplasmic membrane. Well-conserved SbtC-like proteins encoded in the neighborhood of sbtAB exist in many cyanobacterial genomes, pointing toward an important role in the cyanobacterial CCM.

Walke, Peter [Univ. of Rostock (Germany)] (ORCID:0

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits, two catalytic centers, common amino acid sequence, and specific contacts to provide a flip-flop, or alternate site, mechanism of action. Each catalytic center [PP:PYR] is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and aminopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core [PP:PYR]* within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GX@&(G)@XXGQ, and GDGX25-30 within the PP- domain, and the E&(G)@XXG@ within the PYR-domain, where Q, corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, Paulina M.

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity