Search NASASearch

SEARCH · Search NASA

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

AlloSHP: deconvoluting single homeologous polymorphism for phylogenetic analysis of allopolyploids

Background The genomic and evolutionary study of allopolyploid organisms involves multiple copies of homeologous chromosomes, making their assembly, annotation, and phylogenetic analysis challenging. Bioinformatics tools and protocols have been developed to study polyploid genomes, but sometimes require the assembly of their genomes, or at least the genes, limiting their use. Results We have developed AlloSHP, a command-line tool for detecting and extracting single homeologous polymorphisms (SHPs) from the subgenomes of allopolyploid species. This tool integrates three main algorithms, WGA, VCF2ALIGNMENT and VCF2SYNTENY, and allows the detection of SHPs for the study of diploid-polyploid complexes with available diploid progenitor genomes, without assembling and annotating the genomes of the allopolyploids under study. AlloSHP has been validated on three diploid-polyploid plant complexes, Brachypodium, Brassica, and Triticum-Aegilops, and a set of synthetic hybrid yeasts and their progenitors of the genus Saccharomyces. The results and congruent phylogenies obtained from the four datasets demonstrate the potential of AlloSHP for the evolutionary analysis of allopolyploids with a wide range of ploidy and genome sizes. Conclusions AlloSHP combines the strategies of simultaneous mapping against multiple reference genomes and syntenic alignment of these genomes to call SHPs, using as input data a single VCF file and the reference genomes of the known or closest extant diploid progenitor species. This novel approach provides a valuable tool for the evolutionary study of allopolyploid species, both at the interspecific and intraspecific levels, allowing the simultaneous analysis of a large number of accessions and avoiding the complex process of assembling polyploid genomes.

Allopolyploids

Nitrogen limitation causes a seismic shift in redox state and phosphorylation of proteins implicated in carbon flux and lipidome remodeling in Rhodotorula toruloides

Background: Oleaginous yeast are prodigious producers of oleochemicals, offering alternative and secure sources for applications in foodstuff, skincare, biofuels, and bioplastics. Nitrogen starvation is the primary strategy used to induce oil accumulation in oleaginous yeast as part of a global stress response. While research has demonstrated that post-translational modifications (PTMs), including phosphorylation and protein cysteine thiol oxidation (redox PTMs), are involved in signaling pathways that regulate stress responses in metazoa and algae, their role in oleaginous yeast remain understudied and unexplored. Results: Towards linking the yeast oleaginous phenotype to protein function, we integrated lipidomics, redox proteomics, and phosphoproteomics to investigate Rhodotorula toruloides under nitrogen-rich and starved conditions over time. Our lipidomics results unearthed interactions involving sphingolipids and cardiolipins with ER stress and mitophagy. Our redox and phosphoproteomics data highlighted the roles of the AMPK, TOR, and calcium signaling pathways in regulation of lipogenesis, autophagy, and oxidative stress response. As a first, we also demonstrated that lipogenic enzymes including fatty acid synthase are modified as a consequence of shifts in cellular redox states due to nutrient availability. Conclusions: We conclude that lipid accumulation is largely a consequence of carbon rerouting and autophagy governed by changes to PTMs, and not increases in the abundance of enzymes involved in central carbon metabolism and fatty acid biosynthesis. Our systems-level approach sets the stage for acquiring multidimensional data sets for protein structural modeling and predicting the functional relevance of PTMs using Artificial Intelligence/Machine Learning (AI/ML). Coupled to those bioinformatics approaches, the putative PTM switches that we delineate will enable advanced metabolic engineering strategies to decouple lipid accumulation from nitrogen limitation.

Lipid Signalling

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence

Demographic drivers of gut microbiome diversity

Abstract The gut microbiome plays a central role in orchestrating metabolic, immune, and neurological functions essential for human health. While extensive research has explored the effects of diseases and pathological conditions on gut microbiome composition, the influence of demographic factors remains underexplored, limiting our understanding of microbiome variations in disease states. This study addresses this gap by investigating the impact of demographic variables, including age, sex, and geography, on gut microbiome diversity in healthy individuals. Using the American Gut Project’s extensive dataset and the QIIME2 bioinformatics pipeline, we conducted a comprehensive analysis of microbial profiles across diverse demographic groups. Our results revealed significant age-related shifts in microbial richness and composition, and geographic location strongly influenced phylogenetic diversity. In contrast, sex exhibited limited impact on microbial diversity within healthy BMI ranges. These findings highlight the critical role of demographic factors in shaping gut microbiome diversity, providing a foundational framework to better contextualize disease-related microbiome variations and advance personalized healthcare approaches.

Biotechnology & Applied Microbiology

Tunturi virus isolates and metagenome-assembled viral genomes provide insights into the virome of Acidobacteriota in Arctic tundra soils

Arctic soils are climate-critical areas, where microorganisms play crucial roles in nutrient cycling processes. Acidobacteriota are phylogenetically and physiologically diverse bacteria that are abundant and active in Arctic tundra soils. Still, surprisingly little is known about acidobacterial viruses in general and those residing in the Arctic in particular. Here, we applied both culture-dependent and -independent methods to study the virome of Acidobacteriota in Arctic soils. Five virus isolates, Tunturi 1–5, were obtained from Arctic tundra soils, Kilpisjärvi, Finland (69°N), using Tunturiibacter spp. strains originating from the same area as hosts. The new virus isolates have tailed particles with podo- (Tunturi 1, 2, 3), sipho- (Tunturi 4), or myovirus-like (Tunturi 5) morphologies. The dsDNA genomes of the viral isolates are 63–98 kbp long, except Tunturi 5, which is a jumbo phage with a 309-kbp genome. Tunturi 1 and Tunturi 2 share 88% overall nucleotide identity, while the other three are not related to one another. For over half of the open reading frames in Tunturi genomes, no functions could be predicted. To further assess the Acidobacteriota-associated viral diversity in Kilpisjärvi soils, bulk metagenomes from the same soils were explored and a total of 1881 viral operational taxonomic units (vOTUs) were bioinformatically predicted. Almost all vOTUs (98%) were assigned to the class Caudoviricetes. For 125 vOTUs, including five (near-)complete ones, Acidobacteriota hosts were predicted. Acidobacteriota-linked vOTUs were abundant across sites, especially in fens. Terriglobia-associated proviruses were observed in Kilpisjärvi soils, being related to proviruses from distant soils and other biomes. Approximately genus- or higher-level similarities were found between the Tunturi viruses, Kilpisjärvi vOTUs, and other soil vOTUs, suggesting some shared groups of Acidobacteriota viruses across soils. This study provides acidobacterial virus isolates as laboratory models for future research and adds insights into the diversity of viral communities associated with Acidobacteriota in tundra soils. Predicted virus-host links and viral gene functions suggest various interactions between viruses and their host microorganisms. Largely unknown sequences in the isolates and metagenome-assembled viral genomes highlight a need for more extensive sampling of Arctic soils to better understand viral functions and contributions to ecosystem-wide cycling processes in the Arctic.

54 ENVIRONMENTAL SCIENCES

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive full-length, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

Loving, Rebekah K. (ORCID:0000000187250376)

Biophysical and biochemical evidence for the role of acetate kinases (AckAs) in an acetogenic pathway in pathogenic spirochetes

Unraveling the metabolism of Treponema pallidum is a key component to understanding the pathogenesis of the human disease that it causes, syphilis. For decades, it was assumed that glucose was the sole carbon/energy source for this parasitic spirochete. But the lack of citric-acid-cycle enzymes suggested that alternative sources could be utilized, especially in microaerophilic host environments where glycolysis should not be robust. Recent bioinformatic, biophysical, and biochemical evidence supports the existence of an acetogenic energy-conservation pathway in T . pallidum and related treponemal species. In this hypothetical pathway, exogenous D-lactate can be utilized by the bacterium as an alternative energy source. Herein, we examined the final enzyme in this pathway, acetate kinase (named TP0476), which ostensibly catalyzes the generation of ATP from ADP and acetyl-phosphate. We found that TP0476 was able to carry out this reaction, but the protein was not suitable for biophysical and structural characterization. We thus performed additional studies on the homologous enzyme (75% amino-acid sequence identity) from the oral pathogen Treponema vincentii , TV0924. This protein also exhibited acetate kinase activity, and it was amenable to structural and biophysical studies. We established that the enzyme exists as a dimer in solution, and then determined its crystal structure at a resolution of 1.36 Å, showing that the protein has a similar fold to other known acetate kinases. Mutation of residues in the putative active site drastically altered its enzymatic activity. A second crystal structure of TV0924 in the presence of AMP (at 1.3 Å resolution) provided insight into the binding of one of the enzyme’s substrates. On balance, this evidence strongly supported the roles of TP0476 and TV0924 as acetate kinases, reinforcing the hypothesis of an acetogenic pathway in pathogenic treponemes.

Deka, Ranjit K.

Sequencing and analysis of 131 SARS-CoV-2 isolates in previously sampled and unsampled regions of Jordan from 2020 to 2023

The Hashemite Kingdom of Jordan remains an understudied country for next generation sequencing analysis of SARS-CoV-2 genomes collected during the 2019 pandemic. Here we provide 131 additional reference genomes collected between 2020–2023 from SARS-CoV-2-positive patients across Jordan. Phylogenetic analysis supports existing pandemic narratives of changing clade dominance over time and adds genomes in novel Jordanian locations and timepoints to make Jordan SARS-CoV-2 databases more comprehensive. Samples from the less-sequenced cities of Ajloun, Jaresh, Karak, and Madaba identified previously unreported lineages while Amman, Irbid, and Zarqa have existing sequencing efforts bolstered. Despite many incomplete patient records and a relatively small sample size, we observe interesting symptom patterns that support existing global and Jordanian pandemic narratives. We note how in-country COVID-19 pandemic genomic studies showcase Jordan’s efforts to expand next generation sequencing capabilities, especially through the leveraging of EDGE COVID-19, a bioinformatics platform for performing rapid, batched analysis of SARS-CoV-2 sequencing that streamlines sample processing prepared from a network of hospital locations.

60 APPLIED LIFE SCIENCES

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES

Discovering Innovations in Stress Tolerance through Comparative Gene Regulatory Network Analysis and Cell-Type Specific Expression Maps (Final Technical Report with Cover Page)

Through this grant, we developed a comparative framework to elucidate the mechanisms behind variations in environmental stress responses among a diverse group of species within the Brassicaceae family. Our focus was on the differences in physiological and transcriptomic responses to abscisic acid (ABA), a hormone associated with water stress. We examined the differential growth responses of four Brassicaceae species, finding that most exhibited reduced root growth correlated with smaller meristem size. In contrast, Schrenkiella parvula showed accelerated growth due to increased root cell elongation. We employed RNA sequencing to analyze the transcriptional responses to ABA across these species, and innovative bioinformatics techniques were used to pinpoint biological pathways with significant divergence. Additionally, we utilized DAP-seq to map the gene regulatory networks associated with ABAresponsive transcription factors, revealing that variations in the regulation of growth hormone biosynthesis play a critical role in the distinct ABA effects on root growth among the species. This research sets a new standard for comparative physiology by integrating comparative genomics and transcriptomics to uncover pathway divergences.

59 BASIC BIOLOGICAL SCIENCES

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES

Creation of an Acyltransferase Toolbox for Plant Biomass Engineering (Final Report)

The major goal of this project was to expand our understanding of acyl‐CoA ligases and BAHD acyltransferases and their utility in plant engineering. We combined bioinformatic analysis of genes and transcripts with functional fingerprinting of synthesized genes produced by JGI. Best candidates from this experimental pipeline were transferred into bioenergy plants to study their effects on lignin composition. We found combinations of ligase and transferase genes encoding enzymes with interesting catalytic specificities. Our work demonstrated the feasibility of use of acyl-CoA ligases and BAHD acyltransferases to alter the composition of plant cell walls without deleterious effects on the modified plant.

59 BASIC BIOLOGICAL SCIENCES

Comparative Analysis of DNA LLM Classification Techniques Using Intra-Layer Feature Extraction with Autoencoder Stacks [Poster]

This project conducts a comparative analysis of DNA LLM classification techniques using Evo2, Grover, and UTRML, focusing on intra-layer feature extraction in Evo2. By extracting features from multiple layers of Evo2 and integrating them into an autoencoder stack with a binary classification head, we evaluate its effectiveness in classifying genomic sequences compared to smaller DNA language models. My findings demonstrate that Evo2 outperforms Grover and UTRML in classification accuracy on a dataset provided by department 08625, CAO2021, while UTRML offers competitive performance with lower computational costs. This study highlights the potential of advanced embedding techniques in enhancing genomic data analysis and informs future research in bioinformatics.

59 BASIC BIOLOGICAL SCIENCES

2024 International Conference on Microbiome Engineering (ICME)

The 2024 International Conference on Microbiome Engineering (ICME) took place November 12-14 at Tufts University in Medford, MA. ICME connects experts from academia and industry to share the most recent developments in the field of microbiome engineering. This includes genetically engineered organisms that function within microbiomes, control of microbiomes through environmental/nutrient modifications, and inference of engineering principles from analysis of synthetic and natural microbiomes. The conference is unique and distinct from other microbiome conferences in that it specifically highlights the integration of engineering design principles with microbiome research (others are more focused on basic biological principles). The conference thus integrates synthetic biology, systems biology, microbial ecology, and bioinformatics across a range of application spaces from the environment to manufacturing, food, and human health. This project utilized support from the Department of Energy’s (DOE) Office of Biological and Environmental Research (BER) to help trainees and early career faculty attend ICME.

60 APPLIED LIFE SCIENCES

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES