Search NASASearch

SEARCH · Search NASA

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Hydrazinoacetic acid is a biosynthetic precursor of the bacterially produced nitramine, N -nitroglycine

Nitramines [R(R′)N–NO 2 ; R,R′=H or alkyl] are valuable synthetic products, but knowledge of the biosynthetic processes that generate these compounds is limited. This work sought to elucidate the biosynthesis of a nitramine natural product, N-nitroglycine (NNG) by Streptomyces noursei . Stable isotope studies showed that S. noursei cells supplemented with L-(ε- 15 N)lysine, ( 15 N)glycine, or ( 13 C)hydrazinoacetic acid (HAA) incorporated 67%, 88%, and 67% of the isotope label into NNG, respectively, indicating that these compounds are biosynthetic precursors of NNG. Liquid chromatography coupled tandem mass spectrometry (LC-MS/MS) of 15 N-Lys-labeled NNG confirmed that the nitro nitrogen of NNG originates from Lys. Bioinformatics analysis of the S. noursei genome showed evidence for a biosynthetic gene cluster (BGC) that contained machinery for HAA biosynthesis ( nngKLM ), consistent with the results of the isotope labeling. In vitro reconstitution of the gene products produced HAA. The borders of this BGC were defined by cross-referencing the predicted BGC with previously published differential proteomics data. Furthermore, we show that azaserine is produced alongside NNG in S. noursei cultures, linking the two biosynthetic pathways via a proposed nitrosamine biosynthetic intermediate. Finally, the oxygen balance for NNG is −20.2% for the formation of carbon dioxide (CO 2 ), which is comparable to that of hexahydro-1,3,5- trinitro-1,3,5-triazene (common name: RDX; −21.6%). Crystal structure data of NNG indicate that the unit crystalizes as a pure material, not a hydrate, suggesting a favorable energetic crystallization phase. The combined results suggest a route that, with further development, could lead to sustainable production of energetic nitramines via synthetic biology or biocatalytic approaches.

biosynthesis

Cryptic cycling by electroactive bacterioplankton in Trout Bog Lake

The potential for extracellular electron transfer (EET) is a prevailing genomic feature of humic lake bacterioplankton. However, there has been little evidence for the substantial ecological contribution predicted by genetics. We hypothesized that anoxygenic phototrophic electrotrophs and accompanying heterotrophic electrogens cycle dissolved organic matter (DOM) between oxidized and reduced states. We predicted that such bacterioplankton would exhibit diel-scale oscillations due to the light dependency of photosynthesis. Using Trout Bog Lake in Wisconsin, USA, as our model ecosystem, we profiled the water column with depth-discrete metagenomic, physiochemical, and electrochemical analyses. We observed variation in oxidation reduction potential (ORP) in response to sunlight, initiating at depths populated by anoxygenic phototrophs with EET genes. We developed an automated buoy to measure electric current flow between many pairs of electrodes simultaneously, observing correlation in electron consumption to sunlight. Our results, combined with published metatranscriptomic analysis, indicate the occurrence of electron cycling between phototrophic oxidation (electrotrophic metabolism) by Chlorobium and anaerobic respiration (electrogenic metabolism) by Geothrix, involving DOM. We also repeatedly observed gradual seasonal increases in hypolimnion ORP throughout summer. These diel and seasonal patterns imply that electroactive DOM mediates the ecology of electroactive bacteria in lakes, controlling humic lake methane emissions.IMPORTANCEWe investigated the physical, chemical, and redox characteristics of a bog lake and electrodes hung therein to test the hypothesis that dissolved organic matter is being cycled between oxidized and reduced states by electroactive bacterioplankton powered by phototrophy. To do so, we performed field-based analyses on multiple timescales using both established and novel instrumentation. We paired these analyses with recently developed bioinformatics pipelines for metagenomics data to investigate genes that enable electroactive metabolism and accompanying metabolisms. Our results are consistent with our hypothesis and yet upend some of our other expectations. Our findings have implications for understanding greenhouse gas emissions from lakes, including electroactivity as an integral part of lake metabolism throughout more of the anoxic parts of lakes and for a longer portion of the summer than expected. Our results also give a sense of what electroactivity occurs at given depths and provide a strong basis for future studies.

carbon emissions

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES

Bacteria export alarmone synthetases that produce (p)ppApp and (p)ppGpp

Guanosine penta- and tetraphosphate [(p)ppGpp] and their adenosine analogs [(p)ppApp] are bacterial second messengers known as alarmones. Members of the RelA-SpoT homolog (RSH) family synthesize (p)ppGpp to mediate the stringent response during nutrient starvation, whereas (p)ppApp synthetases have been identified as bactericidal toxins in diverse contexts including type VI secretion systems, toxin-antitoxin modules, and phages. Although alarmone synthesis has traditionally been viewed as a cytoplasmic process, early studies in Streptomyces suggested the existence of secreted alarmone synthetases. Here, we identify SaEAS, an exported alarmone synthetase (EAS) from Streptomyces albidoflavus, as the long-mysterious source of extracellular alarmone synthetase activity in Streptomyces. SaEAS produces both (p)ppGpp and (p)ppApp at rates exceeding 100,000 molecules per minute and has kinetic properties adapted to low substrate environments. A broader bioinformatic survey reveals ~600 EASs linked to a range of specialized bacterial secretion systems. Characterization of two additional EASs, VpEAS from Vibrio parahaemolyticus and AaEAS from Amycolatopsis azurea, shows that both produce (p)ppGpp exclusively and inhibit bacterial growth when localized to the cytoplasm. These findings challenge the longstanding view of (p)ppGpp as strictly pro-survival and unveil a diverse family of secreted RSH enzymes with potential roles in interbacterial antagonism and environmental signaling.

Ahmad, Shehryar

Exploring life’s hidden majority: microbial dark matter symposium highlights

The Microbial Dark Matter Symposium held on August 28–29, 2025, in Laguna Beach, Orange County, CA, convened a multidisciplinary group of scientists to address the vast unknowns in microbial life—from uncultured taxa and uncharacterized proteins to elusive viruses and spacefaring microbes. Set against a scenic coastal backdrop, the symposium highlighted advances in single-cell genomics, proximity ligation sequencing, and artificial intelligence-ready bioinformatics, while also probing the limits of microbial persistence, metabolism, and ecological distribution. Sessions explored microbial dark matter from multiple dimensions: cultivability, where new strategies are enabling recovery of elusive microbes; functional ambiguity, where metagenomic dark zones are illuminated by computational annotation; and genomic representation, where single-cell methods bridge gaps left by shotgun community sequencing. Researchers shared breakthroughs in identifying atmospheric microbiomes, “dark oxygen” production in groundwater ecosystems, and microbial survival on the International Space Station. The symposium emphasized integration of methods, disciplines, and ecosystems, advancing a collective push to illuminate the microbial dark matter on Earth and beyond. By highlighting emerging tools, pressing questions, and cross-domain insights, the symposium underscored the need for collaborative, open, and adaptive approaches to study the microbial unknown. The meeting marks a pivotal moment in microbiology, where cultivating knowledge of the uncultivated promises transformative understanding of life, everywhere.

Podar, Mircea [ORNL] (ORCID:0000000327760205)

Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses

ABSTRACT Although long-read sequencing has enabled obtaining high-quality and complete genomes from metagenomes, many challenges still remain to completely decompose a metagenome into its constituent prokaryotic and viral genomes. This study focuses on decomposing an estuarine metagenome to obtain a more accurate estimate of microbial diversity. To achieve this, we developed a new bead-based DNA extraction method, a novel bin refinement method, and obtained 150 Gbp of Nanopore sequencing. We estimate that there are ~500 bacterial and archaeal species in our sample and obtained 68 high-quality bins (>90% complete, <5% contamination, ≤5 contigs, contig length of >100 kbp, and all ribosomal and tRNA genes). We also obtained many contigs of picoeukaryotes, environmental DNA of larger eukaryotes such as mammals, and complete mitochondrial and chloroplast genomes and detected ~40,000 viral populations. Our analysis indicates that there are only a few strains that comprise most of the species abundances. IMPORTANCE Ocean and estuarine microbiomes play critical roles in global element cycling and ecosystem function. Despite the importance of these microbial communities, many species still have not been cultured in the lab. Environmental sequencing is the primary way the function and population dynamics of these communities can be studied. Long-read sequencing provides an avenue to overcome limitations of short-read technologies to obtain complete microbial genomes but comes with its own technical challenges, such as needed sequencing depth and obtaining high-quality DNA. We present here new sampling and bioinformatics methods to attempt decomposing an estuarine microbiome into its constituent genomes. Our results suggest there are only a few strains that comprise most of the species abundances from viruses to picoeukaryotes, and to fully decompose a metagenome of this diversity requires 1 Tbp of long-read sequencing. We anticipate that as long-read sequencing technologies continue to improve, less sequencing will be needed.

Lui, Lauren M.

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan

FAIR Ecosystems for Science at Scale

High Performance Computing (HPC) centers provide resources to users who require greater scale to “get science done”. They deploy infrastructure with singular hardware architectures, cutting-edge software environments, and stricter security measures as compared with users’ own resources. As a result, users often create and configure digital artifacts in ways that are specialized for the unique infrastructure at a given HPC center. Each user of that center will face similar challenges as they develop specialized solutions to take full advantages of the center’s resources, potentially resulting in significant duplication of effort. Much duplicated effort could be avoided, however, if users of these centers found it easier to discover others’ solutions and artifacts as well as share their own. The FAIR principles address this problem by presenting guidelines focused around metadata practices to be implemented by vaguely defined “communities”; in practice, these tend to gather by domain (e.g. bioinformatics, geosciences, agriculture). Domain-based communities can unfortunately end up functioning as silos that tend both to inhibit sharing of solutions and best practices as well as to encourage fragile and unsustainable improvised solutions in the absence of best-practice guidance. We propose that these communities pursuing “science at scale” be nurtured both individually and collectively by HPC centers so that users can take advantage of shared challenges across disciplines and potentially across HPC centers. We describe an architecture based on the EOSC-Life FAIR Workflows Collaboratory, specialized for use with and inside HPC centers such as the Oak Ridge Leadership Computing Facility (OLCF), and we speculate on user incentives to encourage adoption. We note that a focus on FAIR workflow components rather than FAIR workflows is more likely to benefit the users of HPC centers.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector

ATCCfinder - Download and Search the ATCC Genome Portal

Much strain-specific sequence data exists in research conducted before the deployment of large sequencing repositories, making it challenging to identify and validate the identity of strains used in these studies through bioinformatics and phenotyping. The American Type Culture Collection (ATCC) is an organization that sells a wide variety of microbes with strain-level taxonomy classification and associated sequenced reference genomes. Currently, ATCC does not provide a method for searching for sequence similarity between a query sequence and their database of reference genomes. Here I propose the software ATCCfinder, which utilizes ATCC application interface software (API) to generate query-able databases from ATCC Genome resources.

Koehler, Samuel

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora

Dori Tools v2.1.0

Dori Tools is a set of useful programs that support running a mid-range HPC cluster for BioInformatics workloads.

Rath, Georg [Lawrence Berkeley National Laboratory

wastewater_virus

This repo contains software used to clean and assemble high-throughput sequencing data containing viruses. The input is raw illumina sequencing reads and the output is a database of high-quality viral genomes. The specific application is to wastewater viral concentrates but it is not restricted to that sample type. The software is composed of Nextflow workflows and a set of custom Python and bash scripts that call publicly available bioinformatics tools to accomplish obvious tasks in data analysis in a high performance computing environment. For detailed information, please see the repo's README file.

Kantor, Rose [Lawrence Livermore National Laborato

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES