Search NASASearch

SEARCH · Search NASA

Results for “bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Structural Insights into the Mechanism of a Polyketide Synthase Thiocysteine Lyase Domain

Polyketide synthases (PKSs) are renowned for the structural diversity of the polyketide natural products they produce, but sulfur-containing functionalities are rarely installed by PKSs. We previously characterized thiocysteine lyase (SH) domains involved in the biosynthesis of the leinamycin (LNM) family of natural products, exemplified by LnmJ-SH and guangnanmycin (GnmT-SH). Here we report a detailed investigation into the PLP-dependent reaction catalyzed by the SH domains, guided by a 1.8 Å resolution crystal structure of GnmT-SH. A series of elaborate substrate mimics were synthesized to answer specific questions garnered from the crystal structure and from the biosynthetic logic of the LNM family of natural products. Here, through a combination of bioinformatics, molecular modeling, in vitro assays, and mutagenesis, we have developed a detailed model of acyl carrier protein (ACP)-tethered substrate-SH, and interdomain interactions, that contribute to the observed substrate specificity. Comparison of the GnmT-SH structure with archetypical PLP-dependent enzyme structures revealed how Nature, via evolution, has modified a common protein structural motif to accommodate an ACP-tethered substrate, which is significantly larger than any of those previously characterized. Overall, this study demonstrates how PLP-dependent chemistry can be incorporated into the context of PKS assembly lines and sets the stage for engineering PKSs to produce sulfur-containing polyketides.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools

Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses, and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5213 previously published case reports using the Phenopacket Schema, the Human Phenotype Ontology and Mondo disease ontology. Prompts generated from each phenopacket were sent to seven LLMs, including four generalist models and three LLMs specialized for medical applications. The same phenopackets were used as input to a widely used diagnostic tool, Exomiser, in phenotype-only mode. The best LLM ranked the correct diagnosis first in 23.6% of cases, whereas Exomiser did so in 35.5% of cases. While the performance of LLMs for supporting differential diagnosis has been improving, it has not reached the level of commonly used traditional bioinformatics tools. Future research is needed to determine the best approach to incorporate LLMs into diagnostic pipelines.

Reese, Justin T. [Lawrence Berkeley National Labor

Microbial species and intraspecies units exist and are maintained by ecological cohesiveness coupled to high homologous recombination

Abstract Recent genomic analyses have revealed that microbial communities are predominantly composed of persistent, sequence-discrete species and intraspecies units (genomovars), but the mechanisms that create and maintain these units remain unclear. By analyzing closely-related isolate genomes from the same or related samples and identifying recent recombination events using a novel bioinformatics methodology, we show that high ecological cohesiveness coupled to frequent-enough and unbiased (i.e., not selection-driven) horizontal gene flow, mediated by homologous recombination, often underlie these diversity patterns. Ecological cohesiveness was inferred based on greater similarity in temporal abundance patterns of genomes of the same vs. different units, and recombination was shown to affect all sizable segments of the genome (i.e., be genome-wide) and have two times or greater impact on sequence evolution than point mutations. These results were observed in bothSalinibacter ruber, an environmental halophilic organism, andEscherichia coli, the model gut-associated organism and an opportunistic pathogen, indicating that they may be more broadly applicable to the microbial world. Therefore, our results represent a departure compared to previous models of microbial speciation that invoke either ecology or recombination, but not necessarily their synergistic effect, and answer an important question for microbiology: what a species and a subspecies are.

Science & Technology - Other Topics

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES

Characterization of a widespread sugar phosphate-processing bacterial microcompartment

Many prokaryotes form Bacterial Microcompartments (BMCs) that encapsulate segments of specialized metabolic pathways to enhance catalysis. The various functions of metabolosomes, catabolic BMCs, are dictated by the signature enzyme that processes initial substrates of the confined pathway. The components and native functions of several metabolosomes have been experimentally characterized; however one of the most prevalent across all bacteria has yet to be studied. Sugar Phosphate Utilizing (SPU) BMC loci encode enzymes predicted to be involved in sugar phosphate metabolism. The SPU genetic loci are found in organisms occupying habitats ranging from soils to hot springs, highlighting the ubiquity of the SPU BMC. We bioinformatically characterized seven SPU subtypes, all which contain an enzyme unique to SPU BMCs, a deoxyribose 5-phosphate aldolase (DERA). Here, we define the fundamental characteristics of SPU BMCs and have expressed, purified, and characterized a set of SPU core enzymes. These include a protein-protein complex formed between a SPU BMC DERA and a predicted ribose 5-phosphate isomerase. Further, we show that the SPU BMC DERA is catalytically active and propose that it acts as the universal signature enzyme for the SPU BMC, with implications for fundamental understanding and biotechnological applications of SPU BMCs.

59 BASIC BIOLOGICAL SCIENCES

The influence of protein electrostatics on potential inversion in flavoproteins

Biology uses relatively few electron-transfer cofactors, tuning their potentials, electronic couplings, and reorganization energies to carry out the required chemistry. It is remarkable that the potential ordering of two-electron transfer active flavins can be normal (first oxidation at low potential and second oxidation at high potential) or inverted, and the gap between the potentials can be as large as one volt. Analysis based on structural bioinformatics and electrostatics indicates that the ordering of the flavin redox potential is influenced by protein electrostatics. In all 36 flavoproteins examined, the introduction of a negative charge near the flavin in silico increases the extent of potential inversion (by lowering the electrochemical potential of the second electron-transfer step); the introduction of a positive charge near the flavin favors normally ordered potentials. We also find that the addition of positive charges increases the electrochemical potential for the naturally occurring one-electron transition in flavodoxins (between deprotonated hydroquinone and neutral semiquinone) and also increases the second one-electron transition in bifurcating flavins (between anionic semiquinone and fully oxidized flavin). Finally, we find that proximity of a proton acceptor, notably conserved arginine, supports proton-coupled electron transfer because it may act as a proton acceptor, promoting potential inversion. This key arginine residue may enable two-electron transfer chemistry by promoting the proton-coupled electron transfer process over the pure electron transfer process, suggesting how a protein's flavin environment may influence one- or two-electron chemistry in flavoproteins.

Singh, Niven [Duke Univ., Durham, NC (United State

Electron transfer in polysaccharide monooxygenase catalysis

Polysaccharide monooxygenase (PMO) catalysis involves the chemically difficult hydroxylation of unactivated C–H bonds in carbohydrates. The reaction requires reducing equivalents and will utilize either oxygen or hydrogen peroxide as a cosubstrate. Two key mechanistic questions are addressed here: 1) How does the enzyme regulate the timely and tightly controlled electron delivery to the mononuclear copper active site, especially when bound substrate occludes the active site? and 2) How does this electron delivery differ when utilizing oxygen or hydrogen peroxide as a cosubstrate? Using a computational approach, potential paths of electron transfer (ET) to the active site copper ion were identified in a representative AA9 family PMO from Myceliophthora thermophila ( Mt PMO9E). When Y62, a buried residue 12 Å from the active site, is mutated to F, lower activity is observed with O 2 . However, a WT-level activity is observed with H 2 O 2 as a cosubstrate indicating an important role in ET for O 2 activation. To better understand the structural effects of mutations to Y62 and axial copper ligand Y168, crystal structures were solved of the wild type Mt PMO9E and the variants Y62W, Y62F, and Y168F. A bioinformatic analysis revealed that position 62 is conserved as either Y or W in the AA9 family. The Mt PMO9E Y62W variant has restored activity with O 2 . Overall, the use of redox-active residues to supply electrons for the reaction with O 2 appears to be widespread in the AA9 family. Furthermore, the results provide a molecular framework to understand catalysis with O 2 versus H 2 O 2 .

Sayler, Richard I. (ORCID:000000017252707X)

Electrochemical cofactor recycling of bacterial microcompartments

Bacterial microcompartments (BMCs) are prokaryotic organelles that consist of a protein shell which sequesters metabolic reactions in its interior. While most of the substrates and products are relatively small and can permeate the shell, many of the encapsulated enzymes require cofactors that must be regenerated inside. We have analyzed the occurrence of an enzyme previously assigned as a cobalamin (vitamin B 12 ) reductase and, curiously, found it in many unrelated BMC types that do not employ B 12 cofactors. We propose Nicotinamide adenine dinucleotide (NAD+) regeneration as the function of this enzyme and name it Metabolosome Nicotinamide Adenine Dinucleotide Hydrogen (NADH) dehydrogenase (MNdh). Its partner shell protein BMC-T SE (tandem domain BMC shell protein of the single layer type for electron transfer) assists in passing the generated electrons to the outside. We support this hypothesis with bioinformatic analysis, functional assays, Electron Paramagnetic Resonance spectroscopy, protein voltammetry, and structural modeling verified with X-ray footprinting. This finding represents a paradigm for the BMC field, identifying a new, widely occurring route for cofactor recycling and a new function for the shell as separating redox environments.

bacterial microcompartment

Leaky ribosomal scanning enables tunable translation of bicistronic ORFs in green algae

Advances in sequencing technology have unveiled examples of nucleus-encoded polycistrons, once considered rare. Exclusively polycistronic transcripts are prevalent in green algae, although the mechanism by which multiple polypeptides are translated from a single transcript is unknown. Here, we used bioinformatic and in vivo mutational analyses to evaluate competing mechanistic models for translation of bicistronic mRNAs in green algae. High-confidence manually curated datasets of bicistronic loci from two divergent green algae, Chlamydomonas reinhardtii and Auxenochlorella protothecoides, revealed a preference for weak Kozak-like sequences for ORF 1 and an underrepresentation of potential initiation codons before the ORF 2 start codon, which are suitable conditions for leaky ribosome scanning to allow ORF 2 translation. We used mutational analysis in A. protothecoides to test the mechanism. In vivo manipulation of the ORF 1 Kozak-like sequence and start codon altered reporter expression at ORF 2, with a weaker Kozak-like sequence enhancing expression and a stronger one diminishing it. A synthetic bicistronic dual reporter demonstrated inversely adjustable activity of green fluorescent protein expressed from ORF 1 and luciferase from ORF 2, depending on the strength of the ORF 1 Kozak-like sequence. Our findings demonstrate that translation of multiple ORFs in green algal bicistronic transcripts is consistent with episodic leaky scanning of ORF 1 to allow translation at ORF 2. This work has implications for the potential functionality of upstream open reading frames (uORFs) found across eukaryotic genomes and for transgene expression in synthetic biology applications.

59 BASIC BIOLOGICAL SCIENCES

Pyrodictium abyssi AbpX reveals a calcium-responsive family of microbial biomatrix proteins that form thermostable hydrogels

Evolutionary pressure on microbial communities propagating under extreme environmental conditions often results in unique structural adaptations to promote cell survival. In this work, we report an investigation of AbpX, a biomatrix protein identified in cultures of the hyperthermophilic archaeon Pyrodictium abyssi. Under ex vivo and in vitro conditions, AbpX assembles into a paracrystalline lattice composed of semiflexible fibrils. CryoEM analysis of recombinant AbpX fibrils reveals that the precursor protein polymerizes through donor strand complementation (DSC), a process previously reported for chaperone-usher fimbriae in Gram-negative bacteria. Unlike the latter DSC protein polymers, AbpX undergoes chaperone-free polymerization in the presence of calcium ions, which are sequestered at the donor strand-acceptor groove interface between protomers in the fibril. Using a combination of cryoEM and crystallographic information, a structural model is proposed for the AbpX lattice that provides insight into its potential role in biofilm formation. These findings suggest that calcium ion coordination may contribute to fibril assembly and preorganize fibrils for incorporation into the protein lattice. Bioinformatic analysis indicates that AbpX exemplifies a distinct and broadly distributed clade of calcium ion responsive biomatrix proteins within the TasA superfamily that can be fabricated into hydrogel biomaterials in vitro under environmentally benign conditions.

59 BASIC BIOLOGICAL SCIENCES

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.

EEPD1 evolved a unique DNA clamping dimer protecting reversed replication forks

Exonuclease/endonuclease/phosphatase (EEP)-fold hydrolases are canonically monomeric phosphodiesterases exemplified by APE1, DNase I, and TDP2 nucleases. While EEP family domain containing protein 1 (EEPD1) acts in DNA stress responses, its proposed nuclease activities are enigmatic. Here, we integrate hybrid structural methods, evolution, biochemistry, cancer genomics, plus molecular and cell biology to define EEPD1 structure, assembly, and function at stalled DNA replication forks. Results imply EEPD1 surprisingly requires both unique EEP domain dimer and distinctive tandem Helix-hairpin-Helix [(HhH) 2 ] domains to clamp double-stranded (ds) DNA at reversed DNA replication forks for fork protection. Small-angle X-ray Scattering (SAXS), crystal, and cryo-EM structures unveil an unprecedented tryptophan handshake dimer, conserved interface di-Trp-Pro pocket, and adjustable “wrist” enabling an open-closed conformational switch. EEPD1 dimer cooperatively binds complex dsDNA replication fork intermediates but alone lacks nuclease activity due to loss of key EEP catalytic residues during Metazoan evolution and atmospheric oxygen buildup. Instead, EEPD1 prevents nucleolytic degradation of reversed replication forks by MRE11. Furthermore, cancer bioinformatics support oxidative damage-dependent EEPD1 association as a significant modulator of overall patient survival. Collective findings uncover unexpected EEP dimer and fork protection function in clamping, not cleaving, reversed replication forks for metazoan oxidative stress responses controlling genome stability and cancer outcomes.

Shen, Runze [Univ. of Texas, Houston, TX (United S

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES

Engineering quantitative stomatal trait variation and local adaptation potential by cis‐regulatory editing

Summary Cis‐regulatory element editing can generate quantitative trait variation that mitigates extreme phenotypes and harmful pleiotropy associated with coding sequence mutations. Here, we applied a multiplexed CRISPR/Cas9 approach, informed by bioinformatic datasets, to generate genotypic variation in the promoter ofOsSTOMAGEN, a positive regulator of rice stomatal density. Engineered genotypic variation corresponded to broad and continuous variation in stomatal density, ranging from 70% to 120% of wild‐type stomatal density. This panel of stomatal variants was leveraged in physiological assays to establish discrete relationships between stomatal morphological variation and stomatal conductance, carbon assimilation and intrinsic water use efficiency in steady‐state and fluctuating light conditions. Additionally, promoter alleles were subjected to vegetative drought regimes to assay the effects of the edited alleles on developmental response to drought. Notably, the capacity for drought‐responsive stomatal density reprogramming instomagenand two cis‐regulatory edited alleles was reduced. Collectively our data demonstrate that cis‐regulatory element editing can generate near‐isogenic trait variation that can be leveraged for establishing relationships between anatomy and physiology, providing a basis for optimizing traits across diverse environments.

Biotechnology & Applied Microbiology