Search NASA⌕ Search

SEARCH · Search NASA

Results for “annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

A chromosome-level genome assembly of the varied leaved jewelflower, Streptanthus diversifolius, reveals a recent whole genome duplication

Abstract The Streptanthoid complex, a clade of primarily Streptanthus and Caulanthus species in the Thelypodieae (Brassicaceae) is an emerging model system for ecological and evolutionary studies. This complex spans the full range of the California Floristic Province including desert, foothill, and mountain environments. The ability of these related species to radiate into dramatically different environments makes them a desirable study subject for exploring how plant species expand their ranges and adapt to new environments over time. Ecological and evolutionary studies for this complex have revealed fascinating variation in serpentine soil adaptation, defense compounds, germination, flowering, and life history strategies. Until now a lack of publicly available genome assemblies has hindered the ability to relate these phenotypic observations to their underlying genetic and molecular mechanisms. To help remedy this situation, we present here a chromosome-level genome assembly and annotation of Streptanthus diversifolius, a member of the Streptanthoid Complex, developed using Illumina, Hi-C, and HiFi sequencing technologies. Construction of this assembly also provides further evidence to support the previously reported recent whole genome duplication unique to the Thelypodieae. This whole genome duplication may have provided individuals in the Streptanthoid Complex the genetic arsenal to rapidly radiate throughout the California Floristic Province and to occupy commonly inhospitable environments including serpentine soils.

Genetics & Heredity↗

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes↗

A single genomic region controls primocane fruiting in tetraploid blackberry

The fresh-market blackberry ( Rubus subgenus Rubus ) industry has expanded dramatically in the past 2 decades, driven in part by improved cultivars. Introgression of the primocane-fruiting (PF; annual flowering) trait into elite germplasm has enabled dual cropping in a single year, season extension, and cultivation in tropical and subtropical regions. Despite its economic performance, the genetic basis of PF is not well understood. It has been proposed that the PF trait is controlled by a major recessive locus, but its genomic location is unclear. Here, a genome-wide association study (GWAS) of 365 tetraploid blackberry genotypes identified a single genomic region on chromosome Ra03 (∼33 Mb) strongly associated with PF. Genetic linkage analysis in a biparental population confirmed that the same interval (32–35 Mb) was linked to the PF phenotype. Ten putative candidate genes were identified in this region. Allele mining using whole-genome resequencing of 17 genotypes highlighted 2 high-priority candidates: a CCCH-type zinc finger gene and an ubiquitin-specific protease gene. Use of an improved Rubus argutus “Hillquist” genome annotation (v1.2) enabled refined variant interpretation, including identification of regulatory 3′ UTR polymorphisms in the zinc finger homolog. Two diagnostic KASP markers (PF1 and PF2), designed from the most significant GWAS SNPs, predicted the PF phenotype with over 96% accuracy in a validation panel of 494 tetraploid blackberries from multiple breeding programs. Together, these results provide the first high-resolution mapping of the PF locus in blackberry, identify candidate genes for flowering regulation in Rubus , and deliver diagnostic markers that can be immediately deployed in breeding programs.

GWAS↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Actinomycetota isolated from the sponge Hymeniacidon perlevis as a source of novel compounds with pharmacological applications: diversity, bioactivity screening, and metabolomic analysis

Abstract Aims To combat health conditions, such as multi-resistant bacterial infections, cancer, and metabolic diseases, new drugs need to be urgently found and, in this respect, marine Actinomycetota have a high potential to produce secondary metabolites with pharmacological importance. We aimed to study the cultivable Actinomycetota community associated with a marine sponge from the Portuguese coast, Hymeniacidon perlevis, and investigate the potential of the retrieved isolates to produce compounds with antimicrobial, anticancer and anti-obesity properties. Methods and results The analysis of the 16S rRNA gene revealed 79 Actinomycetota isolates affiliated with 12 genera—Brachybacterium, Dietzia, Glutamicibacter, Gordonia, Micrococcus, Micromonospora, Nocardia, Nocardiopsis, Paenoartrhobacter, Rhodococcus, Streptomyces, and Tsukamurella, most of which affiliated with the genus Streptomyces. The screening of antimicrobial activity revealed 13 strains, all belonging to the Streptomyces genus, capable of inhibiting the growth of Candida albicans, Bacillus subtilis, or Staphylococcus aureus. Forty-three extracts exhibited cytotoxic activity against at least one tested cell line (HepG2, HCT-116, and hCMEC-D3). Three extracts that were active against the two cancer cell lines tested, did not reduce the viability of the non-cancer endothelial cell line, hCMEC-D3. One Gordonia strain exhibited anti-obesity activity, revealed by its ability to reduce the neutral lipids in zebrafish larvae. Mass spectrometry-based dereplication analysis of active extracts identified several compounds associated with known Actinomycetota natural products. Nonetheless, five clusters contained metabolites that did not match any annotated natural products, suggesting they may represent new bioactive molecules. Conclusions This work contributed to increase the knowledge on the diversity and bioactive potential of Actinomycetota associated with H. perlevis.

Fonseca, Ana C.↗

Specialization Restricts the Evolutionary Paths Available to Yeast Sugar Transporters

Functional innovation at the protein level is a key source of evolutionary novelties. The constraints on functional innovations are likely to be highly specific in different proteins, which are shaped by their unique histories and the extent of global epistasis that arises from their structures and biochemistries. These contextual nuances in the sequence–function relationship have implications both for a basic understanding of the evolutionary process and for engineering proteins with desirable properties. Here, we have investigated the molecular basis of novel function in a model member of an ancient, conserved, and biotechnologically relevant protein family. These Major Facilitator Superfamily sugar porters are a functionally diverse group of proteins that are thought to be highly plastic and evolvable. By dissecting a recent evolutionary innovation in an α-glucoside transporter from the yeast Saccharomyces eubayanus, we show that the ability to transport a novel substrate requires high-order interactions between many protein regions and numerous specific residues proximal to the transport channel. To reconcile the functional diversity of this family with the constrained evolution of this model protein, we generated new, state-of-the-art genome annotations for 332 Saccharomycotina yeast species spanning ~400 My of evolution. By integrating phylogenetic and phenotypic analyses across these species, we show that the model yeast α-glucoside transporters likely evolved from a multifunctional ancestor and became subfunctionalized. The accumulation of additive and epistatic substitutions likely entrenched this subfunction, which made the simultaneous acquisition of multiple interacting substitutions the only reasonably accessible path to novelty.

59 BASIC BIOLOGICAL SCIENCES↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

BindingDB in 2024: a FAIR knowledgebase of protein-small molecule binding data

Abstract BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB’s positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset and plans for future growth and development.

Liu, Tiqing↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

BioPortal: an open community resource for sharing, searching, and utilizing biomedical ontologies

Abstract BioPortal (https://bioportal.bioontology.org) is the world’s most comprehensive repository of biomedical ontologies. It provides infrastructure for finding, sharing, searching, and utilizing biomedical ontologies. Launched in 2005, BioPortal now includes 1549 ontologies (1182 of them public). Its open, freely accessible website enables anyone (i) to browse the ontology library, (ii) to search for terms across ontologies, (iii) to browse mappings between terms, (iv) to see popularity ratings and recommendations on which ontologies are most relevant to their use cases, (v) to annotate text with ontology terms, (vi) to submit an ontology, and (vii) to request ontology changes. The library of ontologies can be accessed programmatically via a REST application programming interface (API). Recent enhancements include a BioPortal knowledge graph that integrates knowledge from multiple ontologies; a unified data model for interoperability with other knowledge sources; ontology popularity ratings and recommendations for relevant ontologies; and the ability to request ontology changes via a simple user interface that automatically converts user change requests to GitHub Pull Requests that specify the edits that will be made to the ontology upon approval.

Vendetti, Jennifer↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

Targeted genetic manipulation and yeast-like evolutionary genomics in the green alga Auxenochlorella

Auxenochlorella spp. are diploid oleaginous green algae whose streamlined genomes can be readily manipulated by homologous recombination, making them highly amenable to discovery research and bioengineering. Vegetatively diploid organisms experience specific evolutionary phenomena, including allodiploid hybridization, mitotic recombination, loss-of-heterozygosity, and aneuploidy; however, studies of these forces have largely focused on yeasts. Here, we present a telomere-to-telomere phased diploid genome assembly of Auxenochlorella UTEX 250-A (haploid length 22 Mb) and introduce a genetic toolkit for site-specific manipulation of the nuclear genome in multiple strains, featuring several selectable markers, inducible promoters, and fluorescent reporters for protein localization. UTEX 250-A is an allodiploid hybrid of Auxenochlorella protothecoides and Auxenochlorella symbiontica, two species differentiated by extensive chromosomal rearrangements. UTEX 250-A haplotypes are a mosaic of each parental species following mitotic recombination, and two chromosomes are trisomic. Loss-of-heterozygosity events are pervasive across Auxenochlorella and can evolve rapidly in the laboratory. High-quality structural annotation yielded ∼7,500 genes per haplotype. Auxenochlorella have experienced gene family loss and reduction, including core photosynthesis genes, and exhibit periodic adenine and cytosine methylation at promoters and gene bodies, respectively. Approximately 10% of genes, especially those involved in DNA repair and sex, overlap antisense long noncoding RNAs, which may participate in a regulatory mechanism. We demonstrate the utility of Auxenochlorella for fundamental research by knockout of a chlorophyll biosynthesis enzyme, and confirm one trisomy by allele-specific transformation. These results demonstrate the generality of several evolutionary forces associated with vegetative diploidy and provide a foundation for the use of Auxenochlorella as a reference organism.

CHL27↗

The small protein SbtC is a functional component of the CO 2 concentrating mechanism in Synechocystis sp. PCC 6803

Oxygenic phototrophs fix CO 2 via the enzyme ribulose-1,5-bisphosphate carboxylase/oxygenase (RubisCO), which shows relatively low CO 2 affinity and specificity. To circumvent low and fluctuating CO 2 concentrations in aquatic systems, cyanobacteria and algae have evolved sophisticated inorganic carbon (Ci) concentrating mechanisms (CCMs). Bicarbonate transporters such as SbtA play a crucial role in the cyanobacterial CCM and hence display multiple layers of tight regulation. Control of sbtA gene expression and corresponding transporter activity involves the PII-like protein SbtB, whose gene is frequently co-transcribed with sbtA. A previously non-annotated gene located upstream of the sbtAB operon in the model Synechocystis sp. PCC 6803 encodes the small protein SbtC, composed of 80 amino acids. Presence of SbtC was confirmed by immunoblotting of the sbtC-coding sequence fused to a Flag-tag. Similar to sbtAB , transcription of the sbtC locus is induced by low CO 2 availability; however, it is controlled independently. Mutation of the sbtC locus in a wild-type background produced only a mild phenotype, even under low CO 2 , but impaired diurnal growth resembled that of the mutant ΔsbtB . Biochemical analysis indicated a trimeric SbtABC complex in the membrane. Bicarbonate leakage from cells was strongly elevated when either sbtB or sbtC was deleted from recombinant Synechocystis strains harboring only SbtA as single Ci uptake system. Here, our results provide evidence that SbtC contributes to the formation of the SbtAB complex, thereby regulating bicarbonate exchange at the cytoplasmic membrane. Well-conserved SbtC-like proteins encoded in the neighborhood of sbtAB exist in many cyanobacterial genomes, pointing toward an important role in the cyanobacterial CCM.

Walke, Peter [Univ. of Rostock (Germany)] (ORCID:0↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding on the BRIAR Dataset: A Comprehensive Whole Body Biometric Recognition Resource at Extreme Distances and Real-World Scenarios (Collections 1-4)

The state-of-the-art in biometric recognition algorithms and operational systems has advanced quickly in recent years providing high accuracy and robustness in more challenging collection environments and consumer applications. However, the technology still suffers greatly when applied to non-conventional settings such as those seen when performing identification at extreme distances or from elevated cameras on buildings or mounted to UAVs. This paper summarizes an extension to the largest dataset currently focused on addressing these operational challenges, and describes its composition as well as methodologies of collection, curation, and annotation.

Cornett, David [ORNL] (ORCID:0000000222910860)↗

DOC-DICAM: Domain Aware One Class Defect Identification in Composite Aerostructure Material

Fiber-reinforced composites are a common material used in the design of aircraft structures due to their good tensile strength and resistance to compression. During the manufacturing process, these structures are thoroughly inspected for flaws and defects to ensure structural integrity during commercial use. Non-destructive testing (NDT) is a collection of inspection methods that allow inspectors to evaluate material without altering it. Due to the high safety standards in aerospace manufacturing, the NDT process is done manually and can be a significant bottleneck in the development workflow. In this paper, we develop an AI-based assistance tool to drastically reduce inspection time. Typical AI workflows require large amounts of annotated data, but defects rarely occur resulting in strong class imbalance. To overcome this, we formulate the problem of defect identification as an anomaly detection task in which our primary focus is learning non-defect characteristics. To do this, we develop a multi-task self-supervised learning framework that embeds problem specific domain knowledge into the deep learning model. We verify our method using fuselage data generated in a production environment. As a result, we show that our method can effectively identify defects and requires minimal training and inference time.

anomaly detection↗