Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES

wastewater_virus

This repo contains software used to clean and assemble high-throughput sequencing data containing viruses. The input is raw illumina sequencing reads and the output is a database of high-quality viral genomes. The specific application is to wastewater viral concentrates but it is not restricted to that sample type. The software is composed of Nextflow workflows and a set of custom Python and bash scripts that call publicly available bioinformatics tools to accomplish obvious tasks in data analysis in a high performance computing environment. For detailed information, please see the repo's README file.

Kantor, Rose [Lawrence Livermore National Laborato

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University

Transcriptomic data sets for Novosphingobium aromaticivorans DSM12444 and a ΔSARO_RS14285 mutant grown in the presence of glucose and either protocatechuic, vanillic, syringic, or 4-coumaric acid

The SARO_RS14285 gene, encoding a transcription factor, was deleted in Novosphingobium aromaticivorans DSM12444. The transcriptomes of the parent and ΔSARO_RS14285 strains were determined when grown in medium containing glucose with or without protocatechuic, vanillic, syringic, or 4-coumaric acid. We present the raw RNA sequencing data obtained from these cultures.

Novosphingobium aromaticivorans

Transcriptomic data sets for Novosphingobium aromaticivorans grown with the β-5-linked aromatic dimer dehydrodiconiferyl alcohol and the related G-aromatic monomers vanillin and ferulic acid

ABSTRACT The transcriptomes of a 2-pyrone-4,6-dicarboxylic acid-producing strain of Novosphingobium aromaticivorans DSM12444 were determined when grown in minimal medium containing glucose alone or glucose plus vanillin, ferulic acid, or the β-5-linked aromatic dimer dehydrodiconiferyl alcohol as carbon sources. Here, we present the RNA-sequencing data we obtained.

Metz, Fletcher

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean

Multi-omics data resource: Data package 24 (Pck024)

The data package consists of isolated pancreatic islets from 3 human donors treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and IL-1β, IFNγ, IL-1β + IFNγ, IL-1β + IFNγ + NMMA or NMMA for 18 h and submitted for scRNA-seq. This study examines cytokine-stimulated changes in gene expression in human islets using single-cell RNA sequencing. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE251730 Publication: 10.1093/function/zqae015

Sarkar, Soumyadeep [Pacific Northwest National Lab

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection

Exploring Saccharomycotina Yeast Ecology Through an Ecological Ontology Framework

Yeasts in the subphylum Saccharomycotina are found across the globe in disparate ecosystems. A major aim of yeast research is to understand the diversity and evolution of ecological traits, such as carbon metabolic breadth, insect association, and cactophily. This includes studying aspects of ecological traits like genetic architecture or association with other phenotypic traits. Genomic resources in the Saccharomycotina have grown rapidly. Ecological data, however, are still limited for many species, especially those only known from species descriptions where usually only a limited number of strains are studied. Moreover, ecological information is recorded in natural language format limiting high throughput computational analysis. To address these limitations, we developed an ontological framework for the analysis of yeast ecology. A total of 1,088 yeast strains were added to the Ontology of Yeast Environments (OYE) and analyzed in a machine-learning framework to connect genotype to ecology. This framework is flexible and can be extended to additional isolates, species, or environmental sequencing data. Widespread adoption of OYE would greatly aid the study of macroecology in the Saccharomycotina subphylum.

59 BASIC BIOLOGICAL SCIENCES

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks

Ecological connectivity and habitat loss shape patterns of genetic diversity in a threatened salamander

Context The maintenance of genetic diversity is essential for preserving adaptive potential in populations, yet it is increasingly threatened by landscape alteration. The field of landscape genetics offers a framework for assessing how patch-level landscape conditions, modeled at multiple scales, influence genetic diversity. Objectives We sought to assess how local environmental features and connectivity influence genetic diversity across 74 four-toed salamander (Hemidactylium scutatum) breeding wetlands in the southeastern United States. Methods Using next-generation sequencing data and hierarchical Bayesian models, we examined genome-wide heterozygosity in relation to local landscape features and ecological connectivity. We also assessed the scale of effect of landscape features and tested for temporal lag effects. Results Genetic diversity was lower in wetlands with higher levels of historic deforestation and lower connectivity. An interaction between deforestation and connectivity indicated that deforestation had stronger negative effects in isolated wetlands but weaker effects in well-connected wetlands. Accounting for scale of effect and temporal lags was critical for detecting these relationships. Conclusions Our analyses highlight the importance of assessing the spatial scale (scale of effect) and temporal lag of landscape features to detect key drivers of genetic diversity. In line with population genetic theory, our results indicate that the genetic consequences of habitat loss do not affect populations uniformly and are most severe in isolated populations where gene flow cannot buffer against loss of diversity. Altogether, we highlight the importance of considering the interaction of habitat loss and connectivity in conservation genetic management.

Hemidactylium scutatum

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag

Identification and characterization of mono- and bifunctional galactan synthases in the pediatric pathogen Kingella kingae

The emerging pediatric pathogen Kingella kingae elaborates a lipopolysaccharide (LPS) that is extended with a galactofuranose homopolymer called galactan, which is a key virulence determinant that contributes to resistance to complement-mediated and neutrophil-mediated killing. Previous work has demonstrated that the pamABCDE locus is required for galactan synthesis. In this study, mutational studies suggested that the pamC gene product is a UDP-galactofuranose (Galf) transferase and is the galactan synthase. Analysis of genome sequence data revealed two distinct pamC alleles designated pamC1 and pamC2, which correlate with the two galactan structures in K. kingae. Examination of isogenic mutants expressing either pamC1 or pamC2 demonstrated that the pamC alleles are the determinants of galactan structure. Experiments with recombinant PamC1 and PamC2 in vitro established that these proteins are galactan synthases capable of extending synthetic Galf disaccharide acceptors in the presence of UDP-Galf. Homology analysis identified critical amino acids that are essential for PamC1 and PamC2 enzymatic activity both in vitro and in K. kingae. Structural analysis of the in vitro-modified synthetic acceptors implicated PamC1 as a monofunctional enzyme capable of generating a β-(1 → 5) Galf linkage and PamC2 as a bifunctional enzyme capable of generating β-(1 → 3) and β-(1 → 6) Galf linkages. This study advances our understanding of the GT2 family of UDP-galactofuranosyltransferases.

60 APPLIED LIFE SCIENCES