Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks

Ecological connectivity and habitat loss shape patterns of genetic diversity in a threatened salamander

Context The maintenance of genetic diversity is essential for preserving adaptive potential in populations, yet it is increasingly threatened by landscape alteration. The field of landscape genetics offers a framework for assessing how patch-level landscape conditions, modeled at multiple scales, influence genetic diversity. Objectives We sought to assess how local environmental features and connectivity influence genetic diversity across 74 four-toed salamander (Hemidactylium scutatum) breeding wetlands in the southeastern United States. Methods Using next-generation sequencing data and hierarchical Bayesian models, we examined genome-wide heterozygosity in relation to local landscape features and ecological connectivity. We also assessed the scale of effect of landscape features and tested for temporal lag effects. Results Genetic diversity was lower in wetlands with higher levels of historic deforestation and lower connectivity. An interaction between deforestation and connectivity indicated that deforestation had stronger negative effects in isolated wetlands but weaker effects in well-connected wetlands. Accounting for scale of effect and temporal lags was critical for detecting these relationships. Conclusions Our analyses highlight the importance of assessing the spatial scale (scale of effect) and temporal lag of landscape features to detect key drivers of genetic diversity. In line with population genetic theory, our results indicate that the genetic consequences of habitat loss do not affect populations uniformly and are most severe in isolated populations where gene flow cannot buffer against loss of diversity. Altogether, we highlight the importance of considering the interaction of habitat loss and connectivity in conservation genetic management.

Hemidactylium scutatum

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag

Identification and characterization of mono- and bifunctional galactan synthases in the pediatric pathogen Kingella kingae

The emerging pediatric pathogen Kingella kingae elaborates a lipopolysaccharide (LPS) that is extended with a galactofuranose homopolymer called galactan, which is a key virulence determinant that contributes to resistance to complement-mediated and neutrophil-mediated killing. Previous work has demonstrated that the pamABCDE locus is required for galactan synthesis. In this study, mutational studies suggested that the pamC gene product is a UDP-galactofuranose (Galf) transferase and is the galactan synthase. Analysis of genome sequence data revealed two distinct pamC alleles designated pamC1 and pamC2, which correlate with the two galactan structures in K. kingae. Examination of isogenic mutants expressing either pamC1 or pamC2 demonstrated that the pamC alleles are the determinants of galactan structure. Experiments with recombinant PamC1 and PamC2 in vitro established that these proteins are galactan synthases capable of extending synthetic Galf disaccharide acceptors in the presence of UDP-Galf. Homology analysis identified critical amino acids that are essential for PamC1 and PamC2 enzymatic activity both in vitro and in K. kingae. Structural analysis of the in vitro-modified synthetic acceptors implicated PamC1 as a monofunctional enzyme capable of generating a β-(1 → 5) Galf linkage and PamC2 as a bifunctional enzyme capable of generating β-(1 → 3) and β-(1 → 6) Galf linkages. This study advances our understanding of the GT2 family of UDP-galactofuranosyltransferases.

60 APPLIED LIFE SCIENCES

Microbial Community Changes across Time and Space in a Constructed Wetland

Constructed wetlands are artificial ecosystems designed to replicate natural wetland processes. Microbial communities play a pivotal role in cycling essential elements, particularly sulfur, which is crucial for trace metal fixation and remobilization in these ecosystems. By their response to their environment, microbial communities act as biological indicators of the wetland performance. To address knowledge gaps pertinent to the changes in trace metal bioavailability in relation to microbial activities in the H-02 constructed wetland, we performed this study to investigate temporal and spatial variations in microbial communities by using molecular biology tools. Quantitative polymerase chain reaction and next generation sequencing techniques were employed to analyze archaeal and bacterial groups associated with sulfur and methane cycling. Alpha diversity indices were used to assess species richness, evenness, and dominance. Results indicated high gene abundance of Desulfuromonas (5.37 × 10 6 g.cell –1 ), methane oxidizing bacteria (6.92 × 10 6 g.cell –1 ), and methanogenic microorganisms (3.02 × 10 5 g.cell –1 ) during cool months. Warm months were marked by sulfate reducing bacteria dominance (3.31 × 10 6 g.cell –1 ), potentially due to competitive interactions and environmental conditions, higher temperatures, and lower redox potential. Spatial variability among microbial groups was insignificant, but trends in gene abundance indicated complex factors influencing these groups. Next generation sequencing data demonstrated Firmicutes as the most abundant phylum with over 50% regardless of the season or sampling location. Cool months exhibited higher alpha diversity than warm months. Overall, this study showed that seasonal changes significantly impacted the microbial communities in the H-02 constructed wetland that are associated with the sulfur cycle and eventually trace metal biogeochemistry, revealing two distinct mechanisms of the sulfur cycle between the two main seasons, whereas spatial variability effects were not conclusive.

54 ENVIRONMENTAL SCIENCES

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.

Whole genome sequence of the syntrophic benzoate-degrading bacterium, Syntrophus buswellii DSM 102354 T

The syntrophic obligate proton-reducing bacterium Syntrophus buswellii DSM 102354 T , isolated from anaerobic digester sludge, degrades benzoate and possibly hydrocinnamate (phenyl-3-propionate) when grown with a suitable hydrogenotrophic partner. The complete genome sequence data provide insight regarding aromatic compound degradation under thermodynamically limiting conditions when exogenous electron acceptors are absent.

Syntrophus

Identification and characterization of a skin microbiome on Caenorhabditis elegans suggests environmental microbes confer cuticle protection

ABSTRACT In the wild, C. elegans are emersed in environments teeming with a veritable menagerie of microorganisms. The C. elegans cuticular surface serves as a barrier and first point of contact with their microbial environments. In this study, we identify microbes from C. elegans natural habitats that associate with its cuticle, constituting a simple “skin microbiome.” We rear our animals on a modified CeMbio, mCeMbio, a consortium of ecologically relevant microbes. We first combine standard microbiological methods with an adapted micro skin-swabbing tool to describe the skin-resident bacteria on the C. elegans surface. Furthermore, we conduct 16S rRNA gene sequencing studies to identify relative shifts in the proportion of mCeMbio bacteria upon surface-sterilization, implying distinct skin- and gut-microbiomes. We find that some strains of bacteria, including Enterobacter sp. JUb101 , are primarily found on the nematode skin, while others like Stenotrophomonas indicatrix JUb19 and Ochrobactrum vermis MYb71 are predominantly found in the animal’s gut. Finally, we show that this skin microbiome promotes host cuticle integrity in harsh environments. Together, we identify a skin microbiome for the well-studied nematode model and propose its value in conferring host fitness advantages in naturalized contexts. IMPORTANCE The genetic model organism C. elegans has recently emerged as a tool for understanding host–microbiome interactions. Nearly all of these studies either focus on pathogenic or gut-resident microbes. Little is known about the existence of native, nonpathogenic skin microbes or their function. We demonstrate that members of a modified C. elegans model microbiome, mCeMbio, can adhere to the animal's cuticle and confer protection from noxious environments. We combine a novel micro-swab tool, the first 16S microbial sequencing data from relatively unperturbed C. elegans , and physiological assays to demonstrate microbially mediated protection of the skin. This work serves as a foundation to explore wild C. elegans skin microbiomes and use C. elegans as a model for skin research.

16S RNA

data-encoder-circuits v1.0

Lightweight python package built on top of Qiskit to generate quantum circuits that can be used to load classical data (sequence of numbers) on a quantum computer.

Camps, Daan

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms

Isolation and characterization of IgG3 glycan-targeting antibodies with exceptional cross-reactivity for diverse viral families

Broadly reactive antibodies that target sequence-diverse antigens are of interest for vaccine design and monoclonal antibody therapeutic development because they can protect against multiple strains of a virus and provide a barrier to evolution of escape mutants. Using LIBRA-seq (linking B cell receptor to antigen specificity through sequencing) data for the B cell repertoire of an individual chronically infected with human immunodeficiency virus type 1 (HIV-1), we identified a lineage of IgG3 antibodies predicted to bind to HIV-1 Envelope (Env) and influenza A Hemagglutinin (HA). Two lineage members, antibodies 2526 and 546, were confirmed to bind to a large panel of diverse antigens, including several strains of HIV-1 Env, influenza HA, coronavirus (CoV) spike, hepatitis C virus (HCV) E protein, Nipah virus (NiV) F protein, and Langya virus (LayV) F protein. We found that both antibodies bind to complex glycans on the antigenic surfaces. Antibody 2526 targets the stem region of influenza HA and the N-terminal domain (NTD) region of SARS-CoV-2 spike. A crystal structure of 2526 Fab bound to mannose revealed the presence of a glycan-binding pocket on the light chain. Antibody 2526 cross-reacted with antigens from multiple pathogens and displayed no signs of autoreactivity. These features distinguish antibody 2526 from previously described glycan-reactive antibodies. Further study of this antibody class may aid in the selection and engineering of broadly reactive antibody therapeutics and can inform the development of effective vaccines with exceptional breadth of pathogen coverage.

Microbiology

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES

Post Wildfire Soil Bacterial MAGs and Metagenome Analysis

We compare the differences between bacteria in soil affected by a wildfire to an unaffected area from Minnewaska State Park, NY located in the biodiverse northern Shawangunk ridge. We detail our metagenomic sequencing data, relative abundance of bacterial phyla, and the taxonomic classification of three high-quality MAGs.

59 BASIC BIOLOGICAL SCIENCES

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES