Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Identification and characterization of mono- and bifunctional galactan synthases in the pediatric pathogen Kingella kingae

The emerging pediatric pathogen Kingella kingae elaborates a lipopolysaccharide (LPS) that is extended with a galactofuranose homopolymer called galactan, which is a key virulence determinant that contributes to resistance to complement-mediated and neutrophil-mediated killing. Previous work has demonstrated that the pamABCDE locus is required for galactan synthesis. In this study, mutational studies suggested that the pamC gene product is a UDP-galactofuranose (Galf) transferase and is the galactan synthase. Analysis of genome sequence data revealed two distinct pamC alleles designated pamC1 and pamC2, which correlate with the two galactan structures in K. kingae. Examination of isogenic mutants expressing either pamC1 or pamC2 demonstrated that the pamC alleles are the determinants of galactan structure. Experiments with recombinant PamC1 and PamC2 in vitro established that these proteins are galactan synthases capable of extending synthetic Galf disaccharide acceptors in the presence of UDP-Galf. Homology analysis identified critical amino acids that are essential for PamC1 and PamC2 enzymatic activity both in vitro and in K. kingae. Structural analysis of the in vitro-modified synthetic acceptors implicated PamC1 as a monofunctional enzyme capable of generating a β-(1 → 5) Galf linkage and PamC2 as a bifunctional enzyme capable of generating β-(1 → 3) and β-(1 → 6) Galf linkages. This study advances our understanding of the GT2 family of UDP-galactofuranosyltransferases.

60 APPLIED LIFE SCIENCES

Microbial Community Changes across Time and Space in a Constructed Wetland

Constructed wetlands are artificial ecosystems designed to replicate natural wetland processes. Microbial communities play a pivotal role in cycling essential elements, particularly sulfur, which is crucial for trace metal fixation and remobilization in these ecosystems. By their response to their environment, microbial communities act as biological indicators of the wetland performance. To address knowledge gaps pertinent to the changes in trace metal bioavailability in relation to microbial activities in the H-02 constructed wetland, we performed this study to investigate temporal and spatial variations in microbial communities by using molecular biology tools. Quantitative polymerase chain reaction and next generation sequencing techniques were employed to analyze archaeal and bacterial groups associated with sulfur and methane cycling. Alpha diversity indices were used to assess species richness, evenness, and dominance. Results indicated high gene abundance of Desulfuromonas (5.37 × 10 6 g.cell –1 ), methane oxidizing bacteria (6.92 × 10 6 g.cell –1 ), and methanogenic microorganisms (3.02 × 10 5 g.cell –1 ) during cool months. Warm months were marked by sulfate reducing bacteria dominance (3.31 × 10 6 g.cell –1 ), potentially due to competitive interactions and environmental conditions, higher temperatures, and lower redox potential. Spatial variability among microbial groups was insignificant, but trends in gene abundance indicated complex factors influencing these groups. Next generation sequencing data demonstrated Firmicutes as the most abundant phylum with over 50% regardless of the season or sampling location. Cool months exhibited higher alpha diversity than warm months. Overall, this study showed that seasonal changes significantly impacted the microbial communities in the H-02 constructed wetland that are associated with the sulfur cycle and eventually trace metal biogeochemistry, revealing two distinct mechanisms of the sulfur cycle between the two main seasons, whereas spatial variability effects were not conclusive.

54 ENVIRONMENTAL SCIENCES

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology

Novel, active, and uncultured hydrocarbon-degrading microbes in the ocean

ABSTRACT Given the vast quantity of oil and gas input to the marine environment annually, hydrocarbon degradation by marine microorganisms is an essential ecosystem service. Linkages between taxonomy and hydrocarbon degradation capabilities are largely based on cultivation studies, leaving a knowledge gap regarding the intrinsic ability of uncultured marine microbes to degrade hydrocarbons. To address this knowledge gap, metagenomic sequence data from the Deepwater Horizon (DWH) oil spill deep-sea plume was assembled to which metagenomic and metatranscriptomic reads were mapped. Assembly and binning produced new DWH metagenome-assembled genomes that were evaluated along with their close relatives, all of which are from the marine environment (38 total). These analyses revealed globally distributed hydrocarbon-degrading microbes with clade-specific substrate degradation potentials that have not been reported previously. For example, methane oxidation capabilities were identified in all Cycloclasticus . Furthermore, all Bermanella encoded and expressed genes for non-gaseous n -alkane degradation; however, DWH Bermanella encoded alkane hydroxylase, not alkane 1-monooxygenase. All but one previously unrecognized DWH plume member in the SAR324 and UBA11654 have the capacity for aromatic hydrocarbon degradation. In contrast, Colwellia were diverse in the hydrocarbon substrates they could degrade. All clades encoded nutrient acquisition strategies and response to cold temperatures, while sensory and acquisition capabilities were clade specific. These novel insights regarding hydrocarbon degradation by uncultured planktonic microbes provides missing data, allowing for better prediction of the fate of oil and gas when hydrocarbons are input to the ocean, leading to a greater understanding of the ecological consequences to the marine environment. IMPORTANCE Microbial degradation of hydrocarbons is a critically important process promoting ecosystem health, yet much of what is known about this process is based on physiological experiments with a few hydrocarbon substrates and cultured microbes. Thus, the ability to degrade the diversity of hydrocarbons that comprise oil and gas by microbes in the environment, particularly in the ocean, is not well characterized. Therefore, this study aimed to utilize non-cultivation-based ‘omics data to explore novel genomes of uncultured marine microbes involved in degradation of oil and gas. Analyses of newly assembled metagenomic data and previously existing genomes from other marine data sets, with metagenomic and metatranscriptomic read recruitment, revealed globally distributed hydrocarbon-degrading marine microbes with clade-specific substrate degradation potentials that have not been previously reported. This new understanding of oil and gas degradation by uncultured marine microbes suggested that the global ocean harbors a diversity of hydrocarbon-degrading bacteria, which can act as primary agents regulating ecosystem health.

Howe, Kathryn L.

Whole genome sequence of the syntrophic benzoate-degrading bacterium, Syntrophus buswellii DSM 102354 T

The syntrophic obligate proton-reducing bacterium Syntrophus buswellii DSM 102354 T , isolated from anaerobic digester sludge, degrades benzoate and possibly hydrocinnamate (phenyl-3-propionate) when grown with a suitable hydrogenotrophic partner. The complete genome sequence data provide insight regarding aromatic compound degradation under thermodynamically limiting conditions when exogenous electron acceptors are absent.

Syntrophus

Identification and characterization of a skin microbiome on Caenorhabditis elegans suggests environmental microbes confer cuticle protection

ABSTRACT In the wild, C. elegans are emersed in environments teeming with a veritable menagerie of microorganisms. The C. elegans cuticular surface serves as a barrier and first point of contact with their microbial environments. In this study, we identify microbes from C. elegans natural habitats that associate with its cuticle, constituting a simple “skin microbiome.” We rear our animals on a modified CeMbio, mCeMbio, a consortium of ecologically relevant microbes. We first combine standard microbiological methods with an adapted micro skin-swabbing tool to describe the skin-resident bacteria on the C. elegans surface. Furthermore, we conduct 16S rRNA gene sequencing studies to identify relative shifts in the proportion of mCeMbio bacteria upon surface-sterilization, implying distinct skin- and gut-microbiomes. We find that some strains of bacteria, including Enterobacter sp. JUb101 , are primarily found on the nematode skin, while others like Stenotrophomonas indicatrix JUb19 and Ochrobactrum vermis MYb71 are predominantly found in the animal’s gut. Finally, we show that this skin microbiome promotes host cuticle integrity in harsh environments. Together, we identify a skin microbiome for the well-studied nematode model and propose its value in conferring host fitness advantages in naturalized contexts. IMPORTANCE The genetic model organism C. elegans has recently emerged as a tool for understanding host–microbiome interactions. Nearly all of these studies either focus on pathogenic or gut-resident microbes. Little is known about the existence of native, nonpathogenic skin microbes or their function. We demonstrate that members of a modified C. elegans model microbiome, mCeMbio, can adhere to the animal's cuticle and confer protection from noxious environments. We combine a novel micro-swab tool, the first 16S microbial sequencing data from relatively unperturbed C. elegans , and physiological assays to demonstrate microbially mediated protection of the skin. This work serves as a foundation to explore wild C. elegans skin microbiomes and use C. elegans as a model for skin research.

16S RNA

data-encoder-circuits v1.0

Lightweight python package built on top of Qiskit to generate quantum circuits that can be used to load classical data (sequence of numbers) on a quantum computer.

Camps, Daan

Sempervirens: A Fast Reconstruction Algorithm for Noisy and Incomplete Binary Matrix Representations of Trees

Applications such as reconstructing cell lineage trees (represented as phylogenetic trees) from single-cell sequencing data require reconstructing a {0,1}-matrix that has many errors and missing entries. We introduce Sempervirens, a very fast matrix reconstruction algorithm for noisy and incomplete matrix representations of phylogenetic trees. Sempervirens uses an iterative maximum-likelihood approach to determine the topology tree represented by the corrupted data. We show that Sempervirens is at least three orders of magnitude faster than other methods on thousand by thousand matrices, with the speed gap widening with larger matrices. We also show that Sempervirens matches state-of-the-art methods in reconstruction accuracy. The speed of Sempervirens enables it to be tractably applied to reconstructing much larger matrices than those that other methods can reconstruct. In addition to experimental results, we justify the algorithm with a mathematical treatment of its subprocedures.

algorithms

Isolation and characterization of IgG3 glycan-targeting antibodies with exceptional cross-reactivity for diverse viral families

Broadly reactive antibodies that target sequence-diverse antigens are of interest for vaccine design and monoclonal antibody therapeutic development because they can protect against multiple strains of a virus and provide a barrier to evolution of escape mutants. Using LIBRA-seq (linking B cell receptor to antigen specificity through sequencing) data for the B cell repertoire of an individual chronically infected with human immunodeficiency virus type 1 (HIV-1), we identified a lineage of IgG3 antibodies predicted to bind to HIV-1 Envelope (Env) and influenza A Hemagglutinin (HA). Two lineage members, antibodies 2526 and 546, were confirmed to bind to a large panel of diverse antigens, including several strains of HIV-1 Env, influenza HA, coronavirus (CoV) spike, hepatitis C virus (HCV) E protein, Nipah virus (NiV) F protein, and Langya virus (LayV) F protein. We found that both antibodies bind to complex glycans on the antigenic surfaces. Antibody 2526 targets the stem region of influenza HA and the N-terminal domain (NTD) region of SARS-CoV-2 spike. A crystal structure of 2526 Fab bound to mannose revealed the presence of a glycan-binding pocket on the light chain. Antibody 2526 cross-reacted with antigens from multiple pathogens and displayed no signs of autoreactivity. These features distinguish antibody 2526 from previously described glycan-reactive antibodies. Further study of this antibody class may aid in the selection and engineering of broadly reactive antibody therapeutics and can inform the development of effective vaccines with exceptional breadth of pathogen coverage.

Microbiology

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES

Post Wildfire Soil Bacterial MAGs and Metagenome Analysis

We compare the differences between bacteria in soil affected by a wildfire to an unaffected area from Minnewaska State Park, NY located in the biodiverse northern Shawangunk ridge. We detail our metagenomic sequencing data, relative abundance of bacterial phyla, and the taxonomic classification of three high-quality MAGs.

59 BASIC BIOLOGICAL SCIENCES

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES