Search NASA⌕ Search

Engineering topics

Nelson, William C.

Publications and source records attributed to Nelson, William C..

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

A Mineral-Doped Micromodel Platform Demonstrates Fungal Bridging of Carbon Hot Spots and Hyphal Transport of Mineral-Derived Nutrients

Fungal species are foundational members of soil microbiomes, where their contributions in accessing and transporting vital nutrients is key for community resilience. To date, the molecular mechanisms underlying fungal mineral weathering and nutrient translocation in low-nutrient environments remain poorly resolved due to the lack of a platform for spatial analysis of biotic weathering processes.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of co-circulating pathogens and microbiome from COVID-19 infections

Co-infections or secondary infections with SARS-CoV-2 have the potential to affect disease severity and morbidity. Additionally, the potential influence of the nasal microbiome on COVID-19 illness is not well understood. In this study, we analyzed 203 residual samples, originally submitted for SARS-CoV-2 testing, for the presence of viral, bacterial, and fungal pathogens and non-pathogens using a comprehensive microarray technology, the Lawrence Livermore Microbial Detection Array (LLMDA). Eighty-seven percent of the samples were nasopharyngeal samples, and 23% of the samples were oral, nasal and oral pharyngeal swabs. We conducted bioinformatics analyses to examine differences in microbial populations of these samples, as a proxy for the nasal and oral microbiome, from SARS-CoV-2 positive and negative specimens. We found 91% concordance with the LLMDA relative to a diagnostic RT-qPCR assay for detection of SARS-CoV-2. Sixteen percent of all the samples (32/203) revealed the presence of an opportunistic bacterial or frank viral pathogen with the potential to cause co-infections. The two most detected bacteria, Streptococcus pyogenes and Streptococcus pneumoniae , were present in both SARS-CoV-2 positive and negative samples. Human metapneumovirus was the most prevalent viral pathogen in the SARS-CoV-2 negative samples. Sequence analysis of 16S rRNA was also conducted to evaluate bacterial diversity and confirm LLMDA results.

59 BASIC BIOLOGICAL SCIENCES↗

A roadmap for the functional annotation of protein families: a community perspective

Over the last 25 years, biology has entered the genomic era and is becoming a science of ‘big data’. Most interpretations of genomic analyses rely on accurate functional annotations of the proteins encoded by more than 500 000 genomes sequenced to date. By different estimates, only half the predicted sequenced proteins carry an accurate functional annotation, and this percentage varies drastically between different organismal lineages. Such a large gap in knowledge hampers all aspects of biological enterprise and, thereby, is standing in the way of genomic biology reaching its full potential. A brainstorming meeting to address this issue funded by the National Science Foundation was held during 3–4 February 2022. Bringing together data scientists, biocurators, computational biologists and experimentalists within the same venue allowed for a comprehensive assessment of the current state of functional annotations of protein families. Further, major issues that were obstructing the field were identified and discussed, which ultimately allowed for the proposal of solutions on how to move forward.

59 BASIC BIOLOGICAL SCIENCES↗

DNA Viral Diversity, Abundance, and Functional Potential Vary across Grassland Soils with a Range of Historical Moisture Regimes

Soil viruses are abundant, but the influence of the environment and climate on soil viruses remains poorly understood. Here, we addressed this gap by comparing the diversity, abundance, lifestyle, and metabolic potential of DNA viruses in three grassland soils with historical differences in average annual precipitation, low in eastern Washington (WA), high in Iowa (IA), and intermediate in Kansas (KS). Bioinformatics analyses were applied to identify a total of 2,631 viral contigs, including 14 complete viral genomes from three deep metagenomes (1 terabase [Tb] each) that were sequenced from bulk soil DNA. An additional three replicate metagenomes (~0.5 Tb each) were obtained from each location for statistical comparisons. Identified viruses were primarily bacteriophages targeting dominant bacterial taxa. Both viral and host diversity were higher in soil with lower precipitation. Viral abundance was also significantly higher in the arid WA location than in IA and KS. More lysogenic markers and fewer clustered regularly interspaced short palindromic repeats (CRISPR) spacer hits were found in WA, reflecting more lysogeny in historically drier soil. More putative auxiliary metabolic genes (AMGs) were also detected in WA than in the historically wetter locations. The AMGs occurring in 18 pathways could potentially contribute to carbon metabolism and energy acquisition in their hosts. Structural equation modeling (SEM) suggested that historical precipitation influenced viral life cycle and selection of AMGs. The observed and predicted relationships between soil viruses and various biotic and abiotic variables have value for predicting viral responses to environmental change.

59 BASIC BIOLOGICAL SCIENCES↗

The Specific Carbohydrate Diet and Diet Modification as Induction Therapy for Pediatric Crohn’s Disease: A Randomized Diet Controlled Trial

Background: Crohn’s disease (CD) is a chronic inflammatory intestinal disorder associated with intestinal dysbiosis. Diet modulates the intestinal microbiome and therefore has a therapeutic potential. The aim of this study is to determine the potential efficacy of three versions of the specific carbohydrate diet (SCD) in active Crohn’s Disease. Methods: 18 patients with mild/moderate CD (PCDAI 15–45) aged 7 to 18 years were enrolled. Patients were randomized to either SCD, modified SCD(MSCD) or whole foods (WF) diet. Patients were evaluated at baseline, 2, 4, 8 and 12 weeks. PCDAI, inflammatory labs and multi-omics evaluations were assessed. Results: Mean age was 14.3 ± 2.9 years. At week 12, all participants (n = 10) who completed the study achieved clinical remission. The C-reactive protein decreased from 1.3 ± 0.7 at enrollment to 0.9 ± 0.5 at 12 weeks in the SCD group. In the MSCD group, the CRP decreased from 1.6 ± 1.1 at enrollment to 0.7 ± 0.1 at 12 weeks. In the WF group, the CRP decreased from 3.9 ± 4.3 at enrollment to 1.6 ± 1.3 at 12 weeks. In addition, the microbiome composition shifted in all patients across the study period. While the nature of the changes was largely patient specific, the predicted metabolic mode of the organisms increasing and decreasing in activity was consistent across patients. Conclusions: This study emphasizes the impact of diet in CD. Each diet had a positive effect on symptoms and inflammatory burden; the more exclusionary diets were associated with a better resolution of inflammation.

60 APPLIED LIFE SCIENCES↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

Deconstructing the Soil Microbiome into Reduced-Complexity Functional Modules

The soil microbiome is an invaluable component of the biosphere and critical for ecosystem functions, including biogeochemical cycling, soil-atmosphere gas exchange, degradation of toxic compounds, and promotion of plant growth and stress resistance/resilience. An improved understanding of the soil microbiome will help with predicting how these processes respond to external perturbations, and with harnessing beneficial aspects of the soil microbiome for agronomic applications such as crop amendments. However, the extensive taxonomic and functional diversity inherent within the soil microbiome hinders efficient analysis of this system. Microbial biodiversity in soils is orders of magnitude greater than other commonly studied systems such as the human gut microbiome (Blum, Zechmeister-Boltenstern, and Keiblinger 2019; Berendsen, Pieterse, and Bakker 2012). Thousands of microbial taxa may be found in a single gram of soil (Roesch et al. 2007), and high rates of gene flow and mutation further promote microbial diversification (Sergaki et al. 2018). Concomitantly, functional diversity in soil is similarly extensive. Soil is a heterogeneous mixture of microenvironments with defined physical and chemical attributes (Bach et al. 2018), within which numerous microbial guilds of distinct life-strategies and metabolic capacities can be found (Perez-Garcia, Lear, and Singhal 2016; H.-S. Song et al. 2014). Furthermore, the soil microbiome harbors a significant fraction of rare and/or quiescent members (Blagodatskaya and Kuzyakov 2013) that exist below the threshold of detection of current technologies. Such rare taxa may make significant contributions to process rates (Shade and Gilbert 2015; Dawson et al. 2017), but their scarcity complicates their identification and analysis. Finally, the vast amounts of information generated through holistic analyses of soil communities represent an intense computational burden (Scholz, Lo, and Chain 2012; Prosser 2015) that precludes assessment of the complete functional and taxonomic diversity contained in this ecosystem.

Naylor, Dan T.↗