Search NASA⌕ Search

Engineering topics

Henry, Christopher S.

Publications and source records attributed to Henry, Christopher S..

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

Enabling Capabilities and Resources: 2024 Principal Investigator Meeting Proceedings

As a major supporter of basic genome-enabled research, BER’s Biological Systems Science Division (BSSD) fosters scientific discovery by funding - fundamental biological research across disciplines in conjunction with enabling investigational tools and computational capabilities that include world-class user facilities. The overarching goal of BSSD is to provide the necessary fundamental science to understand, predict, manipulate, and design biological systems that underpin innovations for bioenergy and bioproduct production and enhance understanding of natural, DOE-relevant environmental processes (Biological Systems Science Division Strategic Plan, 2021). To accelerate the U.S. bioeconomy, BSSD pursues innovative science underpinning advances in sustainable biofuels and bioproducts and the development of next-generation technologies and computational resources for systems biology research. The 2024 BSSD Enabling Capabilities and Resources (ECR) Principal Investigator (PI) meeting brought together PIs across the BSSD ECR portfolio to confer on shared interests and opportunities. The meeting was held concurrently with the Genomic Science program (GSP) PI meeting to optimize collaboration on research to advance bioenergy and the bioeconomy. Rick Stevens of Argonne National Laboratory gave a keynote on How Generative Artificial Intelligence Can Impact Biological Research (see Keynote: How Generative Artificial Intelligence Can Impact Biological Research, this page). Plenary presentations included several joint sessions that illuminated the integration and understanding of the larger BSSD mission. GSP’s objective is to provide systems-level understanding of plants, microbes, and their communities through its Bioenergy Research, Biosystems Design, and Environmental Microbiome Research portfolios. The objective of the ECR portfolio is to support development of computational and instrumental platforms to advance fundamental GSP research—and BER more broadly— toward the overall goal of understanding the functional principles of living systems and their response to environmental challenges.

59 BASIC BIOLOGICAL SCIENCES↗

TranSyT , an innovative framework for identifying transport systems

The importance and rate of development of genome-scale metabolic models have been growing for the last few years, increasing the demand for software solutions that automate several steps of this process. However, since TRIAGE’s release, software development for the automatic integration of transport reactions into models has stalled. Here, in this paper, we present the Transport Systems Tracker (TranSyT). Unlike other transport systems annotation software, TranSyT does not rely on manual curation to expand its internal database, which is derived from highly curated records retrieved from the Transporters Classification Database and complemented with information from other data sources. TranSyT compiles information regarding transporter families and proteins, and derives reactions into its internal database, making it available for rapid annotation of complete genomes. All transport reactions have GPR associations and can be exported with identifiers from four different metabolite databases. TranSyT is currently available as a plugin for merlin v4.0 and an app for KBase.

59 BASIC BIOLOGICAL SCIENCES↗

kb_DRAM: annotation and metabolic profiling of genomes with DRAM in KBase

Microbial genome annotation is the process of identifying structural and functional elements in DNA sequences and subsequently attaching biological information to those elements. DRAM is a tool developed to annotate bacterial, archaeal, and viral genomes derived from pure cultures or metagenomes. DRAM goes beyond traditional annotation tools by distilling multiple gene annotations to genome level summaries of functional potential. Despite these benefits, a downside of DRAM is the requirement of large computational resources, which limits its accessibility. Further, it did not integrate with downstream metabolic modeling tools that require genome annotation. To alleviate these constraints, DRAM and the viral counterpart, DRAM-v, are now available and integrated with the freely accessible KBase cyberinfrastructure. With kb_DRAM users can generate DRAM annotations and functional summaries from microbial or viral genomes in a point-and-click interface, as well as generate genome-scale metabolic models from DRAM annotations.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-scale metabolic reconstruction of 7,302 human microorganisms for personalized medicine

The human microbiome influences the efficacy and safety of a wide variety of commonly prescribed drugs. Designing precision medicine approaches that incorporate microbial metabolism would require strain- and molecule-resolved, scalable computational modeling. Here, we extend our previous resource of genome-scale metabolic reconstructions of human gut microorganisms with a greatly expanded version. AGORA2 (assembly of gut organisms through reconstruction and analysis, version 2) accounts for 7,302 strains, includes strain-resolved drug degradation and biotransformation capabilities for 98 drugs, and was extensively curated based on comparative genomics and literature searches. The microbial reconstructions performed very well against three independently assembled experimental datasets with an accuracy of 0.72 to 0.84, surpassing other reconstruction resources and predicted known microbial drug transformations with an accuracy of 0.81. We demonstrate that AGORA2 enables personalized, strain-resolved modeling by predicting the drug conversion potential of the gut microbiomes from 616 patients with colorectal cancer and controls, which greatly varied between individuals and correlated with age, sex, body mass index and disease stages. AGORA2 serves as a knowledge base for the human microbiome and paves the way to personalized, predictive analysis of host–microbiome metabolic interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Respiratory energy demands and scope for demand expansion and destruction

Photosynthesis is the primary energy input to plants, but most plant metabolic processes are powered much or all of the time by ATP and NAD(P)H that come from respiratory oxidation of photosynthetically produced carbohydrates, that is, “dark respiration” (Amthor, 1994). The processes of growth, nutrient uptake and assimilation, active transport, and maintenance are thus all clients of respiration and compete for their share of a respiratory energy budget that is capped by photosynthate income. The economic principle of opportunity cost applies to these clients’ competing demands: spending respiratory energy on process A means missing the benefit of investing that energy in process B (Mahmoudabadi et al., 2019). This principle is key to assessing prospects for crop improvement by metabolic engineering. As de Lorenzo (2015) put it: “metabolism…frames and ultimately resolves whether a given genetic program (existing…or engineered) can be deployed or not.” Foreign or reconfigured native processes bolted on to a crop-plant’s metabolic chassis must compete for respiratory energy with native ones without crashing the energy economy. A proviso on plant carbon budgets is that photosynthetic carbon fixation (“source activity”) can in certain cases increase to meet increased carbon demand (“sink activity”), that is, budget envelopes are not always fixed (Smith et al., 2018). However, as there is some consensus that productivity is most often limited or co-limited by carbon supply (Ainsworth and Long, 2005; Körner, 2015; Sonnewald and Fernie, 2018), we make this our basal assumption in the analyses below.

59 BASIC BIOLOGICAL SCIENCES↗

Functional characterization of prokaryotic dark matter: the road so far and what lies ahead

Eight-hundred thousand to one trillion prokaryotic species may inhabit our planet. Yet, fewer than two-hundred thousand prokaryotic species have been described. This uncharted fraction of microbial diversity, and its undisclosed coding potential, is known as the “microbial dark matter” (MDM). Next-generation sequencing has allowed to collect a massive amount of genome sequence data, leading to unprecedented advances in the field of genomics. Still, harnessing new functional information from the genomes of uncultured prokaryotes is often limited by standard classification methods. These methods often rely on sequence similarity searches against reference genomes from cultured species. This hinders the discovery of unique genetic elements that are missing from the cultivated realm. It also contributes to the accumulation of prokaryotic gene products of unknown function among public sequence data repositories, highlighting the need for new approaches for sequencing data analysis and classification. Increasing evidence indicates that these proteins of unknown function might be a treasure trove of biotechnological potential. Here, we outline the challenges, opportunities, and the potential hidden within the functional dark matter (FDM) of prokaryotes. We also discuss the pitfalls surrounding molecular and computational approaches currently used to probe these uncharted waters, and discuss future opportunities for research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolite Damage and Damage Control in a Minimal Genome

Analysis of the genes retained in the minimized Mycoplasma JCVI-Syn3A genome established that systems that repair or preempt metabolite damage are essential to life. Several genes known to have such functions were identified and experimentally validated, including 5-formyltetrahydrofolate cycloligase, coenzyme A (CoA) disulfide reductase, and certain hydrolases. Furthermore, we discovered that an enigmatic YqeK hydrolase domain fused to NadD has a novel proofreading function in NAD synthesis and could double as a MutT-like sanitizing enzyme for the nucleotide pool. Finally, we combined metabolomics and cheminformatics approaches to extend the core metabolic map of JCVI-Syn3A to include promiscuous enzymatic reactions and spontaneous side reactions. This extension revealed that several key metabolite damage control systems remain to be identified in JCVI-Syn3A, such as that for methylglyoxal.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-scale modelling of the primary-specialized metabolism interface

Environmental challenges and development require plants to reallocate resources between primary and specialized metabolites to survive. Genome-scale metabolic models, which map carbon flux through metabolic pathways, are a valuable tool in the study of tradeoffs that arise at this interface. Due to annotation gaps, models that characterize all the enzymatic steps in individual specialized pathways and their linkages to each other and to central carbon metabolism are difficult to construct. Recent studies have successfully curated subsystems of specialized metabolism and characterized the interfaces where flux is diverted to the precursors of glucosinolates, terpenes, and anthocyanins. Although advances in metabolite profiling can help to constrain models at this interface, quantitative analysis remains challenging because of the different timescales on which specialized metabolites from constitutive and reactive pathways accumulate.

59 BASIC BIOLOGICAL SCIENCES↗

Application of the metabolic modeling pipeline in KBase to categorize reactions, predict essential genes, and predict pathways in an isolate genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

DOE knowledgebase↗

Chemical-damage MINE: A database of curated and predicted spontaneous metabolic reactions

Spontaneous reactions between metabolites are often neglected in favor of emphasizing enzyme-catalyzed chemistry because spontaneous reaction rates are assumed to be insignificant under physiological conditions. However, synthetic biology and engineering efforts can raise natural metabolites' levels or introduce unnatural ones, so that previously innocuous or nonexistent spontaneous reactions become an issue. Problems arise when spontaneous reaction rates exceed the capacity of a platform organism to dispose of toxic or chemically active reaction products. While various reliable sources list competing or toxic enzymatic pathways' side-reactions, no corresponding compilation of spontaneous side-reactions exists, nor is it possible to predict their occurrence. We addressed this deficiency by creating the Chemical Damage (CD)-MINE resource. First, we used literature data to construct a comprehensive database of metabolite reactions that occur spontaneously in physiological conditions. We then leveraged this data to construct 148 reaction rules describing the known spontaneous chemistry in a substrate-generic way. We applied these rules to all compounds in the ModelSEED database, predicting 180,891 spontaneous reactions. The resulting (CD)-MINE is available at https://minedatabase.mcs.anl. gov/cdmine/#/home and through developer tools. We also demonstrate how damage-prone intermediates and end products are widely distributed among metabolic pathways, and how predicting spontaneous chemical damage helps rationalize toxicity and carbon loss using examples from published pathways to commercial products. We explain how analyzing damage-prone areas in metabolism helps design effective engineering strategies. Finally, we use the CD-MINE toolset to predict the formation of the novel damage product N-carbamoyl proline, and present mass spectrometric evidence for its presence in Escherichia coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Moderately (D)efficient Enzyme: Catalysis-Related Damage In Vivo and Its Repair

Enzymes have in vivo life spans. Analysis of life spans, i.e., lifetime totals of catalytic turnovers, suggests that nonsurvivable collateral chemical damage from the very reactions that enzymes catalyze is a common but underdiagnosed cause of enzyme death. Analysis also implies that many enzymes are moderately deficient in that their active-site regions are not naturally as hardened against such collateral damage as they could be, leaving room for improvement by rational design or directed evolution. Enzyme life span might also be improved by engineering systems that repair otherwise fatal active-site damage, of which a handful are known and more are inferred to exist. Unfortunately, the data needed to design and execute such improvements are lacking: there are too few measurements of in vivo life span, and existing information about the extent, nature, and mechanisms of active-site damage and repair during normal enzyme operation is too scarce, anecdotal, and speculative to act on. Fortunately, advances in proteomics, metabolomics, cheminformatics, comparative genomics, and structural biochemistry now empower a systematic, data-driven approach for identifying, predicting, and validating instances of active-site damage and its repair. These capabilities would be practically useful in enzyme redesign and improvement of in-use stability and could change our thinking about which enzymes die young in vivo, and why.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ORT: a workflow linking genome-scale metabolic models with reactive transport codes

Abstract Motivation Nutrient and contaminant behavior in the subsurface are governed by multiple coupled hydrobiogeochemical processes which occur across different temporal and spatial scales. Accurate description of macroscopic system behavior requires accounting for the effects of microscopic and especially microbial processes. Microbial processes mediate precipitation and dissolution and change aqueous geochemistry, all of which impacts macroscopic system behavior. As ‘omics data describing microbial processes is increasingly affordable and available, novel methods for using this data quickly and effectively for improved ecosystem models are needed. Results We propose a workflow (‘Omics to Reactive Transport—ORT) for utilizing metagenomic and environmental data to describe the effect of microbiological processes in macroscopic reactive transport models. This workflow utilizes and couples two open-source software packages: KBase (a software platform for systems biology) and PFLOTRAN (a reactive transport modeling code). We describe the architecture of ORT and demonstrate an implementation using metagenomic and geochemical data from a river system. Our demonstration uses microbiological drivers of nitrification and denitrification to predict nitrogen cycling patterns which agree with those provided with generalized stoichiometries. While our example uses data from a single measurement, our workflow can be applied to spatiotemporal metagenomic datasets to allow for iterative coupling between KBase and PFLOTRAN. Availability and implementation Interactive models available at https://pflotranmodeling.paf.subsurfaceinsights.com/pflotran-simple-model/. Microbiological data available at NCBI via BioProject ID PRJNA576070. ORT Python code available at https://github.com/subsurfaceinsights/ort-kbase-to-pflotran. KBase narrative available at https://narrative.kbase.us/narrative/71260 or static narrative (no login required) at https://kbase.us/n/71260/258. Supplementary information Supplementary data are available at Bioinformatics online.

54 ENVIRONMENTAL SCIENCES↗