Search NASA⌕ Search

SEARCH · Search NASA

Results for “Models, Biological”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CryoSegNet: accurate cryo-EM protein particle picking by integrating the foundational AI image segmentation model and attention-gated U-Net

Picking protein particles in cryo-electron microscopy (cryo-EM) micrographs is a crucial step in the cryo-EM-based structure determination. However, existing methods trained on a limited amount of cryo-EM data still cannot accurately pick protein particles from noisy cryo-EM images. The general foundational artificial intelligence–based image segmentation model such as Meta’s Segment Anything Model (SAM) cannot segment protein particles well because their training data do not include cryo-EM images. Here, we present a novel approach (CryoSegNet) of integrating an attention-gated U-shape network (U-Net) specially designed and trained for cryo-EM particle picking and the SAM. The U-Net is first trained on a large cryo-EM image dataset and then used to generate input from original cryo-EM images for SAM to make particle pickings. CryoSegNet shows both high precision and recall in segmenting protein particles from cryo-EM micrographs, irrespective of protein type, shape and size. On several independent datasets of various protein types, CryoSegNet outperforms two top machine learning particle pickers crYOLO and Topaz as well as SAM itself. The average resolution of density maps reconstructed from the particles picked by CryoSegNet is 3.33 Å, 7% better than 3.58 Å of Topaz and 14% better than 3.87 Å of crYOLO. It is publicly available at https://github.com/jianlin-cheng/CryoSegNet

59 BASIC BIOLOGICAL SCIENCES↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Tree rings reveal the transient risk of extinction hidden inside climate envelope forecasts

Given the importance of climate in shaping species’ geographic distributions, climate change poses an existential threat to biodiversity. Climate envelope modeling, the predominant approach used to quantify this threat, presumes that individuals in populations respond to climate variability and change according to species-level responses inferred from spatial occurrence data—such that individuals at the cool edge of a species’ distribution should benefit from warming (the “leading edge”), whereas individuals at the warm edge should suffer (the “trailing edge”). Using 1,558 tree-ring time series of an aridland pine (Pinus edulis) collected at 977 locations across the species’ distribution, we found that trees everywhere grow less in warmer-than-average and drier-than-average years. Ubiquitous negative temperature sensitivity indicates that individuals across the entire distribution should suffer with warming—the entire distribution is a trailing edge. Species-level responses to spatial climate variation are opposite in sign to individual-scale responses to time-varying climate for approximately half the species’ distribution with respect to temperature and the majority of the species’ distribution with respect to precipitation. These findings, added to evidence from the literature for scale-dependent climate responses in hundreds of species, suggest that correlative, equilibrium-based range forecasts may fail to accurately represent how individuals in populations will be impacted by changing climate. A scale-dependent view of the impact of climate change on biodiversity highlights the transient risk of extinction hidden inside climate envelope forecasts and the importance of evolution in rescuing species from extinction whenever local climate variability and change exceeds individual-scale climate tolerances.

54 ENVIRONMENTAL SCIENCES↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Isolation of a Terminal Cobalt Nitride in a Metal–Organic Framework

Transition metal nitrides are reactive intermediates in biological and industrial processes. Chemists have synthesized molecular model complexes of such reactive species to understand their function and electronic requirements for new applications. However, molecular chemistry can suffer from intra- and intermolecular decomposition pathways, which preclude further discovery of unknown reactive species. Metal-organic frameworks offer an opportunity for creating long-lived forms of such species with the vacuum of the pore suppressing degradation while simultaneously enabling substrate access for controlled reactivity studies. Here, in this study, we report the characterization of an elusive terminal cobalt nitride species generated through photolysis or thermolysis of a site-isolated cobalt azide within the evacuated metal-organic framework CoN 3 -MFU-4l. The first crystal structure of such a species is presented, with vibrational, X-ray absorption, and electron paramagnetic resonance spectroscopies providing further direct evidence for its formation while elucidating its electronic structure. The system additionally enables subsequent reactivity studies with selected substrates, revealing a unique ambiphilic behavior for a metal nitride species.

Börgel, Jonas [University of California, Berkeley,↗

Simulating nationwide coupled disease and fear spread in an agent-based model

Human cognitive responses, behavioral responses, and disease dynamics co-evolve over the course of any disease outbreak, and can result in complex feedbacks. We present a dynamic agent-based model that explicitly couples the spread of disease with the spread of fear surrounding the disease, implemented within the EpiCast simulation framework. EpiCast models transmission within a realistic synthetic population, capturing individual-level interactions. In our model, fear propagates through both in-person contact and broadcast media, prompting individuals to adopt protective behaviors that reduce disease spread. In order to better understand these coupled dynamics, we create and compare a range of compartmental models to ensure that introducing additional disease states does not prevent the emergence of multiple waves in these simpler models. Additionally, we compare a range of behavioral scenarios within EpiCast, varying the level and intensity of fear and behavior change. Our results show that the addition of asymptomatic, exposed, and pre-symptomatic disease states can impact both the rate at which an outbreak progresses and its overall trajectory in compartmental models. In EpiCast, the combination of non-local fear spread via broadcasters and strong behavioral responses by fearful individuals generally leads to multiple epidemic waves, an outcome that occurs only within a narrow parameter range when fear spreads purely through local contact. Accounting for the coupled spread of fear and disease is critical for understanding disease dynamics and designing timely, targeted responses to emerging infectious threats.

60 APPLIED LIFE SCIENCES↗

Structural Insights into the Dynamics of Water in SOD1 Catalysis and Drug Interactions

Superoxide dismutase 1 (SOD1) is a crucial enzyme that protects cells from oxidative damage by converting superoxide radicals into H 2 O 2 and O 2 . This detoxification process, essential for cellular homeostasis, relies on a precisely orchestrated catalytic mechanism involving the copper cation, while the zinc cation contributes to the structural integrity of the enzyme. This study presents the 2.3 Å crystal structure of human SOD1 (PDB ID: 9IYK), revealing an assembly of six homodimers and twelve distinct active sites. The water molecules form a complex hydrogen-bonding network that drives proton transfer and sustains active site dynamics. Our structure also uncovers subtle conformational changes that highlight the intrinsic flexibility of SOD1, which is essential for its function. Additionally, we observe how these dynamic structural features may be linked to pathological mutations associated with amyotrophic lateral sclerosis (ALS). By advancing our understanding of hSOD1’s mechanistic intricacies and the influence of water coordination, this study offers valuable insights for developing therapeutic strategies targeting ALS. Our structure’s unique conformations and active site interactions illuminate new facets of hSOD1 function, underscoring the critical role of structural dynamics in enzyme catalysis. Moreover, we conducted a molecular docking analysis using SOD1 for potential radical scavengers and Abelson non-receptor tyrosine kinase (c-Abl, Abl1) inhibitors targeting misfolded SOD1 aggregation along with oxidative stress and apoptosis, respectively. The results showed that CHEMBL1075867, a free radical scavenger derivative, showed the most promising docking results and interactions at the binding site of hSOD1, highlighting its promising role for further studies against SOD1-mediated ALS.

60 APPLIED LIFE SCIENCES↗

In vitro toxicity assessment of uranium particulates on different human lung epithelial cell models

Inhalation of uranium aerosols produced via human activities such as mining can pose a threat to human respiratory systems. Uranium oxide particulates emit short-range alpha particles that elicit DNA and direct damage, beyond associated physiochemical heavy-metal toxicity, to internal epithelial tissues. The availability of reliable in vitro models to study radiation exposure can greatly enhance our ability to understand and combat the biological impacts of exposure. However, the toxicological effects of alpha emissions and/or the oxidation states of uranium particulates vary across different human lung epithelial cell models and have not been systematically compared. We have endeavored to address this limitation by comparing impacts in three different human lung cell models: primary human bronchial and tracheal epithelial cells, primary human small airway epithelial cells, and human adenocarcinoma alveolar basal epithelial cells. Other studies have mainly investigated the toxicity of depleted uranium. Here, we compared the exposure of uranium oxide particulates (U 3 O 8 and UO 3 ) of different enrichment states on the chosen cell systems. Each cell model was exposed to 0.1, 1, 10, 50, 100, and 500 µg/mL of depleted U 3 O 8 , highly-enriched U 3 O 8 , and natural UO 3 particulates for 24 hours in submerged monolayer cultures. We compared viability and superoxide dismutase activity results across cell lines and uranium enrichment/ oxidative states. The results showed that 1) the oxide state of the particulates affected cell viability, implying that uranium’s different oxidation states contribute to different toxicological responses, and 2) each cell model reacts differently when exposed to uranium oxides, which may provide insights into the mechanistic processes associated with the exposure of radiological particulates on different biological systems. For instance, increased uranium enrichment corresponds to increased toxicity for the primary cells, but not for the immortalized cells. Our study shows that a holistic approach that incorporates similarities between model systems and types of radionuclides is required to truly develop empirical solutions for radiation exposure.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput genetics enables identification of nutrient utilization and accessory energy metabolism genes in a model methanogen

Archaea are widespread in the environment and play fundamental roles in diverse ecosystems; however, characterization of their unique biology requires advanced tools. This is particularly challenging when characterizing gene function. Here, we generate randomly barcoded transposon libraries in the model methanogenic archaeon Methanococcus maripaludis and use high-throughput growth methods to conduct fitness assays (RB-TnSeq) across over 100 unique growth conditions. Using our approach, we identified new genes involved in nutrient utilization and response to oxidative stress. We identified novel genes for the usage of diverse nitrogen sources in M. maripaludis including a putative regulator of alanine deamination and molybdate transporters important for nitrogen fixation. Furthermore, leveraging the fitness data, we inferred that M. maripaludis can utilize additional nitrogen sources including $\tiny{L}$-glutamine, $\tiny{D}$-glucuronamide, and adenosine. Under autotrophic growth conditions, we identified a gene encoding a domain of unknown function (DUF166) that is important for fitness and hypothesize that it has an accessory role in carbon dioxide assimilation. Finally, comparing fitness costs of oxygen versus sulfite stress, we identified a previously uncharacterized class of dissimilatory sulfite reductase-like proteins (Dsr-LP; group IIId) that is important during growth in the presence of sulfite. When overexpressed, Dsr-LP conferred sulfite resistance and enabled use of sulfite as the sole sulfur source. The high-throughput approach employed here allowed for generation of a large-scale data set that can be used as a resource to further understand gene function and metabolism in the archaeal domain.

59 BASIC BIOLOGICAL SCIENCES↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change↗

A network-enabled pipeline for gene discovery and validation in non-model plant species

Identifying key regulators of important genes in non-model crop species is challenging due to limited multi-omics resources. To address this, we introduce the network-enabled gene discovery pipeline NEEDLE, a user-friendly tool that systematically generates coexpression gene network modules, measures gene connectivity, and establishes network hierarchy to pinpoint key transcriptional regulators from dynamic transcriptome datasets. After validating its accuracy with two independent datasets, we applied NEEDLE to identify transcription factors (TFs) regulating the expression of cellulose synthase-like F6 ( CSLF6 ), a crucial cell wall biosynthetic gene, in Brachypodium and sorghum. Our analyses uncover regulators of CSLF6 and also shed light on the evolutionary conservation or divergence of gene regulatory elements among grass species. These results highlight NEEDLE’s capability to provide biologically relevant TF predictions and demonstrate its value for non-model plant species with dynamic transcriptome datasets.

59 BASIC BIOLOGICAL SCIENCES↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Submicron immunoglobulin particles exhibit FcγRII-dependent toxicity linked to autophagy in TNFα-stimulated endothelial cells

In intravenous immunoglobulins (IVIG), and some other immunoglobulin products, protein particles have been implicated in adverse events. Role and mechanisms of immunoglobulin particles in vascular adverse effects of blood components and manufactured biologics have not been elucidated. We have developed a model of spherical silica microparticles (SiMPs) of distinct sizes 200–2000 nm coated with different IVIG- or albumin (HSA)-coronas and investigated their effects on cultured human umbilical vein endothelial cells (HUVEC). IVIG products (1–20 mg/mL), bare SiMPs or SiMPs with IVIG-corona, did not display significant toxicity to unstimulated HUVEC. In contrast, in TNFα-stimulated HUVEC, IVIG-SiMPs induced decrease of HUVEC viability compared to HSA-SiMPs, while no toxicity of soluble IVIG was observed. 200 nm IVIG-SiMPs after 24 h treatment further increased ICAM1 (intercellular adhesion molecule 1) and tissue factor surface expression, apoptosis, mammalian target of rapamacin (mTOR)-dependent activation of autophagy, and release of extracellular vesicles, positive for mitophagy markers. Toxic effects of IVIG-SiMPs were most prominent for 200 nm SiMPs and decreased with larger SiMP size. Using blocking antibodies, toxicity of IVIG-SiMPs was found dependent on FcγRII receptor expression on HUVEC, which increased after TNFα-stimulation. Similar results were observed with different IVIG products and research grade IgG preparations. In conclusion, submicron particles with immunoglobulin corona induced size-dependent toxicity in TNFα-stimulated HUVEC via FcγRII receptors, associated with apoptosis and mTOR-dependent activation of autophagy. Testing of IVIG toxicity in endothelial cells prestimulated with proinflammatory cytokines is relevant to clinical conditions. Our results warrant further studies on endothelial toxicity of sub-visible immunoglobulin particles.

59 BASIC BIOLOGICAL SCIENCES↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

Comparative Performance Evaluation of Large Language Models for Extracting Molecular Interactions and Pathway Knowledge

Understanding the interactions and regulatory relationships among biomolecules is essential for deciphering complex biological systems and elucidating the mechanisms behind diverse biological functions. Traditionally, the collection of such molecular interaction data has relied on expert curation, a process that is both time-consuming and labor-intensive. To address these limitations, this study explores the use of large language models (LLMs) to automate the genome-scale extraction of molecular interaction knowledge. Here, we evaluate the performance of various LLMs on key biological tasks, including the identification of protein-protein interactions, detection of genes associated with pathways influenced by low-dose radiation, and inference of gene regulatory relationships. Our findings demonstrate that larger LLMs tend to perform better, particularly in extracting intricate gene and protein interactions. Despite their strengths, these models face challenges in recognizing functionally diverse gene groups and highly correlated regulatory relationships. Through a comprehensive analysis using established molecular interaction and pathway databases, we show that LLMs possess the potential to identify relevant biomolecules and predict their interactions, offering valuable insights and marking a significant step toward AI-driven biological knowledge discovery.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Assessing the evolution of research topics in a biological field using plant science as an example

Scientific advances due to conceptual or technological innovations can be revealed by examining how research topics have evolved. But such topical evolution is difficult to uncover and quantify because of the large body of literature and the need for expert knowledge in a wide range of areas in a field. Using plant biology as an example, we used machine learning and language models to classify plant science citations into topics representing interconnected, evolving subfields. The changes in prevalence of topical records over the last 50 years reflect shifts in major research trends and recent radiation of new topics, as well as turnover of model species and vastly different plant science research trajectories among countries. Our approaches readily summarize the topical diversity and evolution of a scientific field with hundreds of thousands of relevant papers, and they can be applied broadly to other fields.

60 APPLIED LIFE SCIENCES↗

Quantifying microbial roles in environmental iron oxidation via an integrated kinetics, `omics and metabolic modeling study (Final Report)

Iron oxyhydroxides are extremely reactive components of environmental systems, and therefore exert a strong influence on biogeochemical cycles. These oxyhydroxides strongly adsorb many biologically-relevant elements, including organic carbon and phosphate, as well as a wide range of metals including uranium and actinide species. Thus, the formation mechanism of iron oxyhydroxides are key to understanding both nutrient and contaminant cycling. Microorganisms can catalyze iron oxidation and promote the formation of Fe biominerals and thus are increasingly recognized as important players in biogeochemical cycling. However, it is completely unknown how much of environmental iron oxidation is biologically mediated versus abiotic, and various challenges in studying microbial iron oxidation have hindered accurate incorporation into hydrobiogeochemical models. The overarching goal of our work was to quantify and constrain microbial iron oxidation rates and use ‘omics to gain insight into the controls on this process, while developing tools to enable integration of biotic iron oxidation into hydrobiogeochemical models. Our work focused on the Savannah River Site (SRS) in South Carolina, where extensive microbial iron oxidation has been observed. At Tims Branch, part of the Argonne National Laboratory Wetland Hydrobiogeochemistry Science Focus Area (Argonne SFA), where groundwater discharges into a stream, iron-oxidizing microbial mats form and appear to be a major sink of uranium. In the wetlands that surround Tims Branch, there are wide swaths of iron microbial mats and flocs (mobilized mat). We measured biotic and abiotic iron oxidation rates using mats and water sampled from these sites and found that iron oxidation is primarily carried out by chemolithotrophic microorganisms. The resulting rate constants can be incorporated into models. These mats were characterized by metagenomics and metatranscriptomics, which showed that aerobic chemolithotrophs were the dominant iron-oxidizing bacteria (FeOB), and these included Gallionellaceae and Leptothrix, and possibly Rhodoferax, which is known as an Fe-reducer but may also oxidize Fe(II). This demonstrated that diverse FeOB can coexist and suggests that there are a range of niches and therefore drivers of chemolithotrophic iron oxidation. Analysis of reconstructed genomes strongly suggests that a major factor in diversity is carbon source, as genomes contained varied pathways for autotrophy and heterotrophy. We performed an in-depth analysis of Leptothrix ochracea genomes, since this sheath-former is one of the primary mat builders, yet its physiology remained unresolved. A combination of genomics, transcriptomics, and metabolic modeling suggest that L. ochracea grows mixotrophically using a combination of Fe(II) and organics for energy and both inorganic and organic carbon to create biomass. This contrasts with the largely autotrophic Gallionellaceae (Gallionella, Sideroxydans, and Ferriphaselus) also present in the mats and flocs. Remarkably, multiple FeOB, both Leptothrix and Gallionellaceae, showed activity in response to Fe(II) in live mat incubations. We tracked the gene expression of individual MAGs to Fe(II) and found that various autotrophic and heterotrophic FeOB responded to Fe(II), increasing expression of both carbon fixation and organic utilization genes. The results of the integrated field, kinetics, and omics studies give detailed insight into 1) the taxa that oxidize Fe, and 2) how they connect Fe, C, and N cycles. Towards the goal of connecting omics data to hydrobiogeochemical models, we worked with the KBase team to create a template metabolic model for chemolithotrophic iron oxidation. We initially modeled the well-characterized isolate Gallionellaceae Sideroxydans lithotrophicus, and also applied the model to the mixotroph L. ochracea. In all, we have characterized diverse FeOB in a representative wetland system and solved key problems that enable better incorporation of iron-oxidizing microbes into hydrobiogeochemical models.

54 ENVIRONMENTAL SCIENCES↗