Search NASASearch

SEARCH · Search NASA

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua

A genomic view of Earth’s biomes

Microorganisms are essential to all life on Earth through critical roles in key biological processes and diverse interactions with other organisms that shape ecosystems, drive biogeochemical cycles and influence both human health and environmental health. High-throughput sequencing from environmental samples has revolutionized the understanding of microbial diversity and functions. With vast amounts of genomes now available across Earth’s biomes, these data provide a blueprint of microbial life that can be harnessed for a more holistic understanding of microbiome structure and function across the various ecosystems on Earth. Here we review the application of genome-centric approaches, including recent advances in single-cell sequencing and functional profiling, to survey microbial and viral diversity. Furthermore, we highlight some of the most impactful evolutionary and functional discoveries, explore the spatial diversity and temporal dynamics of microorganisms across diverse environments, and discuss genome-enabled insights into host-associated microorganisms.

Ecology

Bleach Rescues Nannochloropsis from an Obligate Parasite and Alters Microbial and Metabolite Signatures of Outdoor Cultures

Chemical agents are commonly used to protect algal crops. Yet, few studies have characterized the effects of these agents on associated microbial communities to understand effects on microbial functions relevant to algal crop production and protection. Here, we used shotgun metagenomic sequencing and untargeted exometabolite profiling to link the application of bleach, a -cidal agent used to protect algae from pests, to changes in community composition, metabolic pathways, and exometabolies - at a whole community level. Bleach protected the algal crop from crashing but altered bacterial diversity. Analysis of metagenome-assembled genomes (MAGs) revealed a classic predator-prey cycle between Oligoflexus and our target alga Nannochloropsis. Olifoflexus genomes from our study were notably similar to a previously identified BALO (Bdellovibrio and like organism), FD111, known to kill Nannochloropsis cultures, providing strong evidence that an FD111-like organism was responsible for the crash. Metabolic pathway composition differed between bleached and unbleached ponds, with abundance of twelve pathways related to stress tolerance, including the superpathway of methylglyoxal degradation, lipid IVA biosynthesis, and ectoine biosynthesis, greater in bleached ponds compared to unbleached ponds. Virulence factors related to adherence, biofilm formation, motility, and pathogenicity increased dramatically in bleached ponds with time, although this increase was not coupled with an increase in pathogens - algal or otherwise - or a decline in algal health. Our study highlights the importance of coupling 16S rRNA gene sequencing with whole genome data and other -omics tools to sketch a larger picture of community structure and function in crop systems. Moreover, our results highlight that continued long-term bleaching may lead to negative effects to crop health or downstream adverse health effects to humans or animals, depending on the algal product (i.e. human supplements or animal feedstocks). Future work on alternative treatment methods that would reduce resistance is necessary in the field.

09 BIOMASS FUELS

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)

From microbial diversity to functional potential using dimensionality reduction

The high dimensionality of microbial diversity data from ‘omics observations can be reduced using Machine Learning, with many recent studies showcasing ML utility for exploratory ecological feature finding and process prediction. Here, we compare the Self Organizing Map (SOM) dimensionality reduction method to the well-documented sample-based Principal Coordinate Analysis (PCoA) and taxa-based Weighted Gene Correlation Network Analysis (WGCNA) using near daily 16S rRNA gene amplicon sequencing data from the 2019 to 2020 MOSAiC International Arctic Drift Expedition. We then map k-means clustering outputs from each method to available metagenomes, extracting functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method better represented expected seasonal transitions and identified a greater number of metabolically distinct functional groups than the more traditional PCoA ordination. Ultimately, we identified four community ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover, highlighting the importance of succession in functional diversity for the central Arctic Ocean. These results reinforce ML dimensionality reduction as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and potentially provide ’omics informed ecotype diversity to leverage in mechanistic biogeochemical models.

Arctic Ocean

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto

Cyote-attack Chain Estimator

Attack Chain Estimator (ACE) Application Overview The Attack Chain Estimator (ACE) Application is a sophisticated tool designed for the ingestion, classification, sequencing, and enrichment of cybersecurity threat reports. This application leverages advanced machine learning models and extensive historical data to provide comprehensive insights into cyber threats, specifically targeting Industrial Control Systems (ICS). Purpose The primary functions of the ACE Application include: Ingestion of Cybersecurity Threat Reporting: Capable of ingesting text-based threat reports in markdown or text file format. Supports ingestion of structured data from other sources in STIX/JSON format. Classification of Report’s Text-Based Events: Utilizes a DeBERTa classifier, specifically trained on cybersecurity data, to map the events to MITRE ATT&CK for ICS Tactics and Techniques. Classification is performed using multiple Jupyter notebooks and machine learning workflows hosted as FastAPI microservices: regex_data deberta_base_35_train_hft_classifier_mlflow.ipynb hft_regex_classifier_mlflow.ipynb param_train_hft_classifier_mlflow.ipynb regex_tactic_tech.ipynb Ordering of Tactics, Techniques, and Observable Events: Sequences the identified tactics, techniques, and events to form a coherent attack chain. Enrichment with Historical Attack Chain Details: Enhances the attack chain with details from historical attacks using a Markov model developed from CyOTE Precursor Analysis Report data. The Markov model is available as a FastAPI endpoint for seamless integration. Enrichment with Adversary Emulation Capabilities Data: Integrates adversary emulation capabilities data using MITRE Caldera for OT adversary abilities UUIDs. Export of Output Files: Provides options to export the enriched attack chain in JSON or CSV formats. Routing of Output to Other Applications: Facilitates routing of output to various platforms and applications, including: Threat Intelligence Platforms COREII Scout for Threat Intelligence Analysis COREII Modeling and Simulation for Adversary Emulation Technical Description The ACE Application is an advanced cybersecurity tool designed to provide detailed threat analysis and sequence generation. It is built on a robust architecture that integrates natural language processing, machine learning, and historical data modeling. Key Components: Data Ingestion Module: Handles the input of threat reports and data from various formats, ensuring flexibility in data sources. Classification Engine: Employs DeBERTa-based classifiers hosted as FastAPI microservices to analyze and classify threat report events in accordance with the MITRE ATT&CK framework for ICS. Sequence Generator: Orders the classified events into a logical attack chain, providing clear insight into the sequence of tactics and techniques used in the threat. Enrichment Engine: Integrates historical data and adversary emulation capabilities to enhance the attack chain with valuable context and additional details. The historical data enrichment is powered by a Markov model, which is available as a FastAPI endpoint. Export and Routing Module: Facilitates the export of the enriched attack chain in multiple formats and routes the output to designated applications for further analysis or emulation.

Paul, Tony [Idaho National Laboratory (INL), Idaho

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

DFT-based insight into finite-temperature properties of ferroelectric perovskites with lone-pair: the case of CsGeX 3 (X = Cl, Br, I)

Ferroelectrics remain in the focus of scientific attention for decades owing to their fundamental and practical appeal. Recently, ferroelectricity has been demonstrated in semiconducting halide perovskites (Zhang et al 2022 Sci. Adv. 8 eabj5881), offering both a rare combination of ferroelectricity and semiconductivity in the same material and a possible alternative to the prevailing perovskite oxide ferroelectrics. We propose a route to simulating such materials at finite temperatures capable of reproducing key experimental and first-principle data, such as Curie temperature, phase transition sequence, spontaneous polarization, and soft mode frequencies. The key methodological finding is the superior performance of hybrid exchange correlation functionals in parametrization of effective Hamiltonians for ferroelectrics with lone pair. The parametrization for effective Hamiltonians for CsGeX 3 (X = Cl, Br, I) is reported. The application of methodology to study polarization reversal in CsGeX 3 allows for the development of a ‘minimalistic’ model for polarization reversal in ferroelectrics that provides an insight into the mechanisms of polarization reversal and its key features, such as the relationship between the coercive field, temperature, and AC field frequency. Importantly, the model reveals the origin of the well-known and ever-puzzling overestimation of coercive fields in computations. Furthermore, we report a variety of finite-temperature properties of CsGeX 3 ferroelectrics, such as dielectric susceptibility, pyroelectric coefficients, and energy storage density, which reveal that these halide perovskites possess properties comparable to their oxide counterparts. Here, we believe that our work provides significant methodological advancements, deepens fundamental understanding of ferroelectrics, and reveals the potential of halide perovskite ferroelectrics.

effective Hamiltonian

Benchtop Autonomous Electrochemical Characterization System for Combinatorial Thin-Film Solid Oxide Electrodes

The design of materials for electrochemical energy conversion is complicated by a vast search space of candidate materials and multifaceted property requirements: multicarrier conductivity, stability, and catalytic activity are all necessary but rarely intersect. Although self-driving laboratories are rapidly rising to address such material optimization problems, the required infrastructure for integrated, large-scale robotic facilities can be cost-prohibitive. Here we develop and evaluate a closed-loop measurement system for efficient screening of proton-conducting oxide electrodes for ceramic fuel cells and electrolyzers, building on top of an existing benchtop instrument and integrating techniques for rapid impedance measurement and automated analysis. This system exemplifies a “minimum viable” self-driving implementation that can deliver substantial benefits with relatively simple infrastructure. Combinatorial thin-film microelectrode libraries are characterized with a recently developed joint time-domain and frequency-domain impedance measurement technique, which provides an order-of-magnitude acceleration relative to conventional impedance spectroscopy. The distribution of relaxation times is extracted from impedance data and analyzed without human intervention. These results feed an active learning and Bayesian optimization process that learns to predict electrochemical impedance as a function of material composition, measurement temperature, oxygen partial pressure, and electrical bias, which further reduces the screening time by tenfold with optimized experimental sequences. We apply this system to Ba⁡(Co,Fe,Zr,Y)⁢O 3−𝛿 combinatorial libraries and evaluate its effectiveness for learning material property trends and optimizing expensive-to-evaluate properties such as activation energy. This offers insights into key methodological aspects of practical autonomous experimentation, including surrogate model validation, cost-aware acquisition functions, and high-throughput data interpretation. Our results demonstrate the efficacy of the system for rapidly gathering information, but also highlight real-world experimental challenges of thin-film degradation and numerical instability in surrogate models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N

Statistical estimates of the binary properties of rotational variables

ABSTRACT We present a model to estimate the average primary masses, companion mass ranges, the inclination limit for recognizing a rotational variable, and the primary mass spreads for populations of binary stars. The model fits a population’s binary mass function distribution and allows for a probability that some mass functions are incorrectly estimated. Using tests with synthetic data, we assess the model’s sensitivity to each parameter, finding that we are most sensitive to the average primary mass and the minimum companion mass, with less sensitivity to the inclination limit and little to no sensitivity to the primary mass spread. We apply the model to five populations of binary spotted rotational variables identified in ASAS-SN, computing their binary mass functions using RV data from APOGEE. Their average primary mass estimates are consistent with our expectations based on their CMD locations ($\sim 0.75 \, {\rm M}_{\odot }$ for lower main sequence primaries and $\sim 0.9$–$1.2 \, {\rm M}_{\odot }$ for RS CVn and sub-subgiants). Their companion mass range estimates allow companion masses down to $M_2/M_1\simeq 0.1$, although the main sequence population may have a higher minimum mass fraction ($\sim 0.4$). We see weak evidence of an inclination limit $\gtrsim 50^{\circ }$ for the main sequence and sub-subgiant groups and no evidence of an inclination limit in the other groups. No groups show strong evidence for a preferred primary mass spread. We conclude by demonstrating that the approach will provide significantly better estimates of the primary mass and the minimum mass ratio and reasonable sensitivity to the inclination limit with 10 times as many systems.

Phillips, Anya (ORCID:000900051914974X)

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES

Response of Subsurface Nitrogen-Cycling Microbial Communities to Environmental Fluctuations (Final Technical Report)

Riparian floodplains are dynamic ecosystems linking terrestrial and riverine systems. These floodplains experience hydrological shifts such as changes in water table height, flooding, and drought and can be ‘hotspots’ of biogeochemical cycling due to shifting sediment moisture (and saturation) and subsurface exchanges of water, nutrients, and other compounds across different sediment layers. Subsurface microbial communities are the primary drivers of biogeochemical processes in floodplains, and thus their structure and function can directly influence both surface and groundwater quality. The microbial nitrogen (N) cycle is particularly important in floodplains as it affects nutrient availability and removal. Two functional guilds of chemoautotrophic (i.e. CO2-fixing) microorganisms are responsible for the first oxidative step of the N cycle, nitrification: ammonia-oxidizing archaea (AOA) and bacteria (AOB) catalyze the oxidation of ammonia to nitrite, while nitrite-oxidizing bacteria (NOB) oxidize nitrite to nitrate. Despite the critical role nitrification plays in N-cycling in both terrestrial and aquatic ecosystems, our understanding of the diversity, ecophysiology, and activity of nitrifying organisms in subsurface floodplain soils/sediments is extremely limited. To help address this critical knowledge gap, the overarching goal of this project was to determine how shifts in key environmental parameters and gradients impact microbial N-cycling communities/processes, with particular emphasis on nitrification, within hydrologically-variable floodplain sediments in the Wind River Basin near Riverton, Wyoming. The three specific objectives of this project were to: (1) to associate in situ environmental drivers of N cycling with distinct functional guilds; (2) determine the guild response to variation in key ecosystem drivers; and (3) develop a dynamic ecosystem model of the microbial N cycle with the Riverton subsurface using community genomic and biogeochemical data collected in the first two objectives. Over the course of this project, we employed both 16S rRNA gene amplicon sequencing and genome-resolved metagenomics to examine the phylogenetic diversity and metabolic potential of subsurface nitrifier communities within 68 samples collected across multiple sites, depths, and time points within the Riverton floodplain, allowing for both spatial and temporal investigations at different scales. This project benefitted tremendously from recent advances in high-throughput sequencing technologies coupled with dramatic improvements in the computational tools and algorithms available for analyzing such large, complex genomic datasets. By pairing these cutting-edge genomic approaches with depth-resolved sampling and detailed geochemical analyses of the Riverton floodplain, we have gained novel insights into the structure and function of subsurface nitrifier communities in relation to both hydrology and biogeochemistry. This project resulted in the most detailed and comprehensive characterization of N-cycling floodplain microbial communities to date and will hopefully inspire and pave the way for future studies using similar approaches in other floodplains. Indeed, such information is critical for understanding subsurface biogeochemical cycling and how elemental stores are altered from perturbations initiated by the water cycle within floodplains. Finally, because of the terrestrial-aquatic nature of the Riverton floodplain, results from this project are also of relevance to disciplines such as soil science, estuarine science, limnology & oceanography, biogeochemistry, geobiology, environmental engineering, as well as genomics and data science.

54 ENVIRONMENTAL SCIENCES

Flaring Stars in a Non-targeted mm-wave Survey with SPT-3G

We present a flare star catalog from four years of non-targeted millimeter-wave survey data from the South Pole Telescope (SPT). The data were taken with the SPT-3G camera and cover a 1500-square-degree region of the sky from $20^{h}40^{m}0^{s}$ to $3^{h}20^{m}0^{s}$ in right ascension and $-42^{\circ}$ to $-70^{\circ}$ in declination. This region was observed on a nearly daily cadence from 2019-2022 and chosen to avoid the plane of the galaxy. A short-duration transient search of this survey yields 111 flaring events from 66 stars, increasing the number of both flaring events and detected flare stars by an order of magnitude from the previous SPT-3G data release. We provide cross-matching to Gaia DR3, as well as matches to X-ray point sources found in the second ROSAT all-sky survey. We have detected flaring stars across the main sequence, from early-type A stars to M dwarfs, as well as a large population of evolved stars. These stars are mostly nearby, spanning 10 to 1000 parsecs in distance. Most of the flare spectral indices are constant or gently rising as a function of frequency at 95/150/220 GHz. The timescale of these events can range from minutes to hours, and the peak $\nu L_{\nu}$ luminosities range from $10^{27}$ to $10^{31}$ erg s$^{-1}$ in the SPT-3G frequency bands.

79 ASTRONOMY AND ASTROPHYSICS