Search NASA⌕ Search

SEARCH · Search NASA

Results for “genome sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Nanoscale Bio-engineering Solutions for Space Exploration: The Nanopore Sequencer

Characterization of biological systems at the molecular level and extraction of essential information for nano-engineering design to guide the nano-fabrication of solid-state sensors and molecular identification devices is a computational challenge. The alpha hemolysin protein ion channel is used as a model system for structural analysis of nucleic acids like DNA. Applied voltage draws a DNA strand and surrounding ionic solution through the biological nanopore. The subunits in the DNA strand block ion flow by differing amounts. Atomistic scale simulations are employed using NASA supercomputers to study DNA translocation, with the aim to enhance single DNA subunit identification. Compared to protein channels, solid-state nanopores offer a better temporal control of the translocation of DNA and the possibility to easily tune its chemistry to increase the signal resolution. Potential applications for NASA missions, besides real-time genome sequencing include astronaut health, life detection and decoding of various genomes.

Stolc, Viktor↗

Nanoscale Bioengineering Solutions for Space Exploration the Nanopore Sequencer

Characterization of biological systems at the molecular level and extraction of essential information for nano-engineering design to guide the nano-fabrication of solid-state sensors and molecular identification devices is a computational challenge. The alpha hemolysin protein ion channel is used as a model system for structural analysis of nucleic acids like DNA. Applied voltage draws a DNA strand and surrounding ionic solution through the biological nanopore. The subunits in the DNA strand block ion flow by differing amounts. Atomistic scale simulations are employed using NASA supercomputers to study DNA translocation. with the aim to enhance single DNA subunit identification. Compared to protein channels, solid-state nanopores offer a better temporal control of the translocation of DNA and the possibility to easily tune its chemistry to increase the signal resolution. Potential applications for NASA missions, besides real-time genome sequencing include astronaut health, life detection and decoding of various genomes. http://phenomrph.arc.nasa.gov/index.php

Ioana, Cozmuta↗

Enrichment of root-associated Streptomyces strains in response to drought is driven by diverse functional traits and does not predict beneficial effects on plant growth

The genus Streptomyces has consistently been found enriched in drought-stressed plant root microbiomes, yet the ecological basis and functional variation underlying this enrichment at the strain and isolate level remain unclear. Using two 16S rRNA sequencing methods with different levels of taxonomic resolution, we confirmed drought-associated enrichment (DE) of Streptomyces in field-grown sorghum roots and identified five closely related but distinct amplicon sequence variants (ASVs) belonging to the genus with variable drought enrichment patterns. From a culture collection of sorghum root endophytes, we selected 12 Streptomyces isolates representing these ASVs for phenotypic and genomic characterization. Whole-genome sequencing revealed substantial variation in gene content, even among closely related isolates, and exometabolomic profiling showed distinct metabolic responses to media supplemented with drought- versus well-watered root tissue. Traits linked to drought survival, including osmotic stress tolerance, siderophore production, and carbon utilization, varied widely among isolates and were not phylogenetically conserved. Using a broader panel of 48 Streptomyces, we demonstrate that DE scores, determined through mono-association experiments in gnotobiotic sorghum systems, showed high variability and lacked correlation with plant growth promotion. Pangenome-wide association identified orthogroups involved in osmolyte transport (e.g., proP) and membrane biosynthesis (e.g., fabG) as positively associated with DE, though most associations lacked phylogenetic signal. Collectively, these results demonstrate that Streptomyces DE is not a conserved genus-level trait but is instead strain-specific and functionally heterogeneous. Furthermore, DE in the root microbiome was shown not to predict beneficial effects on plant growth. This work underscores the need to resolve functional traits at the strain level and highlights the complexity of microbe-host-environment interactions under abiotic stress.

Fonseca-Garcia, Citlali↗

Characterization of Multiple Trichloroethene, cis-Dichloroethene and 1,1-Dichloroethene Degrading Propanotrophic Communities

Aerobic cometabolism offers a viable strategy for the remediation of chlorinated solvent plumes at oxic sites where anaerobic approaches are limited. In this study, propane-enriched mixed cultures (derived from agricultural soils and an impacted site sediment) which previously degraded 1,4-dioxane, were evaluated for their capacity to also degrade trichloroethene (TCE), cis-1,2-dichloroethene (cDCE), and 1,1-dichloroethene (1,1-DCE) over successive transfers. Sustained biodegradation of TCE and cDCE was observed across multiple enrichments, and cultures enriched on one compound generally degraded the other. In contrast, 1,1-DCE biodegradation was restricted to a subset of cultures and removal times increased over transfers. Further, 1,1-DCE removal was absent at elevated concentrations, both trends consistent with inhibitory or toxic effects. Whole genome sequencing analyses revealed pronounced substrate-dependent selection of microbial communities, with cDCE-degrading cultures being dominated by Mycobacterium and Mycolicibacterium, whereas TCE-degrading cultures were dominated by Rhodococcus. Rhodococcus metagenome-assembled genomes (MAGs) in the TCE degrading cultures classified as R. opacus or R. wratislaviensis. 1,1-DCE degrading cultures were dominated by Pseudonocardia, although the associated MAGs contained a truncated propane monooxygenase alpha subunit. Functional gene analysis identified both group 5 (prmABCD) and putative group 6 propane monooxygenases. The following KBase narratives contain the quality controlled reads, MAGs (fasta assemblies) and the prokka annotations for each assembly TCE Site 1A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254918) TCE Soil 2A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254919) TCE Soil T3 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254920) TCE Soil T4 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254921) cDCE Site 1A 1B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254915) cDCE Soils T2 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254927) cDCE Soil T3 A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254928) cDCE Soil 4A and B Propanotrophic MAGs (https://narrative.kbase.us/narrative/254942) 1,1-DCE T2 T3 Propanotrophic MAGs (https://narrative.kbase.us/narrative/254903)

59 BASIC BIOLOGICAL SCIENCES↗

Identification of shared viral sequences in peat moss metagenomes reveals elements of a possible Sphagnum core virome

Viruses are an understudied component of plant microbiomes. Identifying viruses that are shared between individual plants, or members of the “core virome”, could reveal stable viral populations with the potential to modulate the composition and function of the microbiome. Here, we examined the virome associated with Sphagnum mosses, a keystone species that has direct influence over the fate of peatland carbon stores. We analyzed bulk metagenomes and metatranscriptomes generated from Sphagnum field samples collected over a ten-month period to identify virus-like sequences shared among plants. Individual Sphagnum samples harbored distinct DNA and RNA viromes where only a small percentage (< 1%) of the total number of identified viral contigs were shared among all samples. Based on taxonomic classification, the shared viral contigs represent bacterial viruses, or phage (Caudoviricetes), as well as viruses of eukaryotes, namely nucleocytoplasmic large DNA viruses (Nucleocytoviricota) and RNA viruses (Riboviria). We linked the shared phage-like contigs to viral regions within sequenced genomes of bacterial taxa that are members of the Sphagnum core microbiome, suggesting that these contigs represent temperate phage or degraded prophage. The putative nucleocytoplasmic large DNA viruses and RNA viruses were phylogenetically diverse and showed sequence similarity to viruses associated with a broad range of hosts and environmental sources. The identification of shared viral contigs suggested that, despite the compositional heterogeneity between samples, Sphagnum mosses may harbor a core virome. Future work validating the presence of the core virome is warranted as it may aid in understanding how persistent viruses impact microbiome ecology and symbiont evolution within this climatically relevant keystone species.

Metagenomics↗

Luteolibacter sp. strain Populi

Luteolibacter sp. strain Populi is bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6 Mb chromosome was completely sequenced using Oxford Nanopore long-reads and is predicted to encode 5301 proteins and 60 RNAs. The bacteria was isolated from the rhizosphere of a mature Populus trichocarpa from the Tieton riverwatershed of Washington state, USA (Lat: 46°42’9” N, Lon: 120°25 39’36” W). A rhizosphere sample (fine roots and adhering soil) was used to obtain a microbial fraction by centrifugation on Histodenz (12) and stained with 5µM Syto59 (Thermo Fisher Scientific Inc). A Cytopeia Influx cell sorter (BD, Franklin Lakes, NJ) was used to sort and array single cells (100 per plate) based on forward-side scatter and fluorescence intensity on asparagine-glucose nutrient agar (ATCC medium 184). The Luteolibacter sp. Populi genome sequence has been deposited in GenBank under the accession number CP161812. A draft genome annotated with Prokka and DRAM is available in this Narrative as Luteolibacter_sp_Prokka.240711.

59 BASIC BIOLOGICAL SCIENCES↗

Determining divergence times with a protein clock: update and reevaluation

A recent study of the divergence times of the major groups of organisms as gauged by amino acid sequence comparison has been expanded and the data have been reanalyzed with a distance measure that corrects for both constraints on amino acid interchange and variation in substitution rate at different sites. Beyond that, the availability of complete genome sequences for several eubacteria and an archaebacterium has had a great impact on the interpretation of certain aspects of the data. Thus, the majority of the archaebacterial sequences are not consistent with currently accepted views of the Tree of Life which cluster the archaebacteria with eukaryotes. Instead, they are either outliers or mixed in with eubacterial orthologs. The simplest resolution of the problem is to postulate that many of these sequences were carried into eukaryotes by early eubacterial endosymbionts about 2 billion years ago, only very shortly after or even coincident with the divergence of eukaryotes and archaebacteria. The strong resemblances of these same enzymes among the major eubacterial groups suggest that the cyanobacteria and Gram-positive and Gram-negative eubacteria also diverged at about this same time, whereas the much greater differences between archaebacterial and eubacterial sequences indicate these two groups may have diverged between 3 and 4 billion years ago.

NASA Discipline Exobiology↗

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag↗

The genome of the polyextremophilic yeast, Naganishia friedmannii, reveals adaptations involved in stress response pathways, carbohydrate metabolism expansion, and a limited DNA repair repertoire

Here we report the draft genome sequence of Naganishia friedmannii (formerly Cryptococcus friedmannii) isolate, a Basidiomycota yeast commonly found in some of the most extreme environments of the Earth's cryosphere. We isolated N. friedmannii strain Llullensis from soils at 6000 m above sea level on Volcán Llullaillaco, Argentina. The genome was 22.2 Mb with 6251 identified protein coding genes. Proteins known to be associated with thermal, osmotic, and radiation stress were identified in the genome. Comparative analysis with seven other Naganishia genomes revealed unique features underlying its polyextremophilic lifestyle. Naganishia friedmannii showed an expansion of genes involved in breaking down plant-derived carbohydrates, supporting the hypothesis that it survives at high elevations by metabolizing wind-deposited organic matter. Surprisingly, many genes involved in cell-cycle checkpoints and DNA repair were missing, as in several other Naganishia species. This extensive loss may be adaptive in extreme environments prone to abiotic stress, where a high mutation rate could generate advantageous traits, and reduced cell-cycle control may allow for faster reproduction that would be advantageous for rapid growth during brief periods of soil wetting following rare snow events.

Vimercati, Lara↗

Comparative Analysis of DNA LLM Classification Techniques Using Intra-Layer Feature Extraction with Autoencoder Stacks [Poster]

This project conducts a comparative analysis of DNA LLM classification techniques using Evo2, Grover, and UTRML, focusing on intra-layer feature extraction in Evo2. By extracting features from multiple layers of Evo2 and integrating them into an autoencoder stack with a binary classification head, we evaluate its effectiveness in classifying genomic sequences compared to smaller DNA language models. My findings demonstrate that Evo2 outperforms Grover and UTRML in classification accuracy on a dataset provided by department 08625, CAO2021, while UTRML offers competitive performance with lower computational costs. This study highlights the potential of advanced embedding techniques in enhancing genomic data analysis and informs future research in bioinformatics.

59 BASIC BIOLOGICAL SCIENCES↗

Developing a Genetic Variant Calling Pipeline for Quantifying the Complex Mutagenic Load Accumulated in BioNutrients-1 Production Pack Samples

Microorganisms hold great promise for on demand production of labile nutrients and pharmaceuticals as well recycling and in situ resource utilization. The utilization of microorganisms for such tasks on space missions is hindered by the limited data on how microbes respond to spaceflight. For example, the genetic stability of microorganisms, and the genomic engineered traits added to deliver desired functions, over long-term storage in the spacecraft environment is poorly understood. The BioNutrients-1 (BN-1) mission conducted a 5-year study of desiccated storage in Low Earth Orbit (LEO) to evaluate the suitability of eight synthetic biology chassis organisms for long-duration space missions. We are employing high-depth, whole genome sequencing (WGS) to determine the mutagenic load that accumulated during long-term storage. Mutation analysis pipelines are well established for homogenous culture grown from a single colony, but the mutational landscape of the BN-1 samples present a unique analysis challenge, as every cell in the BN-1 samples had a unique genetic journey of DNA damage and repair. Consequently, sequence variants are expected at low allele frequency within samples. To address this genetic complexity, we apply two distinct computational approaches to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO. For reference genome free mutation detection, we utilized DiscoSNP++, which is a de Bruijn graph approach. For reference genome-based mutation detection we utilize GATK for Microbes, which is a Bayesian probabilistic approach. We will benchmark these approaches against the mutations originally identified using CRISP, a method optimized for pooled samples. Ultimately, quantifying the mutation load imposed by storage or growth on the ISS will help identify chassis organisms with both high levels of genome stability and viability, which are desirable traits for implementation of bioproduction in long-duration missions.

SNP↗

Myco-Ed: Mycological curriculum for education and discovery

Fungi are important and hyperdiverse organisms, yet chronically understudied. Most fungal clades have no reference genomes, impeding our understanding of their ecosystem functions and use as solutions in health and biotechnology. Also, opportunities for training in fungal biology and genomics are lacking, creating a bottleneck that hinders the recruitment and cultivation of a talented future mycological workforce. To address these issues, we developed Myco-Ed, an educational program offering training and scientific contributions through genome sequencing and analysis. Myco-Ed empowers students to pursue careers in fungal biology while improving fungal resources. Myco-Ed has been piloted at 12 institutions (15 classrooms) ranging from online e-Campuses to R1 universities, resulting in hundreds of fungal observations and many new high-quality reference genomes.

Branco, Sara↗

Microbial Characteristics of ISS Environmental Surfaces

The microbiome of environmental surfaces from the International Space Station were characterized in order to examine the relationship to crew and hardware maintenance. The Microbial Observatory (ISS-MO) experiment generated a microbial census of ISS environments using advanced molecular microbial community analyses along with traditional culture-based methods. Since the “omics” methodologies generated an extensive microbial census, significant insights into spaceflight-induced changes in the populations of beneficial and/or potentially harmful microbes were gained. Surface samples were collected from several ISS surface locations from three flight opportunities, and were returned to Earth via the Soyuz TMA-14M or the Space X Dragon capsule. In addition to cultivation methods, viable microbial burden, iTag-based sequencing, and metagenome analyses were carried out. The cultivable microbial bioburden differed by location and sampling event. Exploring the ISS environmental microbiome revealed presence of opportunistic pathogens and antibiotic resistant microbes. Genes involved in ATP binding cassette transporters, two component systems, and beta-lactam resistance were among a diverse set of metabolic and genetic information processing pathways. Whole genome sequencing (WGS) of 50 ISS strains exhibiting resistance to various antibiotics was carried out. The antibiotic resistant genes deduced from the WGS were compared with the resistomes generated directly from the gene pool of the environmental samples. Two unique Aspergillus fumigatus strains isolated from the ISS were characterized and compared to the experimentally established clinical isolates Af293 and CEA10. A virulence assessment in a neutrophil-deficient larval zebrafish model of invasive aspergillosis indicated that both ISSFT-021 and IF1SW-F4 were significantly more lethal compared to Af293 and CEA10. The findings from this Environmental “Omics” project should be exploited to enhance human health and well-being of a closed system. In other words, the ISS-MO research aims to "translate" findings in fundamental research into medical practice (pathogen detection) and meaningful health outcomes (countermeasure development).

Perry, Jay↗

Leptothrix ochracea genomes reveal potential for mixotrophic growth on Fe(II) and organic carbon

ABSTRACT Leptothrix ochracea creates distinctive iron-mineralized mats that carpet streams and wetlands. Easily recognized by its iron-mineralized sheaths, L. ochracea was one of the first microorganisms described in the 1800s. Yet it has never been isolated and does not have a complete genome sequence available, so key questions about its physiology remain unresolved. It is debated whether iron oxidation can be used for energy or growth and if L. ochracea is an autotroph, heterotroph, or mixotroph. To address these issues, we sampled L. ochracea -rich mats from three of its typical environments (a stream, wetlands, and a drainage channel) and reconstructed nine high-quality genomes of L. ochracea from metagenomes. These genomes contain iron oxidase genes cyc2 and mtoA, showing that L. ochracea has the potential to conserve energy from iron oxidation. Sox genes confer potential to oxidize sulfur for energy. There are genes for both carbon fixation (RuBisCO) and utilization of sugars and organic acids (acetate, lactate, and formate). In silico stoichiometric metabolic models further demonstrated the potential for growth using sugars and organic acids. Metatranscriptomes showed a high expression of genes for iron oxidation; aerobic respiration; and utilization of lactate, acetate, and sugars, as well as RuBisCO, supporting mixotrophic growth in the environment. In summary, our results suggest that L. ochracea has substantial metabolic flexibility. It is adapted to iron-rich, organic carbon-containing wetland niches, where it can thrive as a mixotrophic iron oxidizer by utilizing both iron oxidation and organics for energy generation and both inorganic and organic carbon for cell and sheath production. IMPORTANCE Winogradsky's observations of L. ochracea led him to propose autotrophic iron oxidation as a new microbial metabolism, following his work on autotrophic sulfur-oxidizers. While much culture-based research has ensued, isolation proved elusive, so most work on L. ochracea has been based in the environment and in microcosms. Meanwhile, the autotrophic Gallionella became the model for freshwater microbial iron oxidation, while heterotrophic and mixotrophic iron oxidation is not well-studied. Ecological studies have shown that Leptothrix overtakes Gallionella when dissolved organic carbon content increases, demonstrating distinct niches. This study presents the first near-complete genomes of L. ochracea , which share some features with autotrophic iron oxidizers, while also incorporating heterotrophic metabolisms. These genome, metabolic modeling, and transcriptome results give us a detailed metabolic picture of how the organism may combine lithoautotrophy with organoheterotrophy to promote Fe oxidation and C cycling and drive many biogeochemical processes resulting from microbial growth and iron oxyhydroxide formation in wetlands.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular Identification of Microbial Contaminants

Microorganisms can have significant impacts on the success of NASA’s missions, including the integrity of materials, the reliability of scientific results, and maintenance of crew health. Robust cleaning and sterilization protocols are currently in place in NASA facilities, but agency experts agree that microbial contamination is unavoidable and its impact on NASA’s missions and science must be minimized. Therefore, it is critical to understand: 1) what specific microorganisms are present, 2) how they may impact scientific objectives, and 3) how to select appropriate mitigation strategies. The Marshall Space Flight Center (MSFC) Planetary Protection (PP) microbiology lab historically relied solely upon enumeration of culturable microbial contamination associated with spacecraft materials or cleanrooms. However, this process is time consuming, many microbes cannot be cultured, and very few can be identified with any fidelity using NASA standard microbiological methods. The work described in this white paper includes the establishment of molecular identification capabilities at MSFC, including DNA isolation, amplification, purification, and Sanger sequencing. This capability will not only improve planetary protection efforts at MSFC (i.e. by identifying contaminating microorganisms in cleanrooms or on spacecraft) but also offers a service center-wide for the identification of contaminants that arise in other projects, processing locations, or during set up and roll out of spacecraft. This work also lays the foundation for higher throughput efforts to identify large populations of microbes across the lifetime of a project and serves as the starting point for future work into whole genome sequencing, non-culture based methods, or additional characterization studies. Ultimately, accurate identification informs appropriate mitigation strategies, increasing the chances of success for NASA’s missions and objectives.

C. D. Cassilly↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗