Search NASA⌕ Search

SEARCH · Search NASA

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

342 records · Page 19

Characterization of Two Microbial Isolates from Andean Lakes in Bolivia

We are currently investigating the biological population present in the highest and least explored perennial lakes on earth in the Bolivian and Chilean Andes, including several volcanic crater lakes of more than 6000 m elevation, in combination of microbiological and molecular biological methods. Our samples were collected in saline lakes of the Laguna Blanca Laguna Verde area in the Bolivian Altiplano and in the Licancabur volcano crater (27 deg. 47 min S/67 deg. 47 min. W) in the ongoing project studying high altitude lakes. The main goal of the project is to look for analogies with Martian paleolakes. These Bolivian lakes can be described as Andean lakes following the classification of Chong. We have attempted to isolate pure cultures and phylogenetically characterize prokaryotes that grew under laboratory conditions. Sediment samples taken from the Licancabur crater lake (LC), Laguna Verde (LV), and Laguna Blanca (LB) were analyzed and cultured using enriched liquid media under both aerobic and anaerobic conditions. All cultures were incubated at room temperature (15 to 20 C) and under light exposure. For the reported isolates, 36 hours incubation were necessary for reaching optimal optical densities to consider them viable cultures. Ten serial dilutions starting from 1% inoculum were required to obtain a suitable enriched cell culture to transfer into solid media. Cultures on solid medium were necessary to verify the formation of colonies in order to isolate pure cultures. Different solid media were prepared using several combinations of both trace minerals and carbohydrates sources in order to fit their nutrient requirements. The microorganisms formed individual colonies on solid media enriched with tryptone, yeast extract and sodium chloride. Cells morphology was studied by optical and electronic microscopy. Rodshape morphologies were observed in most cases. Total bacterial genomic DNA was isolated from 50 ml late-exponential phase culture by using the CTAB miniprep protocol. The 16S rRNA genes were amplified by PCR using both Bacteria- and Archaeauniversal primer sets: 27f and 1492r, 21f and 1492r respectively. Sequences of 16S rRNA gene were determined and initially compared with reference sequences contained in the EMBL nucleotide sequence database by using the BLAST program and were subsequently aligned with 16S rRNA reference sequences in the ARB package (http://www.mikro.biologie.tu-muenchen.de). Aligned sequences were inserted within a stable phylogenetic tree by using the ARB parsimony tool. In this work we report the morphology and phylogenetic characterization of two isolates belonged to Laguna Blanca sediments.

Demergasso, C.↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

A new dynamical atmospheric ionizing radiation (AIR) model for epidemiological studies

A new Atmospheric Ionizing Radiation (AIR) model is currently being developed for use in radiation dose evaluation in epidemiological studies targeted to atmospheric flight personnel such as civilian airlines crewmembers. The model will allow computing values for biologically relevant parameters, e.g. dose equivalent and effective dose, for individual flights from 1945. Each flight is described by its actual three dimensional flight profile, i.e. geographic coordinates and altitudes varying with time. Solar modulated primary particles are filtered with a new analytical fully angular dependent geomagnetic cut off rigidity model, as a function of latitude, longitude, arrival direction, altitude and time. The particle transport results have been obtained with a technique based on the three-dimensional Monte Carlo transport code FLUKA, with a special procedure to deal with HZE particles. Particle fluxes are transformed into dose-related quantities and then integrated all along the flight path to obtain the overall flight dose. Preliminary validations of the particle transport technique using data from the AIR Project ER-2 flight campaign of measurements are encouraging. Future efforts will deal with modeling of the effects of the aircraft structure as well as inclusion of solar particle events. Published by Elsevier Ltd on behalf of COSPAR.

Aviation↗

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic fingerprints of the world’s soil ecosystems

Despite the explosion of soil metagenomic data, we lack a synthesized understanding of patterns in the distribution and functions of soil microorganisms. These patterns are critical to predictions of soil microbiome responses to climate change and resulting feedbacks that regulate greenhouse gas release from soils. To address this gap, we assay 1,512 manually curated soil metagenomes using complementary annotation databases, read-based taxonomy, and machine learning to extract multidimensional genomic fingerprints of global soil microbiomes. Our objective is to uncover novel biogeographical patterns of soil microbiomes across environmental factors and ecological biomes with high molecular resolution. We reveal shifts in the potential for (i) microbial nutrient acquisition across pH gradients; (ii) stress-, transport-, and redox-based processes across changes in soil bulk density; and (iii) greenhouse gas emissions across biomes. We also use an unsupervised approach to reveal a collection of soils with distinct genomic signatures, characterized by coordinated changes in soil organic carbon, nitrogen, and cation exchange capacity and in bulk density and clay content that may ultimately reflect soil environments with high microbial activity. Genomic fingerprints for these soils highlight the importance of resource scavenging, plant-microbe interactions, fungi, and heterotrophic metabolisms. Across all analyses, we observed phylogenetic coherence in soil microbiomes—more closely related microorganisms tended to move congruently in response to soil factors. Collectively, the genomic fingerprints uncovered here present a basis for global patterns in the microbial mechanisms underlying soil biogeochemistry and help beget tractable microbial reaction networks for incorporation into process-based models of soil carbon and nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

A metagenomic perspective on the microbial prokaryotic genome census

Following 30 years of sequencing, we assessed the phylogenetic diversity (PD) of >1.5 million microbial genomes in public databases, including metagenome-assembled genomes (MAGs) of uncultivated microbes. As compared to the vast diversity uncovered by metagenomic sequences, cultivated taxa account for a modest portion of the overall diversity, 9.73% in bacteria and 6.55% in archaea, while MAGs contribute 48.54% and 57.05%, respectively. Therefore, a substantial fraction of bacterial (41.73%) and archaeal PD (36.39%) still lacks any genomic representation. This unrepresented diversity manifests primarily at lower taxonomic ranks, exemplified by 134,966 species identified in 18,087 metagenomic samples. Our study exposes diversity hotspots in freshwater, marine subsurface, sediment, soil, and other environments, whereas human samples yielded minimal novelty within the context of existing datasets. These results offer a roadmap for future genome recovery efforts, delineating uncaptured taxa in underexplored environments and underscoring the necessity for renewed isolation and sequencing.

59 BASIC BIOLOGICAL SCIENCES↗

Fractionation of Filamentous Algae from Mixed Biofilms

Filamentous algae, which grow in long, hair-like filaments within biofilms, play a crucial role in wastewater treatment due to their ability to produce significant biomass and their resistance to predation compared to traditional microalgal treatments. These algae can effectively uptake and utilize pollutants, particularly excessive nitrogen (ammonia, nitrate, nitrite) and phosphorus (phosphate), making filamentous algae valuable for wastewater treatment, as well as bioethanol and biodiesel production due to high lipid productions. However, each algal species possesses different capacities, necessitating a thorough genetic identification and understanding of each community. A major challenge in accurately assessing these communities is the lack of coverage in large sequencing databases which can lead to misrepresentation of the true composition and abundance of organisms and overall sequencing bias. To address this, I evaluated chemical and physical techniques for separating filamentous algae from mixed biofilms to achieve clean genetic sequencing results. I employed pH washing (0.001M HCl, 0.001M HCl, DiH2O, 0.0001M HCl, 0.001M HCl) for chemical treatment, followed by physical separation through centrifugation (5000rpm, 6500rpm) or filtration (2mm, 250um, 75um). The most successful method was deionized water washing, which yielded clear differences across stacked filters; the 2mm filtrate showed high levels of filamentous algae, with microalgae eluting in the 75um filtrate or remaining within agglutinations of algae larger filters. Base washing eluted the highest concentrations of microalgae, with larger filter sizes retaining more filamentous algae, indicating the breakdown of extracellular polymeric substances (EPS). Our downstream plans include sending the high-throughput next-generation sequencing to confirm the purity and ratios of filamentous and non-filamentous algae, as well as bacteria present, thereby validating the success of our treatments. Potential applications include creating community-based fractions for analysis, refining current sequencing data with clearer isolations, and generating designer biofilms to enhance our understanding of community interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Estimating irrigation water use from remotely sensed evapotranspiration data: Accuracy and uncertainties at field, water right, and regional scales

Irrigated agriculture is the dominant user of water globally, but most water withdrawals are not monitored or reported. As a result, it is largely unknown when, where, and how much water is used for irrigation. Here, we evaluated the ability of remotely sensed evapotranspiration (ET) data, integrated with other datasets, to calculate irrigation water withdrawals and applications in an intensively irrigated portion of the United States. We compared irrigation calculations based on an ensemble of satellite-driven ET models from OpenET with reported groundwater withdrawals from hundreds of farmer irrigation application records and a statewide flowmeter database at three spatial scales (field, water right group, and management area). At the field scale, we found that ET-based calculations of irrigation agreed best with reported irrigation when the OpenET ensemble mean was aggregated to the growing season timescale (bias = 1.6–4.9%, R 2 = 0.53–0.74), and agreement between calculated and reported irrigation was better for multi-year averages than for individual years. At the water right group scale, linking pumping wells to specific irrigated fields was the primary source of uncertainty. At the management area scale, calculated irrigation exhibited similar temporal patterns as flowmeter data but tended to be positively biased with more interannual variability. Disagreement between calculated and reported irrigation was strongly correlated with annual precipitation, and calculated and reported irrigation agreed more closely after statistically adjusting for annual precipitation. The selection of an ET model was also an important consideration, as variability across ET models was larger than the potential impacts of conservation measures employed in the region. From these results, we suggest key practices for working with ET-based irrigation data that include accurately accounting for changes in soil moisture, deep percolation, and runoff; careful verification of irrigated area and well-field linkages; and conducting application-specific evaluations of uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

Extreme Ionizing-Radiation-Resistant Bacterium

There is a growing concern that desiccation and extreme radiation-resistant, non-spore-forming microorganisms associated with spacecraft surfaces can withstand space environmental conditions and subsequent proliferation on another solar body. Such forward contamination would jeopardize future life detection or sample return technologies. The prime focus of NASA s planetary protection efforts is the development of strategies for inactivating resistance-bearing micro-organisms. Eradi cation techniques can be designed to target resistance-conferring microbial populations by first identifying and understanding their physiologic and biochemical capabilities that confers its elevated tolerance (as is being studied in Deinococcus phoenicis, as a result of this description). Furthermore, hospitals, food, and government agencies frequently use biological indicators to ensure the efficacy of a wide range of radiation-based sterilization processes. Due to their resistance to a variety of perturbations, the nonspore forming D. phoenicis may be a more appropriate biological indicator than those currently in use. The high flux of cosmic rays during space travel and onto the unshielded surface of Mars poses a significant hazard to the survival of microbial life. Thus, radiation-resistant microorganisms are of particular concern that can survive extreme radiation, desiccation, and low temperatures experienced during space travel. Spore-forming bacteria, a common inhabitant of spacecraft assembly facilities, are known to tolerate these extreme conditions. Since the Viking era, spores have been utilized to assess the degree and level of microbiological contamination on spacecraft and their associated spacecraft assembly facilities. Members of the non-sporeforming bacterial community such as Deinococcus radiodurans can survive acute exposures to ionizing radiation (5 kGy), ultraviolet light (1 kJ/m2), and desiccation (years). These resistive phenotypes of Deinococcus enhance the potential for transfer, and subsequent proliferation, on another solar body such as Mars and Europa. These organisms are more likely to escape planetary protection assays, which only take into account presence of spores. Hence, presences of extreme radiation-resistant Deinococcus in the cleanroom facility where spacecraft are assembled pose a serious risk for integrity of life-detection missions. The microorganism described herein was isolated from the surfaces of the cleanroom facility in which the Phoenix Lander was assembled. The isolated bacterial strain was subjected to a comprehensive polyphasic analysis to characterize its taxonomic position. This bacterium exhibits very low 16SrRNA similarity with any other environmental isolate reported to date. Both phenotypic and phylogenetic analyses clearly indicate that this isolate belongs to the genus Deinococcus and represents a novel species. The name Deinococcus phoenicis was proposed after the Phoenix spacecraft, which was undergoing assembly, testing, and launch operations in the spacecraft assembly facility at the time of isolation. D. phoenicis cells exhibited higher resistance to ionizing radiation (cobalt-60; 14 kGy) than the cells of the D. radiodurans (5 kGy). Thus, it is in the best interest of NASA to thoroughly characterize this organism, which will further assess in determining the potential for forward contamination. Upon the completion of genetic and physiological characteristics of D. phoenicis, it will be added to a planetary protection database to be able to further model and predict the probability of forward contamination.

Vaishampayan, Parag A.↗

Extreme Ionizing-Radiation-Resistant Bacterium

There is a growing concern that desiccation and extreme radiation-resistant, non-spore-forming microorganisms associated with spacecraft surfaces can withstand space environmental conditions and subsequent proliferation on another solar body. Such forward contamination would jeopardize future life detection or sample return technologies. The prime focus of NASA s planetary protection efforts is the development of strategies for inactivating resistance-bearing microorganisms. Eradification techniques can be designed to target resistance-conferring microbial populations by first identifying and understanding their physiologic and biochemical capabilities that confers its elevated tolerance (as is being studied in Deinococcus phoenicis, as a result of this description). Furthermore, hospitals, food, and government agencies frequently use biological indicators to ensure the efficacy of a wide range of radiation- based sterilization processes. Due to their resistance to a variety of perturbations, the non-spore forming D. phoenicis may be a more appropriate biological indicator than those currently in use. The high flux of cosmic rays during space travel and onto the unshielded surface of Mars poses a significant hazard to the survival of microbial life. Thus, radiation-resistant microorganisms are of particular concern that can survive extreme radiation, desiccation, and low temperatures experienced during space travel. Spore-forming bacteria, a common inhabitant of spacecraft assembly facilities, are known to tolerate these extreme conditions. Since the Viking era, spores have been utilized to assess the degree and level of microbiological contamination on spacecraft and their associated spacecraft assembly facilities. Members of the non-spore-forming bacterial community such as Deinococcus radiodurans can survive acute exposures to ionizing radiation (5 kGy), ultraviolet light (1 kJ/sq m), and desiccation (years). These resistive phenotypes of Deinococcus enhance the potential for transfer, and subsequent proliferation, on another solar body such as Mars and Europa. These organisms are more likely to escape planetary protection assays, which only take into account presence of spores. Hence, presences of extreme radiation-resistant Deinococcus in the cleanroom facility where spacecraft are assembled pose a serious risk for integrity of life-detection missions. The microorganism described herein was isolated from the surfaces of the cleanroom facility in which the Phoenix Lander was assembled. The isolated bacterial strain was subjected to a comprehensive polyphasic analysis to characterize its taxonomic position. This bacterium exhibits very low 16SrRNA similarity with any other environmental isolate reported to date. Both phenotypic and phylogenetic analyses clearly indicate that this isolate belongs to the genus Deinococcus and represents a novel species. The name Deinococcus phoenicis was proposed after the Phoenix spacecraft, which was undergoing assembly, testing, and launch operations in the spacecraft assembly facility at the time of isolation. D. phoenicis cells exhibited higher resistance to ionizing radiation (cobalt-60; 14 kGy) than the cells of the D. radiodurans (5 kGy). Thus, it is in the best interest of NASA to thoroughly characterize this organism, which will further assess in determining the potential for forward contamination. Upon the completion of genetic and physiological characteristics of D. phoenicis, it will be added to a planetary protection database to be able to further model and predict the probability of forward contamination.

Vaishampayan, Parag A.↗

A Comprehensive Assessment of Biologicals Contained Within Commercial Airliner Cabin Air

Both culture-based and culture-independent, biomarker-targeted microbial enumeration and identification technologies were employed to estimate total microbial and viral burden and diversity within the cabin air of commercial airliners. Samples from each of twenty flights spanning three commercial carriers were collected via air-impingement. When the total viable microbial population was estimated by assaying relative concentrations of the universal energy carrier ATP, values ranged from below detection limits (BDL) to 4.1 x 106 cells/cubic m of air. The total viable microbial population was extremely low in both of Airline A (approximately 10% samples) and C (approximately 18% samples) compared to the samples collected aboard flights on Airline A and B (approximately 70% samples). When samples were collected as a function of time over the course of flights, a gradual accumulation of microbes was observed from the time of passenger boarding through mid-flight, followed by a sharp decline in microbial abundance and viability from the initiation of descent through landing. It is concluded in this study that only 10% of the viable microbes of the cabin air were cultivable and suggested a need to employ state-of-the art molecular assay that measures both cultivable and viable-but-non-cultivable microbes. Among the cultivable bacteria, colonies of Acinetobacter sp. were by far the most profuse in Phase I, and Gram-positive bacteria of the genera Staphylococcus and Bacillus were the most abundant during Phase II. The isolation of the human pathogens Acinetobacter johnsonii, A. calcoaceticus, Janibacter melonis, Microbacterium trichotecenolyticum, Massilia timonae, Staphylococcus saprophyticus, Corynebacterium lipophiloflavum is concerning, as these bacteria can cause meningitis, septicemia, and a handful of sometimes fatal diseases and infections. Molecular microbial community analyses exhibited presence of the alpha-, beta-, gamma-, and delta- proteobacteria, as well as Gram-positive bacteria, Fusobacteria, Cyanobacteria, Deinococci, Bacterioidetes, Spirochetes, and Planctomyces in varying abundance. Neisseria meningitidis rDNA sequences were retrieved in great abundance from Airline A followed by Streptococcus oralis/mitis sequences. Pseudomonas synxantha sequences dominated Airline B clone libraries, followed by those of N. meningitidis and S. oralis/mitis. In Phase II, Airline C, sequences representative of more than 113 species, enveloping 12 classes of bacteria, were retrieved. Proteobacterial sequences were retrieved in greatest frequency (58% of all clone sequences), followed in short order by those stemming from Gram-positives bacteria (31% of all clone sequences). As for overall phylogenetic breadth, Gram-positive and alpha-proteobacteria seem to have a higher affinity for international flights, whereas beta-and gamma-proteobacteria are far more common about domestic cabin air parcels in Airline C samples. Ultimately, the majority of microbial species circulating throughout the cabin airs of commercial airliners are commensal, infrequently pathogenic normal flora of the human nasopharynx and respiratory system. Many of these microbes likely originate from the oral and nasal cavities, and lungs of passengers and flight crew and are disseminated unknowingly via routine conversation, coughing, sneezing, and stochastic passing of fomites. The data documented in this study will be useful to generate a baseline microbial population database and can be utilized to develop biosensor instrumentation for monitoring microbial quality of cabin or urban air.

microbial diversity↗

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Genomic analysis of Klebsiella aerogenes circulating in New Mexico

Klebsiella aerogenes is an opportunistic pathogen and a growing cause of healthcare-associated infections, characterized by multidrug resistance and the emergence of global high-risk clones. However, regional genomic surveillance data remain limited. Here, we sought to characterize the population structure, transmission dynamics and resistance mechanisms of clinical K. aerogenes in Albuquerque, New Mexico. We sequenced 177 clinical isolates collected between 2021 and 2023. We also developed a novel, species-specific PopPUNK database to facilitate rapid, high-resolution typing. The New Mexico K. aerogenes population was diverse but dominated by two global pandemic lineages, ST93 (47.5%) and ST4 (7.9%), which were significantly enriched for the virulence factors yersiniabactin and colibactin. Genomic evidence for recent local transmission was rare, with only four putative transmission pairs identified. The resistome was characterized by intrinsic and adaptive mutations. Nearly all isolates possessed gyrA mutations associated with decreased fluoroquinolone susceptibility. Mutations in the AmpC regulator AmpD and the outer membrane porin Omp36 were common, particularly within the dominant ST93 lineage. These mutations have been associated with increased AmpC-mediated carbapenem resistance. Our findings underscore the critical importance of genomic surveillance to monitor the transmission and evolution of adaptive resistance.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 350 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab Sequencing Lab. The GLDS contains rich metadata about each experiment and has integrated radiation dosimetry data from experiments flown on the Space Shuttle, International Space Station, and Free Flying spacecrafts. With the increasing amount and complexity of omics data being generated, GeneLab utilizes community-defined, common models for metadata and terminology so that omics data and results are discoverable and reliably reproducible. GeneLab uses the ISA-Tab specification and semantic model for organizing and representing omics metadata. In addition to metadata standards, data files must be open-source file or common exchange formats to ensure accessibility and usability by all users. To ease data ingestion and transfer, the web-based submission tool allows PIs a user-friendly user interface to curate, organize, and publish their space relevant omics data. In the more recent years, data curation and submission portal has incorporated the FAIR principles making data findable, accessible, interoperable, and reusable. To increase reusability of data, GeneLab has implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 200 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. To train the next generation of scientists, NASA offers training programs such as GeneLab 4 High School (GL4HS) and GeneLab 4 Universities. NLM Curation at a Scale Workshop 2022 | NASA GeneLab (GL4U) to teach students bioinformatics and computational biology methods to analyze omics data. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

GeneLab↗

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗