Search NASA⌕ Search

SEARCH · Search NASA

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Genotypic analyses of IncHI2 plasmids from enteric bacteria

Incompatibility (Inc) HI2 plasmids are large (typically > 200 kb), transmissible plasmids that encode antimicrobial resistance (AMR), heavy metal resistance (HMR) and disinfectants/biocide resistance (DBR). To better understand the distribution and diversity of resistance-encoding genes among IncHI2 plasmids, computational approaches were used to evaluate resistance and transfer-associated genes among the plasmids. Complete IncHI2 plasmid (N - 667) sequences were extracted from GenBank and analyzed using AMRFinderPlus, IntegronFinder and Plasmid Transfer Factor database. The most common IncHI2-carrying genera included Enterobacter (N = 209), Escherichia (N = 208), and Salmonella (N = 204). Resistance genes distribution was diverse, with plasmids from Escherichia and Salmonella showing general similarity in comparison to Enterobacter and other taxa, which grouped together. Plasmids from Enterobacter and other taxa had a higher prevalence of multiple mercury resistance genes and arsenic resistance gene, arsC, compared to Escherichia and Salmonella. For sulfonamide resistance, sul1 was more common among Enterobacter and other taxa, compared to sul2 and sul3 for Escherichia and Salmonella. Similar gene diversity trends were also observed for tetracyclines, quinolones, β-lactams, and colistin. Over 99% of plasmids carried at least 25 IncHI2-associated conjugal transfer genes. These findings highlight the diversity and dissemination potential for resistance across different enteric bacteria and value of computational-based approaches for the resistance-gene assessment.

59 BASIC BIOLOGICAL SCIENCES↗

Multiplex detection and identification of viral, bacterial, and protozoan pathogens in human blood and plasma using an expanded high-density resequencing microarray platform

Introduction: Nucleic acid tests for blood donor screening have improved the safety of the blood supply; however, increasing numbers of emerging pathogen tests are burdensome. Multiplex testing platforms are a potential solution. Methods: The Blood Borne Pathogen Resequencing Microarray Expanded (BBP-RMAv.2) can perform multiplex detection and identification of 80 viruses, bacteria and parasites. This study evaluated pathogen detection in human blood or plasma. Samples spiked with selected pathogens, each with one of 6 viruses, 2 bacteria and 5 protozoans were tested on this platform. The nucleic acids were extracted, amplified using multiplexed sets of primers, and hybridized to a microarray. The reported sequences were aligned to a database to identify the pathogen. To directly compare the microarray to an emerging molecular approach, the amplified nucleic acids were also submitted to nanopore next generation sequencing (NGS). Results: The BBP-RMAv.2 detected viral pathogens at a concentration as low as 100 copies/ml and a range of concentrations from 1,000 to 100,000 copies/ml for all the spiked pathogens. Coded specimens were identified correctly demonstrating the effectiveness of the platform. The nanopore sequencing correctly identified most samples and the results of the two platforms were compared. Discussion: These results indicated that the BBP-RMAv.2 could be employed for multiplex detection with potential for use in blood safety or disease diagnosis. The NGS was nearly as effective at identifying pathogens in blood and performed better than BBP-RMAv.2 at identifying pathogen-negative samples.

59 BASIC BIOLOGICAL SCIENCES↗

HIV Molecular Immunology 2025

HIV Molecular Immunology is a companion volume to HIV Sequence Compendium. This publication, the 2025 edition, is the PDF version of Los Alamos Na tional Laboratory’s web-based HIV Molecular Immunology Database (https://www.hiv.lanl.gov/content/ immunology/). The web interface for this relational database has many search interfaces for HIV immunological in formation, as well as interactive tools to help immunologists design reagents and interpret their results.

59 BASIC BIOLOGICAL SCIENCES↗

Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites

Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.

Biochemistry & Molecular Biology↗

Pilot Study on the Investigation of Tear Fluid Biomarkers as an Indicator of Ocular, Neurological, and Immunological Health in Astronauts

The purpose of this pilot study is to investigate the collection, preparation, and analysis of tear biomarkers as a means of assessing ocular, neurological, and immunological health. At present, no published data exists on the cytokine profiles of tears from astronauts exposed to long periods of microgravity and space irradiations. In addition, no published data exist on cytokine (biomarker) profiles of tears that have been collected from irradiated non-human biological systems (primates and other animal models). A goal for the proposed pilot study is to discover novel tear biomarkers which can help inform researchers, clinicians, epidemiologist and healthcare providers about the health status of a living biological system, as well as informing them when a disease state is triggered. This would be done via analysis of the onset of expression of pro-inflammatory cytokines, leading up to the full progression of a disease (i.e. cancer, loss of vision, radiation-induced oxidative stress, cardiovascular disorders, fibrosis in major organs, bone loss). Another goal of this pilot study is to investigate the state of disease against proposed medical countermeasures, in order to determine whether the countermeasures are efficacious in preventing or mitigating these injuries. An example of an up and coming tear biomarker technology, Ascendant Dx, a clinical stage diagnostic company, is developing a screening test to detect breast cancer using proteins from tears. The team utilized Liquid Chromatography -Mass Spectrometry with Mass analysis (LC MS/MS) as a discovery platform followed by validation with ELISA to come up with a panel of protein biomarkers that can differentiate breast cancer samples from control ("cancer free") samples with results far surpassing the results of imaging techniques in use today. Continued research into additional proteins is underway to increase the sensitivity and specificity of the test and development efforts are on the way to transfer the test onto a fast, accurate and inexpensive point of care platform. In conclusion, the expected results from this proposed pilot study are to: a) establish an SOP for retrieving/storing/transporting tear fluid samples from multicentre sites b) establish a normal range for relevant biomarkers in tears; and c) establish a database (biobank) of tears of space naïve versus veteran astronauts, to establish a personal baseline for long-term ocular health monitoring

Morton, Stephen↗

WEBINAR, May 6: New Discoveries Using GeneLab

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 220 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetry data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 120 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

Sylvain V. Costes↗

The Apollo Sample Suite: 50 Years of Solar System Insight

The Apollo program was undoubtable a crowning achievement in human history. In addition to the obvious cultural significance, scientific results from the Apollo program had a lasting impression on a range of scientific fields, none more so that the effect the samples had on the fields of geology and cosmochemistry. The six Apollo missions collected 382 kg of rock, regolith, and core samples from geologically diverse locations on the Moon. In the nearly 50 years since the first samples were returned, there have been over 3000 different requests for samples, each yielding insights into fields as disparate as biology, medicine, astronomy, engineering, material science, and of course geology. Early studies of the Apollo samples revealed primary insights into the origin and evolution of the Moon, and of the Earth-Moon system, but the results also had implications for bodies throughout the solar system, e.g., defining crater counting rates. Over the decades, continued study of the Apollo samples by new generations of scientists using new instruments have continued to yield significant new discoveries, including the presence of endogenous water in the Moon and the possible presence of a lunar cataclysm, that in turn has contributed to new models of solar system formation and evolution. The Apollo samples have often been used as a proxy for studying other bodies like Mercury or asteroids. The Apollo samples have also directly contributed to the interpretation of remotely sensed data sets, including their use as ground truth for both Clementine and Lunar Prospector global geochemical maps. Despite the Apollo samples being a static collection, recent efforts will ensure that investigators continue to have access to new samples. For example, there was a recent solicitation for study of previously unopened Apollo samples in vacuum-sealed containers, as well as new access to samples stored frozen or in a He atmosphere. Similarly, the use of X-ray computed tomography as part of the curation process is identifying new clasts within polymict breccias that are available for study. Finally, the MoonDB project is putting all previously published lunar geochemical analyses into a searchable database, which should facilitate new investigations.

Zeigler, Ryan↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

AstroAmpSeq: Microbial Bioinformatics Education with NASA GeneLab’s Amplicon Pipeline

The prevalence and importance of large sequencing datasets in microbiology has led to a movement to share microbial ecology experimental data through open-access databases. This is particularly true of experiments that are difficult to replicate, such as those conducted in the spaceflight environment and shared via NASA GeneLab. It is now possible and indeed valuable for students to access and re-analyze these shared datasets for educational and research purposes. To provide students with experience utilizing microbial bioinformatics tools, GeneLab for Colleges and Universities (GL4U) has designed AstroAmpSeq, a week-long, virtually implemented project-based learning (PBL) minicourse to instruct undergraduate students on 16S amplicon sequencing. AstroAmpSeq was created to be accessible to students without prior bioinformatics or microbial ecology experience. During the minicourse students work in teams to process, analyze, and visualize a subsample of GeneLab dataset GLDS-280 using GeneLab’s standard amplicon processing pipeline, which is based in R. Students develop a hypothesis related to the dataset then generate and analyze figures to evaluate their hypothesis. Formative assessment of student learning is determined via pre- and post-evaluations, peer feedback, and self-reflection. Project and presentation rubrics serve as a summative assessment of student learning. GL4U AstroAmpSeq not only meets American Society for Microbiology Curriculum Guidelines, but also incites student interest in research by an inquiry-based approach and can be made part of a larger semester-long curriculum. GL4U AstroAmpSeq raises awareness of space microbiology and bioinformatics as a field and career path among undergraduates. Further, by using a GeneLab dataset and nesting microbiology techniques into the real-world application of space biology, AstroAmpSeq enforces deeper and longer-lasting student learning.

microbiology↗

U.S. Pacific Coast Workshop Report on Preconstruction Research Recommendations (U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) Project)

In May 2022, the U.S. Offshore Wind Synthesis of Environmental Effects Research (SEER) project team hosted a stakeholder workshop focused on preconstruction (baseline) research needs for potential floating offshore wind (OSW) energy development on the U.S. Pacific Coast, including California, Oregon, and Washington. Prior to the workshop, the SEER team developed a set of initial synthesized research recommendations that were identified based on a review of relevant, publicly available resources and with advisory group input. The workshop covered three marine life breakout groups on subsequent days to discuss research recommendations related to 1) marine mammals and sea turtles, 2) fish and invertebrates, and 3) birds and bats. As part of the workshop, over a hundred participants from the public and private sectors provided feedback on various aspects of the initial research recommendations, including associated data and knowledge gaps, benefits/limitations of available methods and technologies, and technological advancements or infrastructure needed to address the recommendation. Approximately 1,000 total comments were received on the workshop MURAL boards and were synthesized in this report. Based on workshop feedback, SEER developed a final database of over 500 specific research recommendations based on more than 40 resources. In Fall 2022, the full database and a tool with updated synthesized research recommendations were disseminated on Tethys (https://tethys.pnnl.gov) to assist with informing future funding opportunities and research programming. There is a continued need to improve awareness of the potential environmental effects, monitoring technologies, and management strategies for floating OSW energy development on the U.S. Pacific Coast. Coordination of these activities will require the sustained involvement of multiple stakeholders from across sectors. Beyond the baseline considerations discussed in this workshop, future state-of-the-science activities should be planned to consider research needs across wind energy life cycle phases for all relevant wildlife taxa and associated habitat and ecosystem processes.

17 WIND ENERGY↗

Nuclear model calculations and their role in space radiation research

Proper assessments of spacecraft shielding requirements and concomitant estimates of risk to spacecraft crews from energetic space radiation requires accurate, quantitative methods of characterizing the compositional changes in these radiation fields as they pass through thick absorbers. These quantitative methods are also needed for characterizing accelerator beams used in space radiobiology studies. Because of the impracticality/impossibility of measuring these altered radiation fields inside critical internal body organs of biological test specimens and humans, computational methods rather than direct measurements must be used. Since composition changes in the fields arise from nuclear interaction processes (elastic, inelastic and breakup), knowledge of the appropriate cross sections and spectra must be available. Experiments alone cannot provide the necessary cross section and secondary particle (neutron and charged particle) spectral data because of the large number of nuclear species and wide range of energies involved in space radiation research. Hence, nuclear models are needed. In this paper current methods of predicting total and absorption cross sections and secondary particle (neutrons and ions) yields and spectra for space radiation protection analyses are reviewed. Model shortcomings are discussed and future needs presented. c2002 COSPAR. Published by Elsevier Science Ltd. All right reserved.

NASA Center JSC↗

Warming is Associated With More Encoded Antimicrobial Resistance Genes and Transcriptions Within Five Drug Classes in Soil Bacteria: A Case Study and Synthesis

ABSTRACT The effect of warming on anti‐microbial resistance (AMR) genes in the environment has critical implications for public health but is little studied. We collected published soil bacterial genomes from the BV‐BRC database and tested the correlation between reported optimal growth temperature and the number of encoded AMR genes. Furthermore, we tested the relationship between temperature and AMR gene transcription in a natural ecosystem by analysing soil transcriptomes from a warming manipulation experiment in an Alaskan boreal forest. We hypothesised that there is a positive relationship between warming and AMR prevalence in gene content in bacterial genomes and transcriptomic sequences, and that this effect would vary by drug class. Regarding the bacterial genomes, we found a positive relationship between the fraction of encoded AMR genes and the reported optimal temperature of soil bacteria. The drug classes tetracycline and lincosamide/macrolide/streptogramin had the strongest positive relationship with reported optimal temperature. For the case study in a natural ecosystem, we found 61 significantly upregulated AMR gene‐associated transcripts spanning eight drug classes in warmed plots. In the Alaskan soil samples, we found that warming elicited the strongest positive effect on transcripts targeting lincosamide/streptogramin, beta‐lactam and phenicol/quinolone antibiotics. Overall, higher temperatures were linked to AMR gene prevalence.

Hacopian, Melanie T. [Department of Ecology and Ev↗

Rapid recovery from the Late Ordovician mass extinction

Understanding the evolutionary role of mass extinctions requires detailed knowledge of postextinction recoveries. However, most models of recovery hinge on a direct reading of the fossil record, and several recent studies have suggested that the fossil record is especially incomplete for recovery intervals immediately after mass extinctions. Here, we analyze a database of genus occurrences for the paleocontinent of Laurentia to determine the effects of regional processes on recovery and the effects of variations in preservation and sampling intensity on perceived diversity trends and taxonomic rates during the Late Ordovician mass extinction and Early Silurian recovery. After accounting for variation in sampling intensity, we find that marine benthic diversity in Laurentia recovered to preextinction levels within 5 million years, which is nearly 15 million years sooner than suggested by global compilations. The rapid turnover in Laurentia suggests that processes such as immigration may have been particularly important in the recovery of regional ecosystems from environmental perturbations. However, additional regional studies and a global analysis of the Late Ordovician mass extinction that accounts for variations in sampling intensity are necessary to confirm this pattern. Because the record of Phanerozoic mass extinctions and postextinction recoveries may be compromised by variations in preservation and sampling intensity, all should be reevaluated with sampling-standardized analyses if the evolutionary role of mass extinctions is to be fully understood.

Evolution↗

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

PAVC: The foundation for a Pan-Arctic Vegetation Cover database

Field-measured Arctic vegetation cover data is essential for creating accurate, high-quality vegetation structure and composition maps. Extrapolating field data into high-resolution cover maps provides detailed, function-specific information for use in Earth System Models, vegetation classifications, and monitoring vegetation change over time and space. However, field campaigns that collect plant cover vary substantially in scope, method, and purpose, which makes them difficult to unify across data stores, and they are often not designed to meet remote sensing needs. In this work, we synthesized and harmonized field-based fractional cover data from various data stores to create a high-quality, consistent repository schema for remote sensing-based vegetation cover mapping applications. We developed a reproducible workflow for synthesizing visual estimate and point-intercept fractional cover data. The resultant Pan-Arctic Vegetation Cover (PAVC) database contains synthesized fractional cover at both the species and plant functional type levels. The latter includes absolute foliar cover for deciduous shrubs and trees, evergreen shrubs and trees, forbs, graminoids, lichen, bryophytes, and “other” vegetation, as well as absolute cover for litter and top cover for water and bare ground.

Steckler, Morgan R. [Oak Ridge National Laboratory↗