Search NASA⌕ Search

SEARCH · Search NASA

Results for “biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES↗

The PMIP4 Contribution to CMIP6 - Part 1: Overview and Over-Arching Analysis Plan

This paper is the first of a series of four GMD (Geoscientific Model Development) papers on the PMIP4-CMIP6 (Paleoclimate Modelling Intercomparison Project - Phase 4 -- Coupled Model Intercomparison Project - Phase 6) experiments. Part 2 (Otto-Bliesner et al., 2017) gives details about the two PMIP4-CMIP6 interglacial experiments, Part 3 (Jungclaus et al., 2017) about the last millennium experiment, and Part 4 (Kageyama et al., 2017) about the Last Glacial Maximum experiment. The mid-Pliocene Warm Period experiment is part of the Pliocene Model Intercomparison Project (PlioMIP) - Phase 2, detailed in Haywood et al. (2016). The goal of the Paleoclimate Modelling Intercomparison Project (PMIP) is to understand the response of the climate system to different climate forcings for documented climatic states very different from the present and historical climates. Through comparison with observations of the environmental impact of these climate changes, or with climate reconstructions based on physical, chemical, or biological records, PMIP also addresses the issue of how well state-of-the-art numerical models simulate climate change. Climate models are usually developed using the present and historical climates as references, but climate projections show that future climates will lie well outside these conditions. Palaeoclimates very different from these reference states therefore provide stringent tests for state-of-the-art models and a way to assess whether their sensitivity to forcings is compatible with palaeoclimatic evidence. Simulations of five different periods have been designed to address the objectives of the sixth phase of the Coupled Model Intercomparison Project (CMIP6): the millennium prior to the industrial epoch (CMIP6 name: past1000); the mid-Holocene, 6000 years ago (midHolocene); the Last Glacial Maximum, 21,000 years ago (lgm); the Last Interglacial, 127,000 years ago (lig127k); and the mid-Pliocene Warm Period, 3.2 million years ago (midPliocene-eoi400). These climatic periods are well documented by palaeoclimatic and palaeoenvironmental records, with climate and environmental changes relevant for the study and projection of future climate changes. This paper describes the motivation for the choice of these periods and the design of the numerical experiments and database requests, with a focus on their novel features compared to the experiments performed in previous phases of PMIP and CMIP. It also outlines the analysis plan that takes advantage of the comparisons of the results across periods and across CMIP6 in collaboration with other MIPs.

climate↗

BRAKER3: Fully automated genome annotation using RNA-seq and protein evidence with GeneMark-ETP, AUGUSTUS, and TSEBRA

Gene prediction has remained an active area of bioinformatics research for a long time. Still, gene prediction in large eukaryotic genomes presents a challenge that must be addressed by new algorithms. The amount and significance of the evidence available from transcriptomes and proteomes vary across genomes, between genes, and even along a single gene. User-friendly and accurate annotation pipelines that can cope with such data heterogeneity are needed. The previously developed annotation pipelines BRAKER1 and BRAKER2 use RNA-seq or protein data, respectively, but not both. A further significant performance improvement integrating all three data types was made by the recently released GeneMark-ETP. We here present the BRAKER3 pipeline that builds on GeneMark-ETP and AUGUSTUS, and further improves accuracy using the TSEBRA combiner. BRAKER3 annotates protein-coding genes in eukaryotic genomes using both short-read RNA-seq and a large protein database, along with statistical models learned iteratively and specifically for the target genome. We benchmarked the new pipeline on genomes of 11 species under an assumed level of relatedness of the target species proteome to available proteomes. BRAKER3 outperforms BRAKER1 and BRAKER2. The average transcript-level F1-score is increased by about 20 percentage points on average, whereas the difference is most pronounced for species with large and complex genomes. BRAKER3 also outperforms other existing tools, MAKER2, Funannotate, and FINDER. The code of BRAKER3 is available on GitHub and as a ready-to-run Docker container for execution with Docker or Singularity. Overall, BRAKER3 is an accurate, easy-to-use tool for eukaryotic genome annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Three Conservation Applications of Astronaut Photographs of Earth: Tidal Flat Loss (Japan), Elephant Impacts on Vegetation (Botswana), and Seagrass and Mangrove Monitoring (Australia)

NASA photographs taken from low Earth orbit can provide information relevant to conservation biology. This data source is now more accessible due to improvements in digitizing technology, Internet file transfer, and availability of image processing software. We present three examples of conservation-related projects that benefited from using orbital photographs. (1) A time series of photographs from the Space Shuttle showing wetland conversion in Japan was used as a tool for communicating about the impacts of tidal flat loss. Real-time communication with astronauts about a newsworthy event resulted in acquiring current imagery. These images and the availability of other high resolution digital images from NASA provided timely public information on the observed changes. (2) A Space Shuttle photograph of Chobe National Park in Botswana was digitally classified and analyzed to identify the locations of elephant-impacted woodland. Field validation later confirmed that areas identified on the image showed evidence of elephant impacts. (3) A summary map from intensive field surveys of seagrasses in Shoalwater Bay, Australia was used as reference data for a supervised classification of a digitized photograph taken from orbit. The classification was able to distinguish seagrasses, sediments and mangroves with accuracy approximating that in studies using other satellite remote sensing data. Orbital photographs are in the public domain and the database of nearly 400,000 photographs from the late 1960s to the present is available at a single searchable location on the Internet. These photographs can be used by conservation biologists for general information about the landscape and in quantitative applications.

Lulla, Kamlesh P.↗

BGC Atlas: a web resource for exploring the global chemical diversity encoded in bacterial genomes

Secondary metabolites are compounds not essential for an organism’s development, but provide significant ecological and physiological benefits. These compounds have applications in medicine, biotechnology and agriculture. Their production is encoded in biosynthetic gene clusters (BGCs), groups of genes collectively directing their biosynthesis. The advent of metagenomics has allowed researchers to study BGCs directly from environmental samples, identifying numerous previously unknown BGCs encoding unprecedented chemistry. Here, we present the BGC Atlas (https://bgc-atlas.cs.uni-tuebingen.de), a web resource that facilitates the exploration and analysis of BGC diversity in metagenomes. The BGC Atlas identifies and clusters BGCs from publicly available datasets, offering a centralized database and a web interface for metadata-aware exploration of BGCs and gene cluster families (GCFs). We analyzed over 35 000 datasets from MGnify, identifying nearly 1.8 million BGCs, which were clustered into GCFs. The analysis showed that ribosomally synthesized and post-translationally modified peptides are the most abundant compound class, with most GCFs exhibiting high environmental specificity. We believe that our tool will enable researchers to easily explore and analyze the BGC diversity in environmental samples, significantly enhancing our understanding of bacterial secondary metabolites, and promote the identification of ecological and evolutionary factors shaping the biosynthetic potential of microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

The monitoring system for vibratory disturbance detection in microgravity environment aboard the international space station

Scientists in the Office of Life and Microgravity Sciences and Applications within the Microgravity Research Division oversee studies in important physical, chemical, and biological processes in microgravity environment. Research is conducted in microgravity environment because of the beneficial results that come about for experiments. When research is done in normal gravity, scientists are limited to results that are affected by the gravity of Earth. Microgravity provides an environment where solid, liquid, and gas can be observed in a natural state of free fall and where many different variables are eliminated. One challenge that NASA faces is that space flight opportunities need to be used effectively and efficiently in order to ensure that some of the most scientifically promising research is conducted. Different vibratory sources are continually active aboard the International Space Station (ISS). Some of the vibratory sources include crew exercise, experiment setup, machinery startup (life support fans, pumps, freezer/compressor, centrifuge), thruster firings, and some unknown events. The Space Acceleration Measurement System (SAMs), which acts as the hardware and carefully positioned aboard the ISS, along with the Microgravity Environment Monitoring System MEMS), which acts as the software and is located here at NASA Glenn, are used to detect these vibratory sources aboard the ISS and recognize them as disturbances. The various vibratory disturbances can sometimes be harmful to the scientists different research projects. Some vibratory disturbances are recognized by the MEMS's database and some are not. Mainly, the unknown events that occur aboard the International Space Station are the ones of major concern. To better aid in the research experiments, the unknown events are identified and verified as unknown events. Features, such as frequency, acceleration level, time and date of recognition of the new patterns are stored in an Excel database. My task is to carefully synthesize frequency and acceleration patterns of unknown events within the Excel database into a new file to determine whether or not certain information that is received i s considered a real vibratory source. Once considered as a vibratory source, further analysis is carried out. The resulting information is used to retrain the MEMS to recognize them as known patterns. These different vibratory disturbances are being constantly monitored to observe if, in any way, the disturbances have an effect on the microgravity environment that research experiments are exposed to. If the disturbance has little or no effect on the experiments, then research is continued. However, if the disturbance is harmful to the experiment, scientists act accordingly by either minimizing the source or terminating the research and neither NASA's time nor money is wasted.

Laster, Rachel M.↗

Sensitive and error-tolerant annotation of protein-coding DNA with BATH

We present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.

59 BASIC BIOLOGICAL SCIENCES↗

GeneLab Analysis Working Group Kick-Off Meeting

Goals to achieve for GeneLab AWG - GL vision - Review of GeneLab AWG charter Timeline and milestones for 2018 Logistics - Monthly Meeting - Workshop - Internship - ASGSR Introduction of team leads and goals of each group Introduction of all members Q/A Three-tier Client Strategy to Democratize Data Physiological changes, pathway enrichment, differential expression, normalization, processing metadata, reproducibility, Data federation/integration with heterogeneous bioinformatics external databases The GLDS currently serves over 100 omics investigations to the biomedical community via open access. In order to expand the scope of metadata record searches via the GLDS, we designed a metadata warehouse that collects and updates metadata records from external systems housing similar data. To demonstrate the capabilities of federated search and retrieval of these data, we imported metadata records from three open-access data systems into the GLDS metadata warehouse: NCBI's Gene Expression Omnibus (GEO), EBI's PRoteomics IDEntifications (PRIDE) repository, and the Metagenomics Analysis server (MG-RAST). Each of these systems defines metadata for omics data sets differently. One solution to bridge such differences is to employ a common object model (COM) to which each systems' representation of metadata can be mapped. Warehoused metadata records are then transformed at ETL to this single, common representation. Queries generated via the GLDS are then executed against the warehouse, and matching records are shown in the COM representation (Fig. 1). While this approach is relatively straightforward to implement, the volume of the data in the omics domain presents challenges in dealing with latency and currency of records. Furthermore, the lack of a coordinated has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

GeneLab↗

Tools for Performing SBG hyperspectral Observing System Simulation Experiment

One of NASA’s Decadal Survey mission, Surface Biology and Geology (SBG), will include a hyperspectral remote sensing imager, which has a very high spatial resolution and a wide spectral coverage (from UV to Near IR). Unprecedented large data volumes will be generated by the SBG hyperspectral instrument. Before the launch of the new satellite, an Observing System Simulation Experiment (OSSE) can be used to study different designs of the new satellite system. One of the key components in an OSSE study is a radiative transfer model (RTM) or forward model. In this presentation, we will describe a Principal Component-based Radiative Transfer Model (PCRTM) which is capable of simulating atmospheric (TOA) radiance or reflectance spectra from far IR to visible and UV spectral regions (50 wavenumber to 30000 wavenumber) quickly and accurately. Multiple scattering from multiple layers of clouds/aerosols are included in the model. The PCRTM has a very good accuracy relative to reference line-by-line radiative transfer models (LBLRTM), and it saves 3-4 orders of magnitude computational time relative to LBLRTM or MODTRAN. The PCRTM model has been successfully used to analyze large volumes of data from hyperspectral sensors such as AIRS, CrIS, and IASI. It has also been used to perform OSSE studies for the Climate Absolute Radiance and Refractivity Observatory (CLARREO) mission. Another useful tool for the OSSE is surface Bidirectional Reflectance Distribution Function (BRDF) database. It is very crucial for the SBG OSSE to include realistic BRDF spectra. Currently, most of the surface reflectance spectra such as those in the ECOSIS and ECOSTRESS are measured at specific observation geometries. We have developed a hyperspectral bidirectional reflectance (HSBR) model which combines Ross-Li BRDF model with the existing reflectance spectral libraries using a principal component analysis. This HSBR model can provide realistic BRDF spectra under various observation conditions. It can also be used to generate realistic BRDF spectra using measurements from multi-band imagers or spectrometers such as MODIS or VIIRS.

Xu Liu↗

Shifts in evolutionary lability underlie independent gains and losses of root-nodule symbiosis in a single clade of plants

Abstract Root nodule symbiosis (RNS) is a complex trait that enables plants to access atmospheric nitrogen converted into usable forms through a mutualistic relationship with soil bacteria. Pinpointing the evolutionary origins of RNS is critical for understanding its genetic basis, but building this evolutionary context is complicated by data limitations and the intermittent presence of RNS in a single clade of ca. 30,000 species of flowering plants, i.e., the nitrogen-fixing clade (NFC). We developed the most extensive de novo phylogeny for the NFC and an RNS trait database to reconstruct the evolution of RNS. Our analysis identifies evolutionary rate heterogeneity associated with a two-step process: An ancestral precursor state transitioned to a more labile state from which RNS was rapidly gained at multiple points in the NFC. We illustrate how a two-step process could explain multiple independent gains and losses of RNS, contrary to recent hypotheses suggesting one gain and numerous losses, and suggest a broader phylogenetic and genetic scope may be required for genome-phenome mapping.

59 BASIC BIOLOGICAL SCIENCES↗

Genotypic analyses of IncHI2 plasmids from enteric bacteria

Incompatibility (Inc) HI2 plasmids are large (typically > 200 kb), transmissible plasmids that encode antimicrobial resistance (AMR), heavy metal resistance (HMR) and disinfectants/biocide resistance (DBR). To better understand the distribution and diversity of resistance-encoding genes among IncHI2 plasmids, computational approaches were used to evaluate resistance and transfer-associated genes among the plasmids. Complete IncHI2 plasmid (N - 667) sequences were extracted from GenBank and analyzed using AMRFinderPlus, IntegronFinder and Plasmid Transfer Factor database. The most common IncHI2-carrying genera included Enterobacter (N = 209), Escherichia (N = 208), and Salmonella (N = 204). Resistance genes distribution was diverse, with plasmids from Escherichia and Salmonella showing general similarity in comparison to Enterobacter and other taxa, which grouped together. Plasmids from Enterobacter and other taxa had a higher prevalence of multiple mercury resistance genes and arsenic resistance gene, arsC, compared to Escherichia and Salmonella. For sulfonamide resistance, sul1 was more common among Enterobacter and other taxa, compared to sul2 and sul3 for Escherichia and Salmonella. Similar gene diversity trends were also observed for tetracyclines, quinolones, β-lactams, and colistin. Over 99% of plasmids carried at least 25 IncHI2-associated conjugal transfer genes. These findings highlight the diversity and dissemination potential for resistance across different enteric bacteria and value of computational-based approaches for the resistance-gene assessment.

59 BASIC BIOLOGICAL SCIENCES↗

Multiplex detection and identification of viral, bacterial, and protozoan pathogens in human blood and plasma using an expanded high-density resequencing microarray platform

Introduction: Nucleic acid tests for blood donor screening have improved the safety of the blood supply; however, increasing numbers of emerging pathogen tests are burdensome. Multiplex testing platforms are a potential solution. Methods: The Blood Borne Pathogen Resequencing Microarray Expanded (BBP-RMAv.2) can perform multiplex detection and identification of 80 viruses, bacteria and parasites. This study evaluated pathogen detection in human blood or plasma. Samples spiked with selected pathogens, each with one of 6 viruses, 2 bacteria and 5 protozoans were tested on this platform. The nucleic acids were extracted, amplified using multiplexed sets of primers, and hybridized to a microarray. The reported sequences were aligned to a database to identify the pathogen. To directly compare the microarray to an emerging molecular approach, the amplified nucleic acids were also submitted to nanopore next generation sequencing (NGS). Results: The BBP-RMAv.2 detected viral pathogens at a concentration as low as 100 copies/ml and a range of concentrations from 1,000 to 100,000 copies/ml for all the spiked pathogens. Coded specimens were identified correctly demonstrating the effectiveness of the platform. The nanopore sequencing correctly identified most samples and the results of the two platforms were compared. Discussion: These results indicated that the BBP-RMAv.2 could be employed for multiplex detection with potential for use in blood safety or disease diagnosis. The NGS was nearly as effective at identifying pathogens in blood and performed better than BBP-RMAv.2 at identifying pathogen-negative samples.

59 BASIC BIOLOGICAL SCIENCES↗

HIV Molecular Immunology 2025

HIV Molecular Immunology is a companion volume to HIV Sequence Compendium. This publication, the 2025 edition, is the PDF version of Los Alamos Na tional Laboratory’s web-based HIV Molecular Immunology Database (https://www.hiv.lanl.gov/content/ immunology/). The web interface for this relational database has many search interfaces for HIV immunological in formation, as well as interactive tools to help immunologists design reagents and interpret their results.

59 BASIC BIOLOGICAL SCIENCES↗

Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites

Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.

Biochemistry & Molecular Biology↗

Pilot Study on the Investigation of Tear Fluid Biomarkers as an Indicator of Ocular, Neurological, and Immunological Health in Astronauts

The purpose of this pilot study is to investigate the collection, preparation, and analysis of tear biomarkers as a means of assessing ocular, neurological, and immunological health. At present, no published data exists on the cytokine profiles of tears from astronauts exposed to long periods of microgravity and space irradiations. In addition, no published data exist on cytokine (biomarker) profiles of tears that have been collected from irradiated non-human biological systems (primates and other animal models). A goal for the proposed pilot study is to discover novel tear biomarkers which can help inform researchers, clinicians, epidemiologist and healthcare providers about the health status of a living biological system, as well as informing them when a disease state is triggered. This would be done via analysis of the onset of expression of pro-inflammatory cytokines, leading up to the full progression of a disease (i.e. cancer, loss of vision, radiation-induced oxidative stress, cardiovascular disorders, fibrosis in major organs, bone loss). Another goal of this pilot study is to investigate the state of disease against proposed medical countermeasures, in order to determine whether the countermeasures are efficacious in preventing or mitigating these injuries. An example of an up and coming tear biomarker technology, Ascendant Dx, a clinical stage diagnostic company, is developing a screening test to detect breast cancer using proteins from tears. The team utilized Liquid Chromatography -Mass Spectrometry with Mass analysis (LC MS/MS) as a discovery platform followed by validation with ELISA to come up with a panel of protein biomarkers that can differentiate breast cancer samples from control ("cancer free") samples with results far surpassing the results of imaging techniques in use today. Continued research into additional proteins is underway to increase the sensitivity and specificity of the test and development efforts are on the way to transfer the test onto a fast, accurate and inexpensive point of care platform. In conclusion, the expected results from this proposed pilot study are to: a) establish an SOP for retrieving/storing/transporting tear fluid samples from multicentre sites b) establish a normal range for relevant biomarkers in tears; and c) establish a database (biobank) of tears of space naïve versus veteran astronauts, to establish a personal baseline for long-term ocular health monitoring

Morton, Stephen↗

WEBINAR, May 6: New Discoveries Using GeneLab

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 220 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetry data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 120 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

Sylvain V. Costes↗

The Apollo Sample Suite: 50 Years of Solar System Insight

The Apollo program was undoubtable a crowning achievement in human history. In addition to the obvious cultural significance, scientific results from the Apollo program had a lasting impression on a range of scientific fields, none more so that the effect the samples had on the fields of geology and cosmochemistry. The six Apollo missions collected 382 kg of rock, regolith, and core samples from geologically diverse locations on the Moon. In the nearly 50 years since the first samples were returned, there have been over 3000 different requests for samples, each yielding insights into fields as disparate as biology, medicine, astronomy, engineering, material science, and of course geology. Early studies of the Apollo samples revealed primary insights into the origin and evolution of the Moon, and of the Earth-Moon system, but the results also had implications for bodies throughout the solar system, e.g., defining crater counting rates. Over the decades, continued study of the Apollo samples by new generations of scientists using new instruments have continued to yield significant new discoveries, including the presence of endogenous water in the Moon and the possible presence of a lunar cataclysm, that in turn has contributed to new models of solar system formation and evolution. The Apollo samples have often been used as a proxy for studying other bodies like Mercury or asteroids. The Apollo samples have also directly contributed to the interpretation of remotely sensed data sets, including their use as ground truth for both Clementine and Lunar Prospector global geochemical maps. Despite the Apollo samples being a static collection, recent efforts will ensure that investigators continue to have access to new samples. For example, there was a recent solicitation for study of previously unopened Apollo samples in vacuum-sealed containers, as well as new access to samples stored frozen or in a He atmosphere. Similarly, the use of X-ray computed tomography as part of the curation process is identifying new clasts within polymict breccias that are available for study. Finally, the MoonDB project is putting all previously published lunar geochemical analyses into a searchable database, which should facilitate new investigations.

Zeigler, Ryan↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗