Search NASA⌕ Search

SEARCH · Search NASA

Results for “biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

Physics through the 1990s: Scientific interfaces and technological applications

The volume examines the scientific interfaces and technological applications of physics. Twelve areas are dealt with: biological physics-biophysics, the brain, and theoretical biology; the physics-chemistry interface-instrumentation, surfaces, neutron and synchrotron radiation, polymers, organic electronic materials; materials science; geophysics-tectonics, the atmosphere and oceans, planets, drilling and seismic exploration, and remote sensing; computational physics-complex systems and applications in basic research; mathematics-field theory and chaos; microelectronics-integrated circuits, miniaturization, future trends; optical information technologies-fiber optics and photonics; instrumentation; physics applications to energy needs and the environment; national security-devices, weapons, and arms control; medical physics-radiology, ultrasonics, MNR, and photonics. An executive summary and many chapters contain recommendations regarding funding, education, industry participation, small-group university research and large facility programs, government agency programs, and computer database needs.

Source record↗

VEG: An intelligent workbench for analysing spectral reflectance data

An Intelligent Workbench (VEG) was developed for the systematic study of remotely sensed optical data from vegetation. A goal of the remote sensing community is to infer the physical and biological properties of vegetation cover (e.g. cover type, hemispherical reflectance, ground cover, leaf area index, biomass, and photosynthetic capacity) using directional spectral data. VEG collects together, in a common format, techniques previously available from many different sources in a variety of formats. The decision as to when a particular technique should be applied is nonalgorithmic and requires expert knowledge. VEG has codified this expert knowledge into a rule-based decision component for determining which technique to use. VEG provides a comprehensive interface that makes applying the techniques simple and aids a researcher in developing and testing new techniques. VEG also provides a classification algorithm that can learn new classes of surface features. The learning system uses the database of historical cover types to learn class descriptions of one or more classes of cover types.

Harrison, P. Ann↗

New insight into the global record of the Ediacaran tubular morphotype: a common solution to early multicellularity

The tubular morphogroup is a common component of Earth’s first complex, multicellular communities—the Ediacaran biota—and offers valuable insight into biological traits that are fundamental to animal life because they have intriguing links to metazoan phyla and are highly abundant in Ediacaran ecosystems. Biomineral tubes (e.g. Cloudina) are well described from the Nama assemblage (~550–538 Myr), yielding a relatively detailed understanding of this subset of the morphogroup. Conversely, the non-biomineral tubular taxa of the Nama assemblage, as well as of the older White Sea assemblage (~560–550 Myr), are poorly understood. As a result, the variability of characters that define non-biomineral tubular organisms is unknown and their diversity dynamics throughout the terminal Ediacaran are unconstrained. To test hypotheses related to the diversity, morphological variability and temporal distribution of non-biomineral tubes, a comprehensive database of non-biomineral Ediacaran tubular taxa was compiled. Results demonstrate previously unrecognized morphological disparity in the non-biomineral tubular morphogroup and reveal that it comprises a higher number of genera than all other non-tubular morphogroups in the White Sea and the Nama. Thus, it illustrates that a tubular form dominated Ediacaran ecosystems for considerably longer than previously appreciated and, importantly, was the most common solution to early multicellularity.

Rachel L. Surprenant↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

Sensitive Probing of Exoplanetary Oxygen via Mid-Infrared Collisional Absorption

The collision-induced fundamental vibration–rotation band at 6.4 μm is the strongest absorption feature from O2 in the infrared, yet it has not been previously incorporated into exoplanet spectral analyses for several reasons. Either collision-induced absorptions (CIAs) were not included or incomplete/obsolete CIA databases were used. Also, the current version of HITRAN does not include CIAs at 6.4 μm with other collision partners (O2–X). We include O2–X CIA features in our transmission spectroscopy simulations by parameterizing the 6.4-μm O2–N2 CIA based on ref. 3 and the O2–CO2 CIA based on ref. 4. Here we report that the O2–X CIA may be the most detectable O2 feature for transit observations. For a potential TRAPPIST-1 e analogue system within 5 pc of the Sun, it could be the only O2 signature detectable with the James Webb Space Telescope (JWST) (using MIRI LRS (Mid-Infrared Instrument low-resolution spectrometer)) for a modern Earth-like cloudy atmosphere with biological quantities of O2. Also, we show that the 6.4-μm O2–X CIA would be prominent for O2-rich desiccated atmospheres and could be detectable with JWST in just a few transits. For systems beyond 5 pc, this feature could therefore be a powerful discriminator of uninhabited planets with non-biological ‘false-positive’ O2 in their atmospheres, as they would only be detectable at these higher O2 pressures.

Thomas J Fauchez↗

Sensitive Probing of Exoplanetary Oxygen via Mid-Infrared Collisional Absorption

The collision-induced fundamental vibration–rotation band at 6.4 μm is the strongest absorption feature from O2 in the infrared1,2,3, yet it has not been previously incorporated into exoplanet spectral analyses for several reasons. Either collision-induced absorptions (CIAs) were not included or incomplete/obsolete CIA databases were used. Also, the current version of HITRAN does not include CIAs at 6.4 μm with other collision partners (O2–X). We include O2–X CIA features in our transmission spectroscopy simulations by parameterizing the 6.4-μm O2–N2 CIA based on ref. 3 and the O2–CO2 CIA based on ref. 4. Here we report that the O2–X CIA may be the most detectable O2 feature for transit observations. For a potential TRAPPIST-1 e analogue system within 5 pc of the Sun, it could be the only O2 signature detectable with the James Webb Space Telescope (JWST) (using MIRI LRS (Mid-Infrared Instrument low-resolution spectrometer)) for a modern Earth-like cloudy atmosphere with biological quantities of O2. Also, we show that the 6.4-μm O2–X CIA would be prominent for O2-rich desiccated atmospheres5 and could be detectable with JWST in just a few transits. For systems beyond 5 pc, this feature could therefore be a powerful discriminator of uninhabited planets with non-biological ‘false-positive’ O2 in their atmospheres, as they would only be detectable at these higher O2 pressures.

Thomas Fauchez↗

SIMBIOS Project Data Processing and Analysis Results

The Sensor Intercomparison and Merger for Biological and Interdisciplinary Oceanic Studies (SIMBIOS) Project is concerned with ocean color satellite sensor data intercomparison and merger for biological and interdisciplinary studies of the global oceans. Imagery from different ocean color sensors can now be processed by a single software package using the same algorithms, adjusted by different sensor spectral characteristics, and the same ancillary meteorological and environmental data. This enables cross-comparison and validation of the data derived from satellite sensors and, consequently, creates continuity in ocean color information on both the temporal and spatial scale. The next step in this process is the integration of in situ ocean and atmospheric parameters to enable cross-validation and further refinement of the ocean color methodology. The SIMBIOS Project Office accomplishments during 2000 year are summarized under satellite data processing, data product validation, SeaWiFS Bio-Optical Archive and Storage System (SeaBASS) database, supporting services, sun photometers and calibration activities, and calibration round robins. These accomplishments are described.

Ainsworth, Ewa↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing a Small RNAseq Analysis Pipeline for NASA GeneLab Using Open-Source Tools and Libraries

Small RNA sequencing (small RNAseq) is a powerful tool for studying the regulation of gene expression in various organisms. Small RNAseq has been leveraged in space biology research to study how expression of small RNAs, e.g. micro RNAs (miRNAs), small interfering RNAs (siRNAs), and piwi-interacting RNAs (piRNAs), change upon exposure to the space environment. NASA GeneLab currently hosts small RNAseq raw data derived from space-relevant experiments on the Open Science Data Repository (OSDR). To maximize the accessibility of these data to the scientific community, in addition to hosting raw data, which is only interpretable by bioinformaticians, GeneLab plans to process all small RNAseq datasets and make those processed data available to the scientific community via the OSDR. In this study, we present the development of the GeneLab standardized pipeline for processing small RNAseq datasets. Using human, plant, and synthetic small RNAseq datasets, we interrogate various open-source software and publicly available databases to evaluate their accuracy and reproducibility in each step of the pipeline. For quality control and adapter detection and trimming, we evaluated TrimGalore!, FASTX, SeqKit, and DNApi methods to optimize alignment to reference genomes. We compared BWA, Bowtie, and Bowtie2 to determine the optimal alignment tool. For each alignment tool we also assessed various reference databases, including Ensembl reference genomes and different types of small RNA reference databases, including genome, hairpin, and miRNA references from the miRbase and MirGeneDB databases. To quantify the aligned data, we compared SAMtools, HTSeq, and RSEM for counting alignment events from each alignment tool used. Finally, we evaluated various tools, including DESeq2 and EdgeR, for data normalization and subsequent differential expression analysis. We will present the results from our comparative analyses for each pipeline step and propose a consensus pipeline for processing small RNAseq data derived from various organisms exposed to the space environment.

SmallRNAseq, NASA GeneLab, quality control, adapte↗

Validation of Carbon Flux and Related Products for SIMBIOS: The CARIACO Continental Margin Time Series and the Orinoco River Plume

Between 1997 and 2000, this Sensor Intercomparison and Merger for Biological and Interdisciplinary Oceanic Studies (SIMBIOS) investigation collected bio-optical measurements in the Southeastern Caribbean Sea and the tropical western Atlantic to help understand the color of coastal and continental shelf waters. Specifically, bio-optical data were collected to complement an oceanographic time series maintained within the Cariaco Basin, a site affected by seasonal coastal upwelling. Bio-optical data were also collected within the plume of the Orinoco River during seasonal extremes in discharge. This program focused on providing data to the Sea-Viewing Wide Field-of-view Sensor (SeaWiFS) and SIMBIOS Projects for validating SeaWiFS products. The data are unique in that they provide a substantial number of observations on repeated seasonal cycles for the SeaWIFS Bio-Optical Archive and Storage System (SeaBASS) bio-optical database. An important aspect of this SIMBIOS investigation was a focus on proper interpretation of ocean color remote sensing data from coastal and continental shelf environments. With this goal in mind, ocean color satellite data from a variety and locations and from different satellite sensors were examined to understand spatial and temporal variability in pigment concentrations, and also to conduct an in-depth study of current atmospheric correction and bio-optical algorithms.

Muller-Karger, Frank↗

The Containment Assurance Risk Framework of the Mars Sample Return Program

The Mars Sample Return campaign aims at bringing rock and atmospheric samples from Mars to Earth through a series of robotic missions. These missions would collect the samples being cached and deposited on Martian soil by the Perseverance rover, place them in a container, and launch them into Martian orbit for subsequent capture by an orbiter that would bring them back. Given there exists a non-zero probability that the samples contain biological material, precautions are being taken to design systems that would break the chain of contact between Mars and Earth. These include techniques such as sterilization of Martian particles, redundant containment vessels, and a robust reentry capsule capable of accurate landings without a parachute. Requirements exist that the probability of containment not assured of Martian-contaminated material into Earth’s biosphere be less than one in a million. To demonstrate compliance with this strict requirement, a statistical framework was developed to assess the likelihood of containment loss during each sample return phase and make a statement about the total combined mission probability of containment not assured. The work presented here describes this framework, which considers failure modes or fault conditions that can initiate failure sequences ultimately leading to containment not assured. Reliability estimates are generated from databases, design heritage, component specifications, or expert opinion in the form of probability density functions or point estimates and provided as inputs to the mathematical models that simulate the different failure sequences. The probabilistic outputs are then combined following the logic of several fault trees to compute the ultimate probability of containment not assured. Given the multidisciplinary nature of the problem and the different types of mathematical models used, the statistical tools needed for analysis are required to be computationally efficient. While standard Monte Carlo approaches are used for fast models, a multi-fidelity approach to rare event probabilities is proposed for expensive models. In this paradigm, inexpensive low-fidelity models are developed for computational acceleration purposes while the expensive high-fidelity model is kept in the loop to retain accuracy in the results. This work presents an example of end-to-end application of this framework highlighting the computational benefits of a multi-fidelity approach.

Giuseppe Cataldo↗

The NASA Open Science Data Repository: Biomedical Data, Analysis Tools, and Informatic Collaborations

Increased biomedical risks and challenges associated with deep space missions require knowledge discovery, health countermeasures, and biomedical support capabilities. Maximally open-access and reusable data is needed by developers, scientists, and engineers to develop these systems. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database (ie., findable, accessible, interoperable, and reusable), and meets various scientific, technical, and operational needs. It offers users and submitters the ability to upload, download, search, share, analyze, cite, and visualize data across ‘omics, physiological, phenotypic, payload, hardware, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR is an expanded database, based upon the successes of NASA GeneLab. OSDR has >460 studies with datasets covering model organisms to non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets with raw files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) developed from industry norms. OSDR is collecting and curating biomedical human data from a new sub-orbital research flight and is open to more space life science/biomedical submissions from the international and commercial sectors. OSDR also recently began a collaboration with the European Space Agency (ESA) to collect and curate >200 terabytes of human and model organism data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics and ~50 physiological-phenotypic-imaging assay data types. Tools available for OSDR users include: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, and 3) a Multi-study visualization tool which enables users to look across and combine ‘omics datasets. There are ~600 volunteer OSDR Analysis Working Group (AWG) members providing feedback on scientific data/metadata standards and collaborating to mine-reuse OSDR in research. OSDR/GeneLab has enabled ~60 publications reusing data as of October 2023.

space biology↗

Curating NASA's Past, Present, and Future Extraterrestrial Sample Collections

As codified in NASA Policy Directive 7100.10F, the Astromaterials Acquisition and Curation Office at NASA Johnson Space Center (hereafter JSC Curation) is charged with curation of all extraterrestrial material under NASA control, including future NASA missions. JSC Curation curates all or part of nine astromaterial collections in seven clean room suites: (1) Apollo Samples (1969; ISO 6-7), (2) Luna Samples (from USSR; 1972; ISO 7), (3) Antarctic Meteorites (1976; ISO 7), (4) Cosmic Dust (1981; ISO 5), (5) Microparticle Impact Collection (formerly called Space Exposed Hardware; 1985; ISO 5), (6) Genesis Solar Wind Atoms (2004; ISO 4); (7) Stardust Comet Particles (2006; ISO 5), (8) Stardust Interstellar Particles (2006; ISO 5), (9) Hayabusa Asteroid Particles (from JAXA; 2010; ISO 5). In addition to the labs that house the samples, we have installed and maintained a wide variety of facilities and infrastructure required to support the clean-rooms: more than 10 different HEPA-filtered air-handling systems, ultrapure dry gaseous nitrogen systems, an ultrapure water system (UPW) and cleaning facilities to provide clean tools and equipment for the labs. We also have sample preparation facilities for making thin sections, microtome sections, and even focused ion-beam (FIB) sections to meet the research requirements of scientists. To ensure that we are keeping the samples as pristine as possible, we routinely monitor the cleanliness of our clean rooms and infrastructure systems. This monitoring includes: daily monitoring of the quality of our UPW, weekly airborne particle counts in the labs, monthly monitoring of the stable isotope composition of the gaseous N2 system, and annual measurements of inorganic or organic contamination in processing cabinets. We track within our databases the current and ever-changing characteristics of more than 250,000 individual samples across our various collections (including the 19,141 samples on loan to 433 Principal Investigators in 24 countries). The next sample return missions that NASA will participate in are Hayabusa2 and OSIRIS-REx (Origins Spectral Interpretation Resource Identification Security - Regolith Explorer). The designs for a new state-of-the-art suite of clean rooms to house these samples at JSC have been finalized. This includes separate ISO class 5 clean rooms to house each collection, a common ISO class 7 area for general use, an ISO class 7 microtome laboratory, and a separate thin section lab. Additionally, a new cleaning facility is being designed and procedures developed that will allow for enhanced cleaning of cabinets and tools in an inorganically, organically, and biologically clean manner. We are also designing a large multi-purpose Advanced Curation laboratory that will allow us to develop the techniques necessary to fully support the Hayabusa2 and OSIRIS-REx missions, as well as future possible sample return missions (e.g., Lunar Polar Volatiles, Mars, Comet Surface). A micro-CT (micro Computed Tomography) laboratory dedicated to the study of astromaterials has come online within JSC Curation, and we plan to add additional facilities that will enable non-destructive (or minimally-destructive) analyses of astromaterials in the near future (e.g., micro-XRF (micro X-Ray Fluorescence), confocal imaging Raman Spectroscopy). These facilities will be available to: (1) develop sample handling and storage techniques for future sample return missions, (2) be utilized by PET (Positron Emission Tomography) for future sample return missions, (3) for retroactive PET-style analyses of our existing collections, and (4) for periodic assessments of the existing sample collections.

Zeigler, Ryan A.↗

Effects of light on brain and behavior

It is obvious that light entering the eye permits the sensory capacity of vision. The human species is highly dependent on visual perception of the environment and consequently, the scientific study of vision and visual mechanisms is a centuries old endeavor. Relatively new discoveries are now leading to an expanded understanding of the role of light entering the eye in addition to supporting vision, light has various nonvisual biological effects. Over the past thirty years, animal studies have shown that environmental light is the primary stimulus for regulating circadian rhythms, seasonal cycles, and neuroendocrine responses. As with all photobiological phenomena, the wavelength, intensity, timing and duration of a light stimulus is important in determining its regulatory influence on the circadian and neuroendocrine systems. Initially, the effects of light on rhythms and hormones were observed only in sub-human species. Research over the past decade, however, has confirmed that light entering the eyes of humans is a potent stimulus for controlling physiological rhythms. The aim of this paper is to examine three specific nonvisual responses in humans which are mediated by light entering the eye: light-induced melatonin suppression, light therapy for winter depression, and enhancement of nighttime performance. This will serve as a brief introduction to the growing database which demonstrates how light stimuli can influence physiology, mood and behavior in humans. Such information greatly expands our understanding of the human eye and will ultimately change our use of light in the human environment.

Brainard, George C.↗