Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

MISIP: a data standard for the reuse and reproducibility of any stable isotope probing-derived nucleic acid sequence and experiment

DNA/RNA-stable isotope probing (SIP) is a powerful tool to link in situ microbial activity to sequencing data. Every SIP dataset captures distinct information about microbial community metabolism, process rates, and population dynamics, offering valuable insights for a wide range of research questions. Data reuse maximizes the information derived from the labor and resource-intensive SIP approaches. Yet, a review of publicly available SIP sequencing metadata showed that critical information necessary for reproducibility and reuse was often missing. Here, we outline the Minimum Information for any Stable Isotope Probing Sequence (MISIP) according to the Minimum Information for any (x) Sequence (MIxS) framework and include examples of MISIP reporting for common SIP experiments. Our objectives are to expand the capacity of MIxS to accommodate SIP-specific metadata and guide SIP users in metadata collection when planning and reporting an experiment. The MISIP standard requires 5 metadata fields—isotope, isotopolog, isotopolog label, labeling approach, and gradient position—and recommends several fields that represent best practices in acquiring and reporting SIP sequencing data (e.g., gradient density and nucleic acid amount). The standard is intended to be used in concert with other MIxS checklists to comprehensively describe the origin of sequence data, such as for marker genes (MISIP-MIMARKS) or metagenomes (MISIP-MIMS), in combination with metadata required by an environmental extension (e.g., soil). The adoption of the proposed data standard will improve the reuse of any sequence derived from a SIP experiment and, by extension, deepen understanding of in situ biogeochemical processes and microbial ecology.

Simpson, Abigayle

A standards perspective on genomic data reusability and reproducibility

Genomic and metagenomic sequence data provides an unprecedented ability to re-examine findings, offering a transformative potential for advancing research, developing computational tools, enhancing clinical applications, and fostering scientific collaboration. However, effective and ethical reuse of genomics data is hampered by numerous technical and social challenges. The International Microbiome and Multi’Omics Standards Alliance (IMMSA, https://www.microbialstandards.org/) and the Genomic Standards Consortium (GSC, https://gensc.org) hosted a 5-part seminar series “A Year of Data Reuse” in 2024 to explore challenges and opportunities of data reuse and reproducibility across disparate domains of the genomic sciences. Addressing these challenges will require a multifaceted approach, including common metadata reporting, clear communication, standardized protocols, improved data management infrastructure, ethical guidelines, and collaborative policies that prioritize transparency and accessibility. We offer strategies to enable responsible and technically feasible data reuse, recognition of data reproducibility challenges, and emphasizing the importance of cross-disciplinary efforts in the pursuit of open science and data-driven innovation.

59 BASIC BIOLOGICAL SCIENCES

Evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus

The extension of ecological tolerance limits may be an important mechanism by which microorganisms adapt to novel environments, but it may come at the evolutionary cost of reduced performance under ancestral conditions. We combined a comparative physiological approach with phylogenetic analyses to study the evolution of thermotolerance in hot spring cyanobacteria of the genus Synechococcus. Among the 20 laboratory clones of Synechococcus isolated from collections made along an Oregon hot spring thermal gradient, four different 16S rRNA gene sequences were identified. Phylogenies constructed by using the sequence data indicated that the clones were polyphyletic but that three of the four sequence groups formed a clade. Differences in thermotolerance were observed for clones with different 16S rRNA gene sequences, and comparison of these physiological differences within a phylogenetic framework provided evidence that more thermotolerant lineages of Synechococcus evolved from less thermotolerant ancestors. The extension of the thermal limit in these bacteria was correlated with a reduction in the breadth of the temperature range for growth, which provides evidence that enhanced thermotolerance has come at the evolutionary cost of increased thermal specialization. This study illustrates the utility of using phylogenetic comparative methods to investigate how evolutionary processes have shaped historical patterns of ecological diversification in microorganisms.

Synechococcus Group/classification/growth & develo

Extensions to the Dynamic Aerospace Vehicle Exchange Markup Language

The Dynamic Aerospace Vehicle Exchange Markup Language (DAVE-ML) is a syntactical language for exchanging flight vehicle dynamic model data. It provides a framework for encoding entire flight vehicle dynamic model data packages for exchange and/or long-term archiving. Version 2.0.1 of DAVE-ML provides much of the functionality envisioned for exchanging aerospace vehicle data; however, it is limited in only supporting scalar time-independent data. Additional functionality is required to support vector and matrix data, abstracting sub-system models, detailing dynamics system models (both discrete and continuous), and defining a dynamic data format (such as time sequenced data) for validation of dynamics system models and vehicle simulation packages. Extensions to DAVE-ML have been proposed to manage data as vectors and n-dimensional matrices, and record dynamic data in a compatible form. These capabilities will improve the clarity of data being exchanged, simplify the naming of parameters, and permit static and dynamic data to be stored using a common syntax within a single file; thereby enhancing the framework provided by DAVE-ML for exchanging entire flight vehicle dynamic simulation models.

Brian, Geoffrey J.

Comparative sequence analyses on the 16S rRNA (rDNA) of Bacillus acidocaldarius, Bacillus acidoterrestris, and Bacillus cycloheptanicus and proposal for creation of a new genus, Alicyclobacillus gen. nov

Comparative 16S rRNA (rDNA) sequence analyses performed on the thermophilic Bacillus species Bacillus acidocaldarius, Bacillus acidoterrestris, and Bacillus cycloheptanicus revealed that these organisms are sufficiently different from the traditional Bacillus species to warrant reclassification in a new genus, Alicyclobacillus gen. nov. An analysis of 16S rRNA sequences established that these three thermoacidophiles cluster in a group that differs markedly from both the obligately thermophilic organisms Bacillus stearothermophilus and the facultatively thermophilic organism Bacillus coagulans, as well as many other common mesophilic and thermophilic Bacillus species. The thermoacidophilic Bacillus species B. acidocaldarius, B. acidoterrestris, and B. cycloheptanicus also are unique in that they possess omega-alicylic fatty acid as the major natural membranous lipid component, which is a rare phenotype that has not been found in any other Bacillus species characterized to date. This phenotype, along with the 16S rRNA sequence data, suggests that these thermoacidophiles are biochemically and genetically unique and supports the proposal that they should be reclassified in the new genus Alicyclobacillus.

Non-NASA Center

Archaeal translation initiation revisited: the initiation factor 2 and eukaryotic initiation factor 2B alpha-beta-delta subunit families

As the amount of available sequence data increases, it becomes apparent that our understanding of translation initiation is far from comprehensive and that prior conclusions concerning the origin of the process are wrong. Contrary to earlier conclusions, key elements of translation initiation originated at the Universal Ancestor stage, for homologous counterparts exist in all three primary taxa. Herein, we explore the evolutionary relationships among the components of bacterial initiation factor 2 (IF-2) and eukaryotic IF-2 (eIF-2)/eIF-2B, i.e., the initiation factors involved in introducing the initiator tRNA into the translation mechanism and performing the first step in the peptide chain elongation cycle. All Archaea appear to posses a fully functional eIF-2 molecule, but they lack the associated GTP recycling function, eIF-2B (a five-subunit molecule). Yet, the Archaea do posses members of the gene family defined by the (related) eIF-2B subunits alpha, beta, and delta, although these are not specifically related to any of the three eukaryotic subunits. Additional members of this family also occur in some (but by no means all) Bacteria and even in some eukaryotes. The functional significance of the other members of this family is unclear and requires experimental resolution. Similarly, the occurrence of bacterial IF-2-like molecules in all Archaea and in some eukaryotes further complicates the picture of translation initiation. Overall, these data lend further support to the suggestion that the rudiments of translation initiation were present at the Universal Ancestor stage.

NASA Discipline Exobiology

Lessons learned supporting onboard solid-state recorders

With the advance of semiconductor technology, Solid-State Recorders (SSR) have matured and been accepted as primary onboard data storage devices. Their high reliability, simpler interface and control, and high flexibility have made the SSR's a superb choice in today's spacecraft design. While there are many benefits, the use of SSR's may also add significant complexity to ground data systems. For instance, real-time and playback data may be interleaved into the same data stream, making data sequencing and time ordering difficult. Stored data may be played back out of time order, increasing processing load significantly. Data may also be played back after being sorted by Virtual Channels in the SSR, potentially creating bursts in packet rates that exceed the real-time processing capabilities of the ground systems. This paper presents a summary of lessons learned through the efforts in supporting a number of NASA's missions that employ SSR's. It describes various problems encountered through the design process, and their potential impact on ground system performance, resources, and cost. Recommended approaches to minimizing the impact are demonstrated by examples. The discussion leads to the conclusion that the use of SSR's demands an even higher level of cooperation between spacecraft and ground system designers in order to build the most cost effective end-to-end system.

Shi, Jeff

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi

Data acquisition and processing history for the Explorer 33 (AIMP-D) satellite

The quality control monitoring system, using accounting and quality control data bases, made it possible to perform an in-depth analysis. Results show that the percentage of useable data files for experimenter analysis was 97.7%; only 0.4% of the data sequences supplied to the experimenter exhibited missing data. The 50 percentile probability delay values (referenced to station record data) indicate that the analog tapes arrived within 11 days, the data were digitized within 4.2 weeks, and the experimenter tapes were delivered in 8.95 weeks or less.

Karras, T. J.

Novel Thermotolerant Siderophilic Filamentous Cyanobacterium that Produces Intracellular Iron-Rich Phases

Cyanobacteria are the main producers of organic compounds in iron-depositing hot springs despite photosynthetically generated-oxygen and the abundance of reduced iron (Fe2+) that likely leads to enormous oxidative stress within cyanobacterial cells. Therefore, the study of cyanobacterial diversity, phylogeny, and biogeochemical activity in iron-depositing hot springs will not only provide insights into the contribution of CB to iron redox cycling in these environments, but it could also provide insights into CB evolution. This study characterizes the phylogeny, morphology, and physiology of isolate JSC-1, a novel filamentous CB isolated from an iron-depositing hot spring. While isolate JSC-1 is morphologically similar to the CB genus Leptolyngbya, 16S rDNA sequence data indicated that it shares 95 percent sequence similarity to the type strain L. boryanum. Strain JSC-1 fixes N2 and exhibited an unusually high ratio between photosystem (PS) I and PS II and was capable of complementary chromatic adaptation. Further, it synthesized only chlorophyll a and a unique set of carotenoids. Strain JSC-1 not only required high levels of Fe for growth (greater than or equal to 40 microM), but it also accumulated large amounts of extracellular ferrihydrite and generated intracellular ferric phosphates. Strain JSC-1 was found to secrete 2-oxoglutaric acid and possesses one ortholog and one paralog of bacterioferritin. Surprisingly, the latter has 70.13 % identity with a bacterioferritin in marine-proteobacterium HTCC 2080 and has joint node with bacterioferritins found in enterobacteria. Collectively, these observations provide insights into the physiological strategies that might have allowed CB to develop and proliferate in Fe-rich environments. Based on its genotypic and phenotypic characterization of strain, JSC-1 represents a new operational taxonomical unit (OTU) JSC-1.

Broun, Igor I.

Purification and immunolocalization of an annexin-like protein in pea seedlings

As part of a study to identify potential targets of calcium action in plant cells, a 35-kDa, annexin-like protein was purified from pea (Pisum sativum L.) plumules by a method used to purify animal annexins. This protein, called p35, binds to a phosphatidylserine affinity column in a calcium-dependent manner and binds 45Ca2+ in a dot-blot assay. Preliminary sequence data confirm a relationship for p35 with the annexin family of proteins. Polyclonal antibodies have been raised which recognize p35 in Western and dot blots. Immunofluorescence and immunogold techniques were used to study the distribution and subcellular localization of p35 in pea plumules and roots. The highest levels of immunostain were found in young developing vascular cells producing wall thickenings and in peripheral root-cap cells releasing slime. This localization in cells which are actively involved in secretion is of interest because one function suggested for the animal annexins is involvement in the mediation of exocytosis.

NASA Discipline Number 40-50

Evolution of hematopoiesis: Three members of the PU.1 transcription factor family in a cartilaginous fish, Raja eglanteria

T lymphocytes and B lymphocytes are present in jawed vertebrates, including cartilaginous fishes, but not in jawless vertebrates or invertebrates. The origins of these lineages may be understood in terms of evolutionary changes in the structure and regulation of transcription factors that control lymphocyte development, such as PU.1. The identification and characterization of three members of the PU.1 family of transcription factors in a cartilaginous fish, Raja eglanteria, are described here. Two of these genes are orthologs of mammalian PU.1 and Spi-C, respectively, whereas the third gene, Spi-D, is a different family member. In addition, a PU.1-like gene has been identified in a jawless vertebrate, Petromyzon marinus (sea lamprey). Both DNA-binding and transactivation domains are highly conserved between mammalian and skate PU.1, in marked contrast to lamprey Spi, in which similarity is evident only in the DNA-binding domain. Phylogenetic analysis of sequence data suggests that the appearance of Spi-C may predate the divergence of the jawed and jawless vertebrates and that Spi-D arose before the divergence of the cartilaginous fish from the lineage leading to the mammals. The tissue-specific expression patterns of skate PU.1 and Spi-C suggest that these genes share regulatory as well as structural properties with their mammalian orthologs.

NASA Discipline Evolutionary Biology

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields, in part due to a culture of open data sharing and reuse. AI/ML methodology is well-suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Inexperienced researchers can produce models that perform poorly outside of the training dataset. Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Casaletto

Phylogeny of the ammonia-producing ruminal bacteria Peptostreptococcus anaerobius, Clostridium sticklandii, and Clostridium aminophilum sp. nov

In previous studies, gram-positive bacteria which grew rapidly with peptides or an amino acid as the sole energy source were isolated from bovine rumina. Three isolates, strains C, FT (T = type strain), and SR, were considered to be ecologically important since they produced up to 20-fold more ammonia than other ammonia-producing ruminal bacteria. On the basis of phenotypic criteria, the taxonomic position of these new isolates was uncertain. In this study, the 16S rRNA sequences of these isolates and related bacteria were determined to establish the phylogenetic positions of the organisms. The sequences of strains C, FT, and SR and reference strains of Peptostreptococcus anaerobius, Clostridium sticklandii, Clostridium coccoides, Clostridium aminovalericum, Acetomaculum ruminis, Clostridium leptum, Clostridium lituseburense, Clostridium acidiurici, and Clostridium barkeri were determined by using a modified Sanger dideoxy chain termination method. Strain C, a large coccus purported to belong to the genus Peptostreptococcus, was closely related to P. anaerobius, with a level of sequence similarity of 99.6%. Strain SR, a heat-resistant, short, rod-shaped organism, was closely related to C. sticklandii, with a level of sequence similarity of 99.9%. However, strain FT, a heat-resistant, pleomorphic, rod-shaped organism, was only distantly related to some clostridial species and P. anaerobius. On the basis of the sequence data, it was clear that strain FT warranted designation as a separate species. The closest known relative of strain FT was C. coccoides (level of similarity, only 90.6%). Additional strains that are phenotypically similar to strain FT were isolated in this study.(ABSTRACT TRUNCATED AT 250 WORDS).

NASA Program Exobiology

wastewater_virus

This repo contains software used to clean and assemble high-throughput sequencing data containing viruses. The input is raw illumina sequencing reads and the output is a database of high-quality viral genomes. The specific application is to wastewater viral concentrates but it is not restricted to that sample type. The software is composed of Nextflow workflows and a set of custom Python and bash scripts that call publicly available bioinformatics tools to accomplish obvious tasks in data analysis in a high performance computing environment. For detailed information, please see the repo's README file.

Kantor, Rose [Lawrence Livermore National Laborato

A new version of the RDP (Ribosomal Database Project)

The Ribosomal Database Project (RDP-II), previously described by Maidak et al. [ Nucleic Acids Res. (1997), 25, 109-111], is now hosted by the Center for Microbial Ecology at Michigan State University. RDP-II is a curated database that offers ribosomal RNA (rRNA) nucleotide sequence data in aligned and unaligned forms, analysis services, and associated computer programs. During the past two years, data alignments have been updated and now include >9700 small subunit rRNA sequences. The recent development of an ObjectStore database will provide more rapid updating of data, better data accuracy and increased user access. RDP-II includes phylogenetically ordered alignments of rRNA sequences, derived phylogenetic trees, rRNA secondary structure diagrams, and various software programs for handling, analyzing and displaying alignments and trees. The data are available via anonymous ftp (ftp.cme.msu. edu) and WWW (http://www.cme.msu.edu/RDP). The WWW server provides ribosomal probe checking, approximate phylogenetic placement of user-submitted sequences, screening for possible chimeric rRNA sequences, automated alignment, and a suggested placement of an unknown sequence on an existing phylogenetic tree. Additional utilities also exist at RDP-II, including distance matrix, T-RFLP, and a Java-based viewer of the phylogenetic trees that can be used to create subtrees.

Non-NASA Center

Developing Open-Source Training Materials for AI/ML and Space Biological Sciences Using NASA Cloud-Based Data

Artificial Intelligence (AI) and Machine Learning (ML) has gained significant traction in the biological and biomedical research fields in the last two decades, in part thanks to an increasing culture of open data sharing and reuse. Due to its capability for identifying complex relationships and patterns, AI/ML methodology is particularly well suited to recognize and predict biological patterns from high-dimensional next-generation sequencing data (e.g. whole genome sequencing, transcriptomic sequencing), as well as from biological or medical imaging data (e.g. microscopy, computed tomography, ultrasound, magnetic resonance imaging, radiography). These methodologies hold particular promise for space biosciences research and automated space health monitoring systems. However, there are many key considerations for properly training, validating, and testing a machine learning model in biological research or clinical application. Even with the positive culture of Open Science and data sharing, inexperienced researchers working quickly without proper checks can produce models that perform poorly outside of the immediate training dataset. Lessons learned from biological AI/ML research indicate that Open Science principles such as data sharing and open-source code must go hand-in-hand with publicly available, high-quality training curricula in best practices, with modules centered on real-life scientific use cases and data so future AI/ML practitioners gain experience on real problems. Here we present the development of open-source training materials for AI/ML and space biosciences, as part of the NASA Transform to Open Science Training (TOPST) initiative. We develop 4 independent training programs, focused on the following topics: 1) Fundamentals of Machine Learning and Space Biosciences Domain, 2) Open Science, Artificial Intelligence, and Ethical Best Practices for Data Sharing and Analysis, 3) Using AI/ML Classification to Identify Gene Networks Affected By Space Exposure in Mouse Liver, and 4) Using Neural Networks to Find DNA Damage Patterns in Immune Cells after Radiation. All programs leverage cloud-based NASA biological datasets. The curriculum we present will enable worldwide access to training in AI/ML and scientific analysis.

James Andrew Casaletto