Search NASASearch

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Development of Computational Environmental Microbiome Workflows for the Laboratory and the International Space Station

Identification of microorganisms in the spaceflight environment is critical for crew health risk assessment on the International Space Station (ISS). Since 2017, nanopore sequencing technology has been used to support thein situ identification of microbial species during spaceflight. Beginning in 2018, a culture-independent, swab-to-sequencer method was implemented onboard the ISS to provide a more thorough insight of the ISS microbiome. Eliminating microbial culture enables identification of difficult-to-culture organisms, reduces risks associated with potentially pathogenic cultures, and could significantly reduce the time from sample-to-answer. However, this molecular-based approach generates large metagenomic datasets that require substantial computational resources for analysis. To process nanopore-generated sequencing data, the JSC Microbiology Laboratory established a bioinformatics workflow on Amazon EC2 under the security guidance of the NASA Science Managed Cloud Environment (SMCE).This resource allows for the development, testing, and accessing of computational tools for processing large and complex datasets. The work described here will address the downlinking of data from the ISS, the automated pipeline developed to identify targeted bacterial and fungal organisms, and the time from sampling onboard to microbial identification. The pipelines have been enhanced to address high and low biomass samples using optimization based on sample source (air, water, or surface) and type of collection (filter, colony, or swab).The resulting microbiome data can be assessed beyond microbial identifications to gain understanding toward population changes over time, potential selective environmental pressures, and evaluating correlations with a wide range of additional data sets. Metagenome analysis pipelines in development could allow for simultaneous identification of microbial species, gene function, and gene pathways present in the environment. Beyond the ground processing, the developed analysis pipeline is currently deployed onboard the ISS to allow for near real-time assessments of the ISS microbiome. This study serves as a critical foundation for exploration missions, where rapid microbiome analyses will be required.

G. Marie Sharp

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey

Optimizing Single Nuclei Sequencing of Brain Samples From Space Flown Mice Across Age and Strain

The NASA GeneLab Sample Processing Laboratory offers high-throughput sequencing services to NASA-funded space biology researchers. Space biology studies have specific challenges such as low sample numbers, introducing susceptibility to batch effects from sample handling. These issues are compounded by complex protocols such as single-nuclei isolation and sequencing, which has recently become an attractive methodology for assessing the cellular diversity within spaceflight samples. High quality single-nuclei sequencing requires reproducible protocols to dissociate tissue and generate clean suspension of intact single nuclei. Producing single-nuclei suspension from brain tissue is particularly challenging due to cell type heterogeneity and the myelin sheath that carries over into the nuclei suspension as debris. Current procedures tend to be time consuming and sometimes include steps that can alter gene expression and create cell-type bias. Commercially available nuclei isolation kits, such as the 10X Genomics nuclei isolation kit, offers a streamlined way to process samples for nuclei isolation, thereby minimizing batch effects and enabling reproducibility. In this study, we report on the performance of the 10X Genomics nuclei isolation kit and Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit to generate sequencing libraries from space-flown mouse brain samples. Single nuclei sequencing was performed on frozen mouse brain tissue from two spaceflight missions, Rodent Research-10 (RR-10) and RR Reference Mission-2 (RRRM-2). RR-10 mice were female B6129SF2/J, euthanized at 18-19 weeks whereas RRRM-2 mice were female C57BL/6NTac, euthanized at 20 or 37 weeks. Sequencing data was processed using standard GeneLab data processing pipelines. We report evaluation of the performance of the 10X Genomics nuclei isolation kit for spaceflight samples from mouse brain, and evaluation of reproducibility across different mouse strains and age groups. We also report preliminary scientific results including cell type inference, cell clustering, and differentially expressed genes and pathways between spaceflight and ground control samples.

RR-10

Revolutionizing Earth Science with Generalized AI Models

Foundation Models (FM) are generalized Artificial Intelligence (AI) models that are designed to replace a task or an application-specific model and can be used for many downstream applications. These FM can be built on any sequence data and are trained utilizing self-supervised approaches. The obstacle of creating a sizable labeled dataset for training is removed by using self-supervised learning. Most FM employ transformer design that takes advantage of the idea of self-attention, allowing the network to represent the impact of distant data points on one another in space and time. The FM models show emergent qualities that are induced from the data. FM can become a valuable tool for Earth science researchers. Due to the size of these models, downstream applications built fine-tuning these FM perform better and exhibit greater accuracy than models created from scratch. FM significantly lowers the entry barrier in terms of both the time and effort required to develop various downstream applications. For some scientific datasets, such as optical remote sensing data, FM can speed up processes like classification, object detection and prediction. By eliminating the training data bottleneck and maximizing the usage of science data, FM can make it simpler to integrate AI into scientific research. Initial results for three different FMs will be presented.

Rahul Ramachandran

AI Foundation Models for Science: An Open Collaborative Initiative

Foundation Models (FMs), AI models designed to replace task-specific models, are increasingly being recognized for their versatility across numerous downstream applications. These models, trained using self-supervised techniques on any type of sequence data, circumvent the need for large annotated datasets, a major bottleneck in traditional AI model development. FMs can be applied to downstream tasks using few-shot learning and fine-tuning, significantly reducing the need for large labeled training datasets and computational resources. However, the development of FMs requires substantial resources, including access to data and compute power, expertise in the latest models, and specialized scientific knowledge for systematic evaluation. It is challenging for a single group to possess all these capabilities. To address this, NASA IMPACT has initiated an open collaborative effort, leveraging partnerships with the private sector and other groups within and outside NASA, to jointly build FMs. The overarching goal is to develop a consistent and collaborative approach to building FMs for high-value science datasets. This initiative has fostered collaboration within NASA and with external partners, including IBM Research, Clark University, DOE’s ORNL, ESA, and USGS. The effort focuses on identifying key datasets with a wide range of downstream applications, pretraining and building FMs using modified transformer architectures, evaluating compute infrastructure needs, and sharing models, pretraining and fine-tuning code, and data with the community. Furthermore, it aims to train the Earth science community to fine-tune these models for various downstream applications. Our initial effort resulted in the creation of a 100 million parameter HLS Geospatial Model within six months, which was released on HuggingFace. We are now expanding our scope to include data from weather and climate models and investigating multimodal models. We invite those interested in participating in this effort to join us by sharing their use cases, expertise, or data.

Rahul Ramachandran

Somatic Mutation Analysis in Spaceflight: NASA Twins Genome Study

The NASA Twins Genome Study investigates the effects of spaceflight on somatic mutation accumulation by comparing genome-wide sequence data from a spaceflight astronaut and his Earth-bound twin. Utilizing advanced computational software on high performance computers, this study identifies and maps somatic mutations, with implications for understanding spaceflight-associated health risks, including cancer, neurodegeneration, and cardiovascular disease. The findings aim to bridge rodent and human space research, offering insights into tissue-specific pathophysiology, risk models, and potential therapeutic interventions.

somatic mutation

Improvements and modifications to the NASA microwave signature acquisition system

A user oriented description of the modified and upgraded Microwave Signature Acquisition System is provided. The present configuration of the sensor system and its operating characteristics are documented and a step-by-step operating procedure provides instruction for mounting the antenna truss assembly, readying the system for data acquisition, and for controlling the system during the data collection sequence. The resulting data products are also identified.

Jean, B. R.

Synchronization Technique For Reception Of Coded Data

Shortest sequence of bits likely to be filled with error bursts examined. Algorithm improves synchronization of frames of noisy binary-coded data signals after Viterbi decoding (recovery from "inner" convolutional code used in transmission channel) and before Reed-Solomon or other decoding (recovery from "outer" error-correcting block code). Based on comparisons of sequences of correct and erroneous Viterbi-decoded received bits with known marker sequence denoting beginning of frame of data. Does not require count of number of bits in received sequence disagreeing with corresponding bits in marker sequence.

Shahshahani, Mehrdad M.

Data compression of discrete sequence: A tree based approach using dynamic programming

A dynamic programming based approach for data compression of a ID sequence is presented. The compression of an input sequence of size N to that of a smaller size k is achieved by dividing the input sequence into k subsequences and replacing the subsequences by their respective average values. The partitioning of the input sequence is carried with the intention of reducing the mean squared error in the reconstructed sequence. The complexity involved in finding the partitions which would result in such an optimal compressed sequence is reduced by using the dynamic programming approach, which is presented.

Shivaram, Gurusrasad

Systematic analysis of coding and noncoding DNA sequences using methods of statistical linguistics

We compare the statistical properties of coding and noncoding regions in eukaryotic and viral DNA sequences by adapting two tests developed for the analysis of natural languages and symbolic sequences. The data set comprises all 30 sequences of length above 50 000 base pairs in GenBank Release No. 81.0, as well as the recently published sequences of C. elegans chromosome III (2.2 Mbp) and yeast chromosome XI (661 Kbp). We find that for the three chromosomes we studied the statistical properties of noncoding regions appear to be closer to those observed in natural languages than those of coding regions. In particular, (i) a n-tuple Zipf analysis of noncoding regions reveals a regime close to power-law behavior while the coding regions show logarithmic behavior over a wide interval, while (ii) an n-gram entropy measurement shows that the noncoding regions have a lower n-gram entropy (and hence a larger "n-gram redundancy") than the coding regions. In contrast to the three chromosomes, we find that for vertebrates such as primates and rodents and for viral DNA, the difference between the statistical properties of coding and noncoding regions is not pronounced and therefore the results of the analyses of the investigated sequences are less conclusive. After noting the intrinsic limitations of the n-gram redundancy analysis, we also briefly discuss the failure of the zeroth- and first-order Markovian models or simple nucleotide repeats to account fully for these "linguistic" features of DNA. Finally, we emphasize that our results by no means prove the existence of a "language" in noncoding DNA.

NASA Discipline Number 14-10

The dynamics of herpesvirus and polyomavirus reactivation and shedding in healthy adults: a 14-month longitudinal study

Humans are infected with viruses that establish long-term persistent infections. To address whether immunocompetent individuals control virus reactivation globally or independently and to identify patterns of sporadic reactivation, we monitored herpesviruses and polyomaviruses in 30 adults, over 14 months. Epstein-Barr virus (EBV) DNA was quantitated in saliva and peripheral blood mononuclear cells (PBMCs), cytomegalovirus (CMV) was assayed in urine, and JC virus (JCV) and BK virus (BKV) DNAs were assayed in urine and PBMCs. All individuals shed EBV in saliva, whereas 67% had >or=1 blood sample positive for EBV. Levels of EBV varied widely. CMV shedding occurred infrequently but occurred more commonly in younger individuals (P<.03). JCV and BKV virurias were 46.7% and 0%, respectively. JCV shedding was age dependent and occurred commonly in individuals >or=40 years old (P<.03). Seasonal variation was observed in shedding of EBV and JCV, but there was no correlation among shedding of EBV, CMV, and JCV (P>.50). Thus, adults independently control persistent viruses, which display discordant, sporadic reactivations.

Non-NASA Center

Archaebacterial phylogeny: perspectives on the urkingdoms

Comparisons of complete 16S ribosomal RNA sequences have been used to confirm, refine and extend earlier concepts of archaebacterial phylogeny. The archaebacteria fall naturally into two major branches or divisions, I--the sulfur-dependent thermophilic archaebacteria, and II--the methanogenic archaebacteria and their relatives. Division I comprises a relatively closely related and phenotypically homogeneous collection of thermophilic sulfur-dependent species--encompassing the genera Sulfolobus, Thermoproteus, Pyrodictium and Desulfurococcus. The organisms of Division II, however, form a less compact grouping phylogenetically, and are also more diverse in phenotype. All three of the (major) methanogen groups are found in Division II, as are the extreme halophiles and two types of thermoacidophiles, Thermoplasma acidophilum and Thermococcus celer. This last species branches sufficiently deeply in the Division II line that it might be considered to represent a separate, third Division. However, both the extreme halophiles and Tp. acidophilum branch within the cluster of methanogens. The extreme halophiles are specifically related to the Methanomicrobiales, to the exclusion of both the Methanococcales and the Methanobacteriales. Tp. acidophilum is peripherally related to the halophile-Methanomicrobiales group. By 16S rRNA sequence measure the archaebacteria constitute a phylogenetically coherent grouping (clade), which excludes both the eubacteria and the eukaryotes--a conclusion that is supported by other sequence evidence as well. Alternative proposals for archaebacterial phylogeny, not based upon sequence evidence, are discussed and evaluated. In particular, proposals to rename (reclassify) various subgroups of the archaebacteria as new kingdoms are found wanting, for both their lack of proper experimental support and the taxonomic confusion they introduce.

NASA Discipline Exobiology

A definition of the domains Archaea, Bacteria and Eucarya in terms of small subunit ribosomal RNA characteristics

The number of small subunit rRNA sequences is now great enough that the three domains Archaea, Bacteria and Eucarya (Woese et al., 1990) can be reliably defined in terms of their sequence "signatures". Approximately 50 homologous positions (or nucleotide pairs) in the small subunit rRNA characterize and distinguish among the three. In addition, the three can be recognized by a variety of nonhomologous rRNA characters, either individual positions and/or higher-order structural features. The Crenarchaeota and the Euryarchaeota, the two archaeal kingdoms, can also be defined and distinguished by their characteristic compositions at approximately fifteen positions in the small subunit rRNA molecule.

NASA Program Exobiology

PCR amplification of 16S rDNA from lyophilized cell cultures facilitates studies in molecular systematics

The sequence of the major portion of a Bacillus cycloheptanicus strain SCH(T) 16S rRNA gene is reported. This sequence suggests that B. cycloheptanicus is genetically quite distinct from traditional Bacillus strains (e.g., B. subtilis) and may be properly regarded as belonging to a different genus. The sequence was determined from DNA that was produced by direct amplification of ribosomal DNA from a lyophilized cell pellet with straightforward polymerase chain reaction (PCR) procedures. By obviating the need to revive cell cultures from the lyophile pellet, this approach facilitates rapid 16S rDNA sequencing and thereby advances studies in molecular systematics.

Non-NASA Center

Molecular cloning and characterization of a tomato cDNA encoding a systemically wound-inducible bZIP DNA-binding protein

Localized wounding of one leaf in intact tomato (Lycopersicon esculentum Mill.) plants triggers rapid systemic transcriptional responses that might be involved in defense. To better understand the mechanism(s) of intercellular signal transmission in wounded tomatoes, and to identify the array of genes systemically up-regulated by wounding, a subtractive cDNA library for wounded tomato leaves was constructed. A novel cDNA clone (designated LebZIP1) encoding a DNA-binding protein was isolated and identified. This clone appears to be encoded by a single gene, and belongs to the family of basic leucine zipper domain (bZIP) transcription factors shown to be up-regulated by cold and dark treatments. Analysis of the mRNA levels suggests that the transcript for LebZIP1 is both organ-specific and up-regulated by wounding. In wounded wild-type tomatoes, the LebZIP1 mRNA levels in distant tissue were maximally up-regulated within only 5 min following localized wounding. Exogenous abscisic acid (ABA) prevented the rapid wound-induced increase in LebZIP1 mRNA levels, while the basal levels of LebZIP1 transcripts were higher in the ABA mutants notabilis (not), sitiens (sit), and flacca (flc), and wound-induced increases were greater in the ABA-deficient mutants. Together, these results suggest that ABA acts to curtail the wound-induced synthesis of LebZIP1 mRNA.

Non-NASA Center

Arabidopsis chloroplast chaperonin 10 is a calmodulin-binding protein

Calcium regulates diverse cellular activities in plants through the action of calmodulin (CaM). By using (35)S-labeled CaM to screen an Arabidopsis seedling cDNA expression library, a cDNA designated as AtCh-CPN10 (Arabidopsis thaliana chloroplast chaperonin 10) was cloned. Chloroplast CPN10, a nuclear-encoded protein, is a functional homolog of E. coli GroES. It is believed that CPN60 and CPN10 are involved in the assembly of Rubisco, a key enzyme involved in the photosynthetic pathway. Northern analysis revealed that AtCh-CPN10 is highly expressed in green tissues. The recombinant AtCh-CPN10 binds to CaM in a calcium-dependent manner. Deletion mutants revealed that there is only one CaM-binding site in the last 31 amino acids of the AtCh-CPN10 at the C-terminal end. The CaM-binding region in AtCh-CPN10 has higher homology to other chloroplast CPN10s in comparison to GroES and mitochondrial CPN10s, suggesting that CaM may only bind to chloroplast CPN10s. Furthermore, the results also suggest that the calcium/CaM messenger system is involved in regulating Rubisco assembly in the chloroplast, thereby influencing photosynthesis. Copyright 2000 Academic Press.

NASA Discipline Plant Biology