Search NASASearch

SEARCH · Search NASA

Results for “sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data compression of discrete sequence: A tree based approach using dynamic programming

A dynamic programming based approach for data compression of a ID sequence is presented. The compression of an input sequence of size N to that of a smaller size k is achieved by dividing the input sequence into k subsequences and replacing the subsequences by their respective average values. The partitioning of the input sequence is carried with the intention of reducing the mean squared error in the reconstructed sequence. The complexity involved in finding the partitions which would result in such an optimal compressed sequence is reduced by using the dynamic programming approach, which is presented.

Shivaram, Gurusrasad

Reporting Differences Between Spacecraft Sequence Files

A suite of computer programs, called seq diff suite, reports differences between the products of other computer programs involved in the generation of sequences of commands for spacecraft. These products consist of files of several types: replacement sequence of events (RSOE), DSN keyword file [DKF (wherein DSN signifies Deep Space Network)], spacecraft activities sequence file (SASF), spacecraft sequence file (SSF), and station allocation file (SAF). These products can include line numbers, request identifications, and other pieces of information that are not relevant when generating command sequence products, though these fields can result in the appearance of many changes to the files, particularly when using the UNIX diff command to inspect file differences. The outputs of prior software tools for reporting differences between such products include differences in these non-relevant pieces of information. In contrast, seq diff suite removes the fields containing the irrelevant pieces of information before processing to extract differences, so that only relevant differences are reported. Thus, seq diff suite is especially useful for reporting changes between successive versions of the various products and in particular flagging difference in fields relevant to the sequence command generation and review process.

Khanampompan, Teerapat

Program Synthesizes UML Sequence Diagrams

A computer program called "Rational Sequence" generates Universal Modeling Language (UML) sequence diagrams of a target Java program running on a Java virtual machine (JVM). Rational Sequence thereby performs a reverse engineering function that aids in the design documentation of the target Java program. Whereas previously, the construction of sequence diagrams was a tedious manual process, Rational Sequence generates UML sequence diagrams automatically from the running Java code.

Barry, Matthew R.

[Characterization of Black and Dichothrix Cyanobacteria Based on the 16S Ribosomal RNA Gene Sequence]

My project focuses on characterizing different cyanobacteria in thrombolitic mats found on the island of Highborn Cay, Bahamas. Thrombolites are interesting ecosystems because of the ability of bacteria in these mats to remove carbon dioxide from the atmosphere and mineralize it as calcium carbonate. In the future they may be used as models to develop carbon sequestration technologies, which could be used as part of regenerative life systems in space. These thrombolitic communities are also significant because of their similarities to early communities of life on Earth. I targeted two cyanobacteria in my research, Dichothrix spp. and whatever black is, since they are believed to be important to carbon sequestration in these thrombolitic mats. The goal of my summer research project was to molecularly identify these two cyanobacteria. DNA was isolated from each organism through mat dissections and DNA extractions. I ran Polymerase Chain Reactions (PCR) to amplify the 16S ribosomal RNA (rRNA) gene in each cyanobacteria. This specific gene is found in almost all bacteria and is highly conserved, meaning any changes in the sequence are most likely due to evolution. As a result, the 16S rRNA gene can be used for bacterial identification of different species based on the sequence of their 16S rRNA gene. Since the exact sequence of the Dichothrix gene was unknown, I designed different primers that flanked the gene based on the known sequences from other taxonomically similar cyanobacteria. Once the 16S rRNA gene was amplified, I cloned the gene into specialized Escherichia coli cells and sent the gene products for sequencing. Once the sequence is obtained, it will be added to a genetic database for future reference to and classification of other Dichothrix sp.

Ortega, Maya

MSLICE Sequencing

MSLICE Sequencing is a graphical tool for writing sequences and integrating them into RML files, as well as for producing SCMF files for uplink. When operated in a testbed environment, it also supports uplinking these SCMF files to the testbed via Chill. This software features a free-form textural sequence editor featuring syntax coloring, automatic content assistance (including command and argument completion proposals), complete with types, value ranges, unites, and descriptions from the command dictionary that appear as they are typed. The sequence editor also has a "field mode" that allows tabbing between arguments and displays type/range/units/description for each argument as it is edited. Color-coded error and warning annotations on problematic tokens are included, as well as indications of problems that are not visible in the current scroll range. "Quick Fix" suggestions are made for resolving problems, and all the features afforded by modern source editors are also included such as copy/cut/paste, undo/redo, and a sophisticated find-and-replace system optionally using regular expressions. The software offers a full XML editor for RML files, which features syntax coloring, content assistance and problem annotations as above. There is a form-based, "detail view" that allows structured editing of command arguments and sequence parameters when preferred. The "project view" shows the user s "workspace" as a tree of "resources" (projects, folders, and files) that can subsequently be opened in editors by double-clicking. Files can be added, deleted, dragged-dropped/copied-pasted between folders or projects, and these operations are undoable and redoable. A "problems view" contains a tabular list of all problems in the current workspace. Double-clicking on any row in the table opens an editor for the appropriate sequence, scrolling to the specific line with the problem, and highlighting the problematic characters. From there, one can invoke "quick fix" as described above to resolve the issue. Once resolved, saving the file causes the problem to be removed from the problem view.

Crockett, Thomas M.

Spaceflight Autonomous Multigenerational Microbial Sequencer in Support of Plant-Growth Systems

The CubeSat platform has proven successful in obtaining meaningful life science information when biological payloads are incorporated. Examples include: 1) the first-ever CubeSat with a biological payload, GeneSat-1, which demonstrated decreased growth rates for flight samples of Escherichia coli in low Earth orbit (Parra et al. 2008); 2) PharmaSat, demonstrated that Saccharomyces cerevisiae in the microgravity environment exhibits a significant level of metabolic activity even at high doses of applied antifungal (Ricco et al. 2011); 3) O/OREOS, which used Bacillus subtilis(bacteria) to demonstrate for the first time that microorganisms can be loaded in a dried, dormant form and then rehydrated and grown in orbit months after launch (Nicholson et al. 2011; Ehrenfreund et al. 2014; 4) the SporeSat payload, which investigated Ceratopteris richardii(fern spores) using lab-on-a-chip devices (BioCDs) and minicentrifuges to produce artificial gravitational forces in ground studies (Park et al. 2017), with demonstration of the BioCD and minicentrifuge in space; 5) EcAMSat, the first CubeSat to be directly deployed from the ISS for an experiment assessing antibiotic resistance of E. coli in the microgravity environment (Padgen et al. 2020); 6) BioSentinel, exposed a culture of yeast to galactic cosmic radiation (GCR) and solar particle events while in heliocentric orbit to measure the rate of double-strand-break repair using DNA-repair-deficient mutants. This effort measures the metabolic parameters of yeast in a deep-space environment compared to Earth ambient conditions using a 3-color LED detection system (Ricco et al. 2020; Padgen et al. 2021). We aim to expand this list to include a Spaceflight Autonomous Multigenerational Microbial Sequencer (SAMMS). SAMMS will allow for the genome level understanding of changes in growth and metabolic activity for any organism. While microbes are suitable for early studies in our proposed platform because of their small size, small and relatively less-complicated genomes, fast generation times, and relevance to life support systems; multicellular organisms can similarly be evaluated for their genetic response to the spaceflight environment. The Spaceflight Autonomous Multigenerational Microbial Sequencer (SAMMS) will enable autonomous sequencing of biological samples in plant production units, cislunar orbit and on the lunar surface to examine spaceflight effects (ie. radiation, altered gravity, reduced pressures) on plant and microbial genomes.On this team a Kennedy Space Center (KSC) space crop production and water systems microbiologist/molecular biologist works with a Johnson Space Center (JSC) International Space Station (ISS) microbial sequencing expert and an Ames Research Center (ARC) CubeSat Engineering team to convert an automated Oxford Nanopore librarypreparation and sequencing method to a fluidic CubeSat payload system. The Oxford Nanopore MinION sequencing platform has proven successful in the spaceflight environment onboard the ISS (Stahl-Rommel et al. 2021). Further long-duration spaceflight and exposure to high levels of radiation will cause genotypic effects in biological organisms that may affect their function. Monitoring the adaption of a population to the spaceflight environment and any subsequent beneficial mutations will allow for the harnessing of organisms best suited for use in life support systems. This will ensure that the selected life support-essential microorganisms maintain their intended specified function over generations of culturing in the relevant spaceflight environment without becoming hazardous to crew or spacecraft systems.

Aubrie E Orourke

Size and Structure of the Sequence Space of Repeat Proteins

The coding space of protein sequences is shaped by evolutionary constraints set by requirements of function and stability. We show that the coding space of a given protein family— the total number of sequences in that family—can be estimated using models of maximum entropy trained on multiple sequence alignments of naturally occurring amino acid sequences. We analyzed and calculated the size of three abundant repeat proteins families, whose members are large proteins made of many repetitions of conserved portions of *30 amino acids. While amino acid conservation at each position of the alignment explains most of the reduction of diversity relative to completely random sequences, we found that correlations between amino acid usage at different positions significantly impact that diversity. We quantified the impact of different types of correlations, functional and evolutionary, on sequence diversity. Analysis of the detailed structure of the coding space of the families revealed a rugged landscape, with many local energy minima of varying sizes with a hierarchical structure, reminiscent of frustrated energy landscapes of spin glass in physics. This clustered structure indicates a multiplicity of subtypes within each family and suggests new strategies for protein design.

Jacopo Marchi

On the Natural Structure of Amino Acid Patterns in Families of Protein Sequences

All known terrestrial proteins are coded as continuous strings of ≈20 amino acids. The patterns formed by the repetitions of elements in groups of finite sequences describes the natural architectures of protein families. We present a method to search for patterns and groupings of patterns in protein sequences using a mathematically precise definition for “repetition”, an efficient algorithmic implementation and a robust scoring system with no adjustable parameters. We show that the sequence patterns can be well-separated into disjoint classes according to their recurrence in nested structures. The statistics of the occurrences of patterns indicate that short repetitions are sufficient to account for the differences between natural families and randomized groups of sequences by more than 10 standard deviations, while contiguous sequence patterns shorter than 5 residues are effectively random in their occurrences. A small subset of patterns is sufficient to account for a robust ”familiarity” definition between arbitrary sets of sequences.

Pablo Turjanski

Optimizing Single Nuclei Sequencing of Brain Samples From Space Flown Mice Across Age and Strain

The NASA GeneLab Sample Processing Laboratory offers high-throughput sequencing services to NASA-funded space biology researchers. Space biology studies have specific challenges such as low sample numbers, introducing susceptibility to batch effects from sample handling. These issues are compounded by complex protocols such as single-nuclei isolation and sequencing, which has recently become an attractive methodology for assessing the cellular diversity within spaceflight samples. High quality single-nuclei sequencing requires reproducible protocols to dissociate tissue and generate clean suspension of intact single nuclei. Producing single-nuclei suspension from brain tissue is particularly challenging due to cell type heterogeneity and the myelin sheath that carries over into the nuclei suspension as debris. Current procedures tend to be time consuming and sometimes include steps that can alter gene expression and create cell-type bias. Commercially available nuclei isolation kits, such as the 10X Genomics nuclei isolation kit, offers a streamlined way to process samples for nuclei isolation, thereby minimizing batch effects and enabling reproducibility. In this study, we report on the performance of the 10X Genomics nuclei isolation kit and Chromium Next GEM Single Cell Multiome ATAC + Gene Expression kit to generate sequencing libraries from space-flown mouse brain samples. Single nuclei sequencing was performed on frozen mouse brain tissue from two spaceflight missions, Rodent Research-10 (RR-10) and RR Reference Mission-2 (RRRM-2). RR-10 mice were female B6129SF2/J, euthanized at 18-19 weeks whereas RRRM-2 mice were female C57BL/6NTac, euthanized at 20 or 37 weeks. Sequencing data was processed using standard GeneLab data processing pipelines. We report evaluation of the performance of the 10X Genomics nuclei isolation kit for spaceflight samples from mouse brain, and evaluation of reproducibility across different mouse strains and age groups. We also report preliminary scientific results including cell type inference, cell clustering, and differentially expressed genes and pathways between spaceflight and ground control samples.

RR-10

Statistical properties of DNA sequences

We review evidence supporting the idea that the DNA sequence in genes containing non-coding regions is correlated, and that the correlation is remarkably long range--indeed, nucleotides thousands of base pairs distant are correlated. We do not find such a long-range correlation in the coding regions of the gene. We resolve the problem of the "non-stationarity" feature of the sequence of base pairs by applying a new algorithm called detrended fluctuation analysis (DFA). We address the claim of Voss that there is no difference in the statistical properties of coding and non-coding regions of DNA by systematically applying the DFA algorithm, as well as standard FFT analysis, to every DNA sequence (33301 coding and 29453 non-coding) in the entire GenBank database. Finally, we describe briefly some recent work showing that the non-coding sequences have certain statistical features in common with natural and artificial languages. Specifically, we adapt to DNA the Zipf approach to analyzing linguistic texts. These statistical properties of non-coding sequences support the possibility that non-coding regions of DNA may carry biological information.

Non-NASA Center

Systematic analysis of coding and noncoding DNA sequences using methods of statistical linguistics

We compare the statistical properties of coding and noncoding regions in eukaryotic and viral DNA sequences by adapting two tests developed for the analysis of natural languages and symbolic sequences. The data set comprises all 30 sequences of length above 50 000 base pairs in GenBank Release No. 81.0, as well as the recently published sequences of C. elegans chromosome III (2.2 Mbp) and yeast chromosome XI (661 Kbp). We find that for the three chromosomes we studied the statistical properties of noncoding regions appear to be closer to those observed in natural languages than those of coding regions. In particular, (i) a n-tuple Zipf analysis of noncoding regions reveals a regime close to power-law behavior while the coding regions show logarithmic behavior over a wide interval, while (ii) an n-gram entropy measurement shows that the noncoding regions have a lower n-gram entropy (and hence a larger "n-gram redundancy") than the coding regions. In contrast to the three chromosomes, we find that for vertebrates such as primates and rodents and for viral DNA, the difference between the statistical properties of coding and noncoding regions is not pronounced and therefore the results of the analyses of the investigated sequences are less conclusive. After noting the intrinsic limitations of the n-gram redundancy analysis, we also briefly discuss the failure of the zeroth- and first-order Markovian models or simple nucleotide repeats to account fully for these "linguistic" features of DNA. Finally, we emphasize that our results by no means prove the existence of a "language" in noncoding DNA.

NASA Discipline Number 14-10

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Non-NASA Center

Comparison of the spatial statistics of random and defined-sequence photoresist films

The resolution-line edge roughness-sensitivity tradeoff has motivated the exploration of potential improvements using defined sequence polymers and polymer-bound photoacid generators and quenchers. We characterize the internal structures of positive tone photoresist polymer films formed from defined sequence polymers and compare them with random copolymers of the same composition. We model their imaging to connect initially to developable film structures. We use a polymer packing algorithm to simulate films of diverse compositions and locations of photoacid generators and quenchers, using the composition of an ESCAP photoresist. We use a simple extreme ultraviolet exposure-deprotection algorithm to model developable image formation within them. In all cases, the spatial distribution of chemical moieties in the film for defined sequence polymers is nearly indistinguishable from random copolymers. We evaluate several exposure-deprotection scenarios and find that a defined sequence copolymer has a distinctive developable image under certain circumstances. The use of defined sequence polymers within a photoresist layer does not automatically result in improved imaging; however, they do have some characteristics different from random polymers of the same composition. Further study of these characteristics may provide a route to improved control over the nanoscale imaging process.

36 MATERIALS SCIENCE

Spitzer Space Telescope Sequencing Operations Software, Strategies, and Lessons Learned

The Space Infrared Telescope Facility (SIRTF) was launched in August, 2003, and renamed to the Spitzer Space Telescope in 2004. Two years of observing the universe in the wavelength range from 3 to 180 microns has yielded enormous scientific discoveries. Since this magnificent observatory has a limited lifetime, maximizing science viewing efficiency (ie, maximizing time spent executing activities directly related to science observations) was the key operational objective. The strategy employed for maximizing science viewing efficiency was to optimize spacecraft flexibility, adaptability, and use of observation time. The selected approach involved implementation of a multi-engine sequencing architecture coupled with nondeterministic spacecraft and science execution times. This approach, though effective, added much complexity to uplink operations and sequence development. The Jet Propulsion Laboratory (JPL) manages Spitzer s operations. As part of the uplink process, Spitzer s Mission Sequence Team (MST) was tasked with processing observatory inputs from the Spitzer Science Center (SSC) into efficiently integrated, constraint-checked, and modeled review and command products which accommodated the complexity of non-deterministic spacecraft and science event executions without increasing operations costs. The MST developed processes, scripts, and participated in the adaptation of multi-mission core software to enable rapid processing of complex sequences. The MST was also tasked with developing a Downlink Keyword File (DKF) which could instruct Deep Space Network (DSN) stations on how and when to configure themselves to receive Spitzer science data. As MST and uplink operations developed, important lessons were learned that should be applied to future missions, especially those missions which employ command-intensive operations via a multi-engine sequence architecture.

missions operations

Monitoring Astronaut Health with DNA Sequencing

In recent years microbe a plethora of microbe populations have been identified onboard the ISS (International Space Station). Approaches for real-time tracking of microbes for routine housekeeping and food/water safety monitoring will be critical for mission safety and crew health on future longer duration missions to the Moon or Mars. This work is a proof-of-concept study demonstrating an end-to-end phylogenetic identification and full genome sequencing effort of multiple microbial populations. Our methodology utilized the ISS flight-certified WetLab-2 molecular toolbox and the Biomolecule Sequencer projects for real-time end-to-end on-orbit microbial biological samples processing and molecular analysis with real time results generated utilizing only field "offline" analytic software. For this experiment we colony-cultured several ISS isolated microorganisms before generation of the pre-sequencing library via the automated VolTRAX device which enabled high library turnover with little wet-bench activity or potential future costly astronaut time. The pre-sequencing library is diluted in loading buffer and injected into the MinION sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown. 16S and full genome alignment, nucleotide matching, gene identification, and phylogenetic sorting was accomplished utilizing the Epi2me software and the offline NCBI Blast viral, microbiome, and human somatic databases. In short, the methodologies developed herein replace the myriad of specific, often highly targeted microbiological tests used in the clinical laboratory, which would be difficult if not impossible to currently implement aboard the ISS or in deep space, with a single metagenomics test.

genomics

Novel End-to-End Molecular Biology Approach for Direct Nanopore 1D cDNA Sequencing of Reverse Transcribed mRNAs Purified from Cell Cultures by the NASA ISS WetLab2 SPM

Continued space bioscience research onboard the International Space Station (ISS) and future long-duration flight missions to the Moon or Mars will require the ability to conduct on-orbit molecular analysis of biological samples independently from Earth. In the last year two new molecular analytic technologies have been installed and the technologies demonstrated onboard the ISS: The Sample Prep Module (SPM) WetLab-2 (WL2) qRT-PCR toolbox and the Oxford Nanopore MinIon Biomolecule Sequencer. Here we describe protocol development and integration into existing ISS technology for end-to-end on-orbit biological sample processing and molecular analysis with real time results generated utilizing only field offline analytic software. For this experiment we isolated primary cells from bone marrow flushes of wild type B6129SF2 mice (Jackson Labs) long bones. The cell isolate was then processed using the SPM to produce total 147nanograms of RNA. The total RNA was purified to only messenger RNA (mRNA) and transferred to Smartcycler Thermocycle ISS kit consumable tube using Eppendorf gel loading pipette tips for further processing. Complementary first strand cDNA was synthesized using OLIGO dT priming followed by addition of SuperScript II Reverse Transcriptase and thermal cycling as per manufacturers instruction. All thermal cycling was conducted using the ISS WetLab-2 Cephid Smarcycler real time thermal cycler. Our protocol takes advantage of mRNAs native poly(A) tail, synthesized in vivo to protect the mRNA from degradation by endonucleases, to eliminate end-prep for adapter ligation. The adapted library is purified using MyOne C1 Streptavidin beads before elution in buffer. The pre-sequencing library is diluted in the loading buffer and injected into the MinIon sample port, drawn into the nanopore window by capillary action, and sequenced using the MinKnown software with local basecalling. The sequencing read produced 34.5 million events and local basecalling produced 117,301 successful reads. NCBI Blast of the data for the mouse genome resulted in 2,462 successful nucleotide collection matches (gene sequences) exceeding 70 homology. These results demonstrate the viability of this novel flight ready end-to-end sample analytic methodology and provide a real time homolog for flight experimentation utilizing supply kits and technologies that have already been demonstrated on ISS.

MinIon

Iterative pass optimization of sequence data

The problem of determining the minimum-cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete. This "tree alignment" problem has motivated the considerable effort placed in multiple sequence alignment procedures. Wheeler in 1996 proposed a heuristic method, direct optimization, to calculate cladogram costs without the intervention of multiple sequence alignment. This method, though more efficient in time and more effective in cladogram length than many alignment-based procedures, greedily optimizes nodes based on descendent information only. In their proposal of an exact multiple alignment solution, Sankoff et al. in 1976 described a heuristic procedure--the iterative improvement method--to create alignments at internal nodes by solving a series of median problems. The combination of a three-sequence direct optimization with iterative improvement and a branch-length-based cladogram cost procedure, provides an algorithm that frequently results in superior (i.e., lower) cladogram costs. This iterative pass optimization is both computation and memory intensive, but economies can be made to reduce this burden. An example in arthropod systematics is discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

NASA Discipline Evolutionary Biology

5S ribosomal ribonucleic acid sequences in Bacteroides and Fusobacterium: evolutionary relationships within these genera and among eubacteria in general

The 5S ribosomal ribonucleic acid (rRNA) sequences were determined for Bacteroides fragilis, Bacteroides thetaiotaomicron, Bacteroides capillosus, Bacteroides veroralis, Porphyromonas gingivalis, Anaerorhabdus furcosus, Fusobacterium nucleatum, Fusobacterium mortiferum, and Fusobacterium varium. A dendrogram constructed by a clustering algorithm from these sequences, which were aligned with all other hitherto known eubacterial 5S rRNA sequences, showed differences as well as similarities with respect to results derived from 16S rRNA analyses. In the 5S rRNA dendrogram, Bacteroides clustered together with Cytophaga and Fusobacterium, as in 16S rRNA analyses. Intraphylum relationships deduced from 5S rRNAs suggested that Bacteroides is specifically related to Cytophaga rather than to Fusobacterium, as was suggested by 16S rRNA analyses. Previous taxonomic considerations concerning the genus Bacteroides, based on biochemical and physiological data, were confirmed by the 5S rRNA sequence analysis.

NASA Discipline Exobiology