Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

BAD2matrix: Phylogenomic matrix concatenation, indel coding, and more

Common steps in phylogenomic matrix production include biological sequence concatenation, morphological data concatenation, insertion/deletion (indel) coding, gene content (presence/absence) coding, removing uninformative characters for parsimony analysis, recording with reduced amino acid alphabets, and occupancy filtering. Existing software does not accomplish these tasks on a phylogenomic scale using a single program. BAD2matrix is a Python script that performs the above-mentioned steps in phylogenomic matrix construction for DNA or amino acid sequences as well as morphological data. The script works in UNIX-like environments (e.g., LINUX, MacOS, Windows Subsystem for LINUX).

59 BASIC BIOLOGICAL SCIENCES↗

Algorithm for Compressing Time-Series Data

An algorithm based on Chebyshev polynomials effects lossy compression of time-series data or other one-dimensional data streams (e.g., spectral data) that are arranged in blocks for sequential transmission. The algorithm was developed for use in transmitting data from spacecraft scientific instruments to Earth stations. In spite of its lossy nature, the algorithm preserves the information needed for scientific analysis. The algorithm is computationally simple, yet compresses data streams by factors much greater than two. The algorithm is not restricted to spacecraft or scientific uses: it is applicable to time-series data in general. The algorithm can also be applied to general multidimensional data that have been converted to time-series data, a typical example being image data acquired by raster scanning. However, unlike most prior image-data-compression algorithms, this algorithm neither depends on nor exploits the two-dimensional spatial correlations that are generally present in images. In order to understand the essence of this compression algorithm, it is necessary to understand that the net effect of this algorithm and the associated decompression algorithm is to approximate the original stream of data as a sequence of finite series of Chebyshev polynomials. For the purpose of this algorithm, a block of data or interval of time for which a Chebyshev polynomial series is fitted to the original data is denoted a fitting interval. Chebyshev approximation has two properties that make it particularly effective for compressing serial data streams with minimal loss of scientific information: The errors associated with a Chebyshev approximation are nearly uniformly distributed over the fitting interval (this is known in the art as the "equal error property"); and the maximum deviations of the fitted Chebyshev polynomial from the original data have the smallest possible values (this is known in the art as the "min-max property").

Hawkins, S. Edward, III↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

A novel beta-glucosidase from the cell wall of maize (Zea mays L.): rapid purification and partial characterization

Plants have a variety of glycosidic conjugates of hormones, defense compounds, and other molecules that are hydrolyzed by beta-glucosidases (beta-D-glucoside glucohydrolases, E.C. 3.2.1.21). Workers have reported several beta-glucosidases from maize (Zea mays L.; Poaceae), but have localized them mostly by indirect means. We have purified and partly characterized a 58-Ku beta-glucosidase from maize, which we conclude from a partial sequence analysis, from kinetic data, and from its localization is not identical to any of those already reported. A monoclonal antibody, mWP 19, binds this enzyme, and localizes it in the cell walls of maize coleoptiles. An earlier report showed that mWP19 inhibits peroxidase activity in crude cell wall extracts and can immunoprecipitate peroxidase activity from these extracts, yet purified preparations of the 58 Ku protein had little or no peroxidase activity. The level of sequence similarity between beta-glucosidases and peroxidases makes it unlikely that these enzymes share epitopes in common. Contrary to a previous conclusion, these results suggest that the enzyme recognized by mWP19 is not a peroxidase, but there is a wall peroxidase closely associated with the 58 Ku beta-glucosidase in crude preparations. Other workers also have co-purified distinct proteins with beta-glucosidases. We found no significant charge in the level of immunodetectable beta-glucosidase in mesocotyls or coleoptiles that precedes the red light-induced changes in the growth rate of these tissues.

Non-NASA Center↗

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES↗

Monitoring technology

A process for infrared spectroscopic monitoring of insitu compositional changes in a polymeric material comprises the steps of providing an elongated infrared radiation transmitting fiber that has a transmission portion and a sensor portion, embedding the sensor portion in the polymeric material to be monitored, subjecting the polymeric material to a processing sequence, applying a beam of infrared radiation to the fiber for transmission through the transmitting portion to the sensor portion for modification as a function of properties of the polymeric material, monitoring the modified infrared radiation spectra as the polymeric material is being subjected to the processing sequence to obtain kinetic data on changes in the polymeric material during the processing sequence, and adjusting the processing sequence as a function of the kinetic data provided by the modified infrared radiation spectra information.

Stevenson, William A.↗

Monitoring technology

A process for infrared spectroscopic monitoring of insitu compositional changes in a polymeric material comprises the steps of providing an elongated infrared radiation transmitting fiber that has a transmission portion and a sensor portion, embedding the sensor portion in the polymeric material to be monitored, subjecting the polymeric material to a processing sequence, applying a beam of infrared radiation to the fiber for transmission through the transmitting portion to the sensor portion for modification as a function of properties of the polymeric material, monitoring the modified infrared radiation spectra as the polymeric material is being subjected to the processing sequence to obtain kinetic data on changes in the polymeric material during the processing sequence, and adjusting the processing sequence as a function of the kinetic data provided by the modified infrared radiation spectra information.

Stevenson, William A.↗

Usefulness of LANDSAT data for monitoring plant development and range conditions in California's annual grassland

A network of sampling sites throughout the annual grassland region of California was established to correlate plant growth stages and forage production to climatic and other environmental factors. Plant growth and range conditions were further related to geographic location and seasonal variations. A sequence of LANDSAT data was obtained covering critical periods in the growth cycle. This was analyzed by both photointerpretation and computer aided techniques. Image characteristics and spectral reflectance data were then related to forage production, range condition, range site and changing growth conditions. It was determined that repeat sequences with LANDSAT color composite images do provide a means for monitoring changes in range condition. Spectral radiance data obtained from magnetic tape can be used to determine quantitatively the critical stages in the forage growth cycle. A computer ratioing technique provided a sensitive indicator of changes in growth stages and an indication of the relative differences in forage production between range sites.

Carneggie, D. M.↗

A parallel VLSI architecture for a digital filter using a number theoretic transform

The advantages of a very large scalee integration (VLSI) architecture for implementing a digital filter using fermat number transforms (FNT) are the following: It requires no multiplication. Only additions and bit rotations are needed. It alleviates the usual dynamic range limitation for long sequence FNT's. It utilizes the FNT and inverse FNT circuits 100% of the time. The lengths of the input data and filter sequences can be arbitraty and different. It is regular, simple, and expandable, and as a consequence suitable for VLSI implementation.

Truong, T. K.↗

Weibull Distribution From Interval Inspection Data

Most likely failure sequence assumed. Memorandum discusses application of Weibull distribution to statistics of failures of turbopump blades. Is generalization of well known exponential random probability distribution and useful in describing component-failure modes including aging effects. Parameters found from experimental data by method of maximum likelihood.

Rheinfurth, Mario H.↗

Determining a Prony Series for a Viscoelastic Material From Time Varying Strain Data

In this study a method of determining the coefficients in a Prony series representation of a viscoelastic modulus from rate dependent data is presented. Load versus time test data for a sequence of different rate loading segments is least-squares fitted to a Prony series hereditary integral model of the material tested. A nonlinear least squares regression algorithm is employed. The measured data includes ramp loading, relaxation, and unloading stress-strain data. The resulting Prony series which captures strain rate loading and unloading effects, produces an excellent fit to the complex loading sequence.

Tzikang, Chen↗

Galileo Spacecraft Modeling for Orbital Operations

The Galileo Jupiter orbital mission using the Low Gain Antenna (LGA) requires a higher degree of spacecraft state knowledge than was originally anticipated. Key elements of the revised design include onboard buffering of science and engineering data and extensive processing of data prior to downlink. In order to prevent loss of data resulting from overflow of the buffers and to allow efficient use of the spacecraft resources, ground based models of the spacecraft processes will be implemented. These models will be integral tools in the development of satellite encounter sequences and the cruise/playback sequences where recorded data is retrieved.

aerospace↗

The Mark 3 data base handler

A data base handler which would act to tie Mark 3 system programs together is discussed. The data base handler is written in FORTRAN and is implemented on the Hewlett-Packard 21MX and the IBM 360/91. The system design objectives were to (1) provide for an easily specified method of data interchange among programs, (2) provide for a high level of data integrity, (3) accommodate changing requirments, (4) promote program accountability, (5) provide a single source of program constants, and (6) provide a central point for data archiving. The system consists of two distinct parts: a set of files existing on disk packs and tapes; and a set of utility subroutines which allow users to access the information in these files. Users never directly read or write the files and need not know the details of how the data are formatted in the files. To the users, the storage medium is format free. A user does need to know something about the sequencing of his data in the files but nothing about data in which he has no interest.

Ryan, J. W.↗

Mosaic organization of DNA nucleotides

Long-range power-law correlations have been reported recently for DNA sequences containing noncoding regions. We address the question of whether such correlations may be a trivial consequence of the known mosaic structure ("patchiness") of DNA. We analyze two classes of controls consisting of patchy nucleotide sequences generated by different algorithms--one without and one with long-range power-law correlations. Although both types of sequences are highly heterogenous, they are quantitatively distinguishable by an alternative fluctuation analysis method that differentiates local patchiness from long-range correlations. Application of this analysis to selected DNA sequences demonstrates that patchiness is not sufficient to account for long-range correlation properties.

NASA Discipline Number 14-10↗

Test Input Generation for Red-Black Trees using Abstraction

We consider the problem of test input generation for code that manipulates complex data structures. Test inputs are sequences of method calls from the data structure interface. We describe test input generation techniques that rely on state matching to avoid generation of redundant tests. Exhaustive techniques use explicit state model checking to explore all the possible test sequences up to predefined input sizes. Lossy techniques rely on abstraction mappings to compute and store abstract versions of the concrete states; they explore under-approximations of all the possible test sequences. We have implemented the techniques on top of the Java PathFinder model checker and we evaluate them using a Java implementation of red-black trees.

Visser, Willem↗

The Ribosomal Database Project

The Ribosomal Database Project (RDP) complies ribosomal sequences and related data, and redistributes them in aligned and phylogenetically ordered form to its user community. It also offers various software packages for handling, analyzing and displaying sequences. In addition, the RDP offers (or will offer) certain analytic services. At present the project is in an intermediate stage of development.

Non-NASA Center↗

A genomic timescale of prokaryote evolution: insights into the origin of methanogenesis, phototrophy, and the colonization of land

BACKGROUND: The timescale of prokaryote evolution has been difficult to reconstruct because of a limited fossil record and complexities associated with molecular clocks and deep divergences. However, the relatively large number of genome sequences currently available has provided a better opportunity to control for potential biases such as horizontal gene transfer and rate differences among lineages. We assembled a data set of sequences from 32 proteins (approximately 7600 amino acids) common to 72 species and estimated phylogenetic relationships and divergence times with a local clock method. RESULTS: Our phylogenetic results support most of the currently recognized higher-level groupings of prokaryotes. Of particular interest is a well-supported group of three major lineages of eubacteria (Actinobacteria, Deinococcus, and Cyanobacteria) that we call Terrabacteria and associate with an early colonization of land. Divergence time estimates for the major groups of eubacteria are between 2.5-3.2 billion years ago (Ga) while those for archaebacteria are mostly between 3.1-4.1 Ga. The time estimates suggest a Hadean origin of life (prior to 4.1 Ga), an early origin of methanogenesis (3.8-4.1 Ga), an origin of anaerobic methanotrophy after 3.1 Ga, an origin of phototrophy prior to 3.2 Ga, an early colonization of land 2.8-3.1 Ga, and an origin of aerobic methanotrophy 2.5-2.8 Ga. CONCLUSIONS: Our early time estimates for methanogenesis support the consideration of methane, in addition to carbon dioxide, as a greenhouse gas responsible for the early warming of the Earths' surface. Our divergence times for the origin of anaerobic methanotrophy are compatible with highly depleted carbon isotopic values found in rocks dated 2.8-2.6 Ga. An early origin of phototrophy is consistent with the earliest bacterial mats and structures identified as stromatolites, but a 2.6 Ga origin of cyanobacteria suggests that those Archean structures, if biologically produced, were made by anoxygenic photosynthesizers. The resistance to desiccation of Terrabacteria and their elaboration of photoprotective compounds suggests that the common ancestor of this group inhabited land. If true, then oxygenic photosynthesis may owe its origin to terrestrial adaptations.

Methane/metabolism↗

Crop classification using multidate/multifrequency radar data

Both C- and L-band radar data acquired over a test site near Colby, Kansas during the summer of 1978 were used to identify three types of vegetation cover and bare soil. The effects of frequency, polarization, and the look angle on the overall accuracy of recognizing the four types of ground cover were analyzed. In addition, multidate data were used to study the improvement in recognition accuracy possible with the addition of temporal information. The soil moisture conditions had changed considerably during the temporal sequence of the data; hence, the effects of soil moisture on the ability to discriminate between cover types were also analyzed. The results provide useful information needed for selecting the parameters of a radar system for monitoring crops.

Ulaby, F. T.↗