Search NASASearch

SEARCH · Search NASA

Results for “Sequence Annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

30 records · Page 2

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits, two catalytic centers, common amino acid sequence, and specific contacts to provide a flip-flop, or alternate site, mechanism of action. Each catalytic center [PP:PYR] is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and aminopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core [PP:PYR]* within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GX@&(G)@XXGQ, and GDGX25-30 within the PP- domain, and the E&(G)@XXG@ within the PYR-domain, where Q, corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, Paulina M.

Local frustration around enzyme active sites

Conflicting biological goals often meet in the specification of protein sequences for structure and function. Overall, strong energetic conflicts are minimized in folded native states according to the principle of minimal frustration, so that a sequence can spontaneously fold, but local violations of this principle open up the possibility to encode the complex energy landscapes that are required for active biological functions. We survey the local energetic frustration patterns of all protein enzymes with known structures and experimentally annotated catalytic residues. In agreement with previous hypotheses, the catalytic sites themselves are often highly frustrated regardless of the protein oligomeric state, overall topology, and enzymatic class. At the same time a secondary shell of more weakly frustrated interactions surrounds the catalytic site itself. We evaluate the conservation of these energetic signatures in various family members of major enzyme classes, showing that local frustration is evolutionarily more conserved than the primary structure itself.

Maria I. Freiberger

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits and two catalytic centers. Each catalytic center (PP:PYR) is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and amhopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core (PP:PYR)(sub 2) within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GXPhiX(sub 4)(G)PhiXXGQ and GDGX(sub 25-30)NN in the PP-domain, and the EX(sub 4)(G)PhiXXGPhi in the PYR-domain, where Phi corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, P.

Verification of Numerical Algorithms

The following strategy is suggested for specification and proof: (1) Defer the construction of a formal program specification with respect to I/O assertions unit the correctness of the program with respect to an abstract mathematical model of program intent is demonstrated. (2) Prove that an abstract machine (using infinite precision arithmetic) would compute that object exactly. (3) Prove that the computational sequences of arithmetic operations that occur in the abstract machine must be precisely the same at every step as those occurring on an actual machine (with finite precision arithmetic), executing the same program. (4) Use a Verification Conditions VC-generator that knows about the semantics of arithmetic operations to annotate the program with assertions that bound (or in some circumstances estimate) the difference between the actual machine state variables and the corresponding ones of the abstract machine. Construct the formal program specification by combining the verification conditions into theorems about computational error that can be proved with mechanical assistance.

Source record

Identification of transcribed sequences in Arabidopsis thaliana by using high-resolution genome tiling arrays

Using a maskless photolithography method, we produced DNA oligonucleotide microarrays with probe sequences tiled throughout the genome of the plant Arabidopsis thaliana. RNA expression was determined for the complete nuclear, mitochondrial, and chloroplast genomes by tiling 5 million 36-mer probes. These probes were hybridized to labeled mRNA isolated from liquid grown T87 cells, an undifferentiated Arabidopsis cell culture line. Transcripts were detected from at least 60% of the nearly 26,330 annotated genes, which included 151 predicted genes that were not identified previously by a similar genome-wide hybridization study on four different cell lines. In comparison with previously published results with 25-mer tiling arrays produced by chromium masking-based photolithography technique, 36-mer oligonucleotide probes were found to be more useful in identifying intron-exon boundaries. Using two-dimensional HPLC tandem mass spectrometry, a small-scale proteomic analysis was performed with the same cells. A large amount of strongly hybridizing RNA was found in regions "antisense" to known genes. Similarity of antisense activities between the 25-mer and 36-mer data sets suggests that it is a reproducible and inherent property of the experiments. Transcription activities were also detected for many of the intergenic regions and the small RNAs, including tRNA, small nuclear RNA, small nucleolar RNA, and microRNA. Expression of tRNAs correlates with genome-wide amino acid usage.

Arabidopsis/genetics

Tissue Photolithography

Tissue lithography will enable physicians and researchers to obtain macromolecules with high purity (greater than 90 percent) from desired cells in conventionally processed, clinical tissues by simply annotating the desired cells on a computer screen. After identifying the desired cells, a suitable lithography mask will be generated to protect the contents of the desired cells while allowing destruction of all undesired cells by irradiation with ultraviolet light. The DNA from the protected cells can be used in a number of downstream applications including DNA sequencing. The purity (i.e., macromolecules isolated form specific cell types) of such specimens will greatly enhance the value and information of downstream applications. In this method, the specific cells are isolated on a microscope slide using photolithography, which will be faster, more specific, and less expensive than current methods. It relies on the fact that many biological molecules such as DNA are photosensitive and can be destroyed by ultraviolet irradiation. Therefore, it is possible to protect the contents of desired cells, yet destroy undesired cells. This approach leverages the technologies of the microelectronics industry, which can make features smaller than 1 micrometer with photolithography. A variety of ways has been created to achieve identification of the desired cell, and also to designate the other cells for destruction. This can be accomplished through chrome masks, direct laser writing, and also active masking using dynamic arrays. Image recognition is envisioned as one method for identifying cell nuclei and cell membranes. The pathologist can identify the cells of interest using a microscopic computerized image of the slide, and appropriate custom software. In one of the approaches described in this work, the software converts the selection into a digital mask that can be fed into a direct laser writer, e.g. the Heidelberg DWL66. Such a machine uses a metalized glass plate (with chrome metallization) on which there is a thin layer of photoresist. The laser transfers the digital mask onto the photoresist by direct writing, with typical best resolution of 2 micrometers. The plate is then developed to remove the exposed photoresist, which leaves the exposed areas susceptible to chemical chrome etch. The etch removes the unprotected chrome. The rest of the photoresist is then removed, by either ultraviolet organic solvent or over-development. The remaining chrome pattern is quickly oxidized by atmospheric exposure (typically within 30 seconds). The ready chrome mask is now applied to the tissue slide and aligned manually, or using automatic software and pre-designed alignment marks. The slide plate sandwich is then exposed to UV to destroy the DNA of the unwanted cells. The slide and plate are separated and the slide is processed in a standard way to prepare for polymerase chain reaction (PCR) and potential identification of cancer sequences.

Wade, Lawrence A.

The Conservation of Structure and Mechanism of Catalytic Action in a Family of Thiamin Pyrophosphate (TPP)-dependent Enzymes

Thiamin pyrophosphate (TPP)-dependent enzymes are a divergent family of TPP and metal ion binding proteins that perform a wide range of functions with the common decarboxylation steps of a -(O=)C-C(OH)- fragment of alpha-ketoacids and alpha- hydroxyaldehydes. To determine how structure and catalytic action are conserved in the context of large sequence differences existing within this family of enzymes, we have carried out an analysis of TPP-dependent enzymes of known structures. The common structure of TPP-dependent enzymes is formed at the interface of four alpha/beta domains from at least two subunits, which provide for two metal and TPP-binding sites. Residues around these catalytic sites are conserved for functional purpose, while those further away from TPP are conserved for structural reasons. Together they provide a network of contacts required for flip-flop catalytic action within TPP-dependent enzymes. Thus our analysis defines a TPP-action motif that is proposed for annotating TPP-dependent enzymes for advancing functional proteomics.

Dominiak, P.

Decoding the effects of synonymous variants

Synonymous single nucleotide variants (sSNVs) are common in the human genome but are often overlooked. However, sSNVs can have significant biological impact and may lead to disease. Existing computational methods for evaluating the effect of sSNVs suffer from the lack of gold-standard training/evaluation data and exhibit over-reliance on sequence conservation signals. We developed synVep (synonymous Variant effect predictor), a machine learning-based method that overcomes both of these limitations. Our training data was a combination of variants reported by gnomAD (observed) and those unreported, but possible in the human genome (generated). We used positive-unlabeled learning to purify the generated variant set of any likely unobservable variants. We then trained two sequential extreme gradient boosting models to identify subsets of the remaining variants putatively enriched and depleted in effect. Our method attained 90% precision/recall on a previously unseen set of variants. Furthermore, although synVep does not explicitly use conservation, its scores correlated with evolutionary distances between orthologs in cross-species variation analysis. synVep was also able to differentiate pathogenic vs. benign variants, as well as splice-site disrupting variants (SDV) vs. non-SDVs. Thus, synVep provides an important improvement in annotation of sSNVs, allowing users to focus on variants that most likely harbor effects.

Zishuo Zeng

SequenceL: Automated Parallel Algorithms Derived from CSP-NT Computational Laws

With the introduction of new parallel architectures like the cell and multicore chips from IBM, Intel, AMD, and ARM, as well as the petascale processing available for highend computing, a larger number of programmers will need to write parallel codes. Adding the parallel control structure to the sequence, selection, and iterative control constructs increases the complexity of code development, which often results in increased development costs and decreased reliability. SequenceL is a high-level programming language that is, a programming language that is closer to a human s way of thinking than to a machine s. Historically, high-level languages have resulted in decreased development costs and increased reliability, at the expense of performance. In recent applications at JSC and in industry, SequenceL has demonstrated the usual advantages of high-level programming in terms of low cost and high reliability. SequenceL programs, however, have run at speeds typically comparable with, and in many cases faster than, their counterparts written in C and C++ when run on single-core processors. Moreover, SequenceL is able to generate parallel executables automatically for multicore hardware, gaining parallel speedups without any extra effort from the programmer beyond what is required to write the sequen tial/singlecore code. A SequenceL-to-C++ translator has been developed that automatically renders readable multithreaded C++ from a combination of a SequenceL program and sample data input. The SequenceL language is based on two fundamental computational laws, Consume-Simplify- Produce (CSP) and Normalize-Trans - pose (NT), which enable it to automate the creation of parallel algorithms from high-level code that has no annotations of parallelism whatsoever. In our anecdotal experience, SequenceL development has been in every case less costly than development of the same algorithm in sequential (that is, single-core, single process) C or C++, and an order of magnitude less costly than development of comparable parallel code. Moreover, SequenceL not only automatically parallelizes the code, but since it is based on CSP-NT, it is provably race free, thus eliminating the largest quality challenge the parallelized software developer faces.

Cooke, Daniel

Adaptive Distributed Environment for Procedure Training (ADEPT)

ADEPT (Adaptive Distributed Environment for Procedure Training) is designed to provide more effective, flexible, and portable training for NASA systems controllers. When creating a training scenario, an exercise author can specify a representative rationale structure using the graphical user interface, annotating the results with instructional texts where needed. The author's structure may distinguish between essential and optional parts of the rationale, and may also include "red herrings" - hypotheses that are essential to consider, until evidence and reasoning allow them to be ruled out. The system is built from pre-existing components, including Stottler Henke's SimVentive instructional simulation authoring tool and runtime. To that, a capability was added to author and exploit explicit control decision rationale representations. ADEPT uses SimVentive's Scalable Vector Graphics (SVG)- based interactive graphic display capability as the basis of the tool for quickly noting aspects of decision rationale in graph form. The ADEPT prototype is built in Java, and will run on any computer using Windows, MacOS, or Linux. No special peripheral equipment is required. The software enables a style of student/ tutor interaction focused on the reasoning behind systems control behavior that better mimics proven Socratic human tutoring behaviors for highly cognitive skills. It supports fast, easy, and convenient authoring of such tutoring behaviors, allowing specification of detailed scenario-specific, but content-sensitive, high-quality tutor hints and feedback. The system places relatively light data-entry demands on the student to enable its rationale-centered discussions, and provides a support mechanism for fostering coherence in the student/ tutor dialog by including focusing, sequencing, and utterance tuning mechanisms intended to better fit tutor hints and feedback into the ongoing context.

Domeshek, Eric

Flight Mechanics Modeling and Simulation of the Earth Entry System

Introduction: The Mars Sample Return (MSR) Campaign being planned by NASA and ESA has the ambitious goal to return Mars samples back to Earth. This international collaboration had developed a concept of operations that included a ESA-designed Earth Return Orbiter (ERO) and NASA-designed Capture, Containment, and Return System (CCRS). The Earth Entry System (EES), consisting of a protective aeroshell that houses the samples as well as sample containment vessels, would conduct entry, descent, and landing (EDL) on a direct Earth trajectory. The EES would enter on a spin-stabilized ballistic trajectory with the goal to passively achieve aerodynamic stability throughout all regions of flight. The EDL sequence would end with the EES impacting the soft playa soil of the Utah Test and Training Range (UTTR). As of the submission of this abstract, the MSR campaign is undergoing a re-architecture leading to a pause in EES development. However, the novel approaches developed in flight mechanics modeling and simulation can significantly benefit the greater IPPW community in the development of Earth return vehicles. This paper will present the latest state of EES flight mechanics modeling and simulation. The paper will highlight the simulation architecture developed and key lessons learned from understanding of EDL trajectory sensitivities. Modeling and Simulation: Figure 1 provides a high-level concept of operations for the approach, entry, descent, and landing (AEDL) phase of the CCRS-portion of MSR. The objective of EES flight mechanics is to model and simulate the EES trajectory from ERO separation to ground impact at UTTR. A variety of flight mechanics simulation models were utilized to model both exo-atmopsheric and atmospheric portions of flight. 42, a 6-DOF simulation developed at Goddard Space Flight Center, is utilized for propagating the attitude of EES during exo-atmospheric flight. 42 allows for a variety of spin eject mechanism scenarios to be simulated for analysis. 10 minutes prior to entry, the 42 states are handed off to the EDL sims. The prime EDL sim utilized by EES is the Program to Optimize Simulated Trajectories II (POST2), a 6-DOF sim developed at Langley Research Center, and the independent verification and validation EDL sim utilized is DSENDS, a 6-DOF sim developed at Jet Propulsion Laboratory. Figure 2 provides a visualization of the flight mechanics simulation model flow through various points in the AEDL phase. Due to the existence of a variety of sim models, the EES flight mechanics team developed processes for data hand-off. These processes included the development of a centralized coordinate frame document, utilization of a single, centralized simulation input document for all sims to reference, and hand-off files containing both the technical data to be ingested by other flight mechanics sims as well as annotations of modeling assumptions utilized to generate the data. Figure~\ref{fig:post2simarchitecture} provides an overview of the POST2 sim architecture wherein POST2 ingests numerous subsystem models and input files. The dispersed state file generated by MONTE provides the position/velocity state of the trajectory while the 42 Handoff file provides the attitude. The aerodynamics database, delivered by the EES aeroscience team, is utilized to simulate the aerodynamic forces and moments experienced during EDL. A custom atmosphere model, developed by EES atmosphere team, is utilized to simulate the anticipated atmosphere environment around the region of Earth through which the EES trajectory flys. These inputs and subsystem models can be varied depending on the AEDL flight mechanics scenario being simulated. Monte Carlo simulations are utilized to generate statistical AEDL performance metrics in the form of scorecards and violin plots. Furthermore, outputs from the POST2 simulation are utilized for follow-on analyses including aerothermal and landing performance. \section{Flight Mechanics Lessons Learned} Though the EES flight mechanics team uncovered a variety of lessons learned through the analysis conducted to support CCRS through preliminary design review, this paper will highlight the most important lessons. A key AEDL performance goal is to ensure the landing footprint of EES remains on the UTTR south range. A common modeling strategy used in EDL analysis is One-Variable-At-a-Time (OVAT). OVAT analysis provides insight into the key drivers that affect AEDL performance metrics. Figure 3 shows the landing ellipses for single dispersion sources as compared to the baseline aggregate of all dispersions. The figure shows that atmosphere winds alone dominate the size of the footprint ellipse (note: EES does not use a parachute unlike previous Earth-return missions and is in wind-driven free fall for ~5min). The significance of the wind led the EES flight mechanics team to pursue the development of a Custom Atmosphere Model [4], in lieu of EarthGRAM [1], built on actual radiosonde wind measurements around the UTTR-region. This decision was driven by the realism in the generated footprint ellipses and lessons-learned from Stardust [5]. These findings will be invaluable for future Earth-return missions in providing an early understanding of the key drivers affecting footprint size and modeling considerations for which to account. Another lesson learned is tied to the AEDL performance goal of achieving passive stability throughout all regions of flight. It is well understood that blunt-body aeroshells are less stable as they transition from supersonic to subsonic. Eliminating a backshell does help improvestability; however, other phenomena such as roll-induced instability during terminal descent can still arise. The EES flight mechanics team developed stability metrics as tools to better understand the causes of and better predict the onset of dynamic instability. These tools were built upon analytical models developed by Jaffe [3] and Murphy [2]. The tools were shown to both be very accurate in correlation with actual unstable cases and useful in developing stability margin policies based on the vehicle design and simulation considerations (e.g. sphere-cone angle change, mass change, wind turbulence). These tools allowed for the current EES design to demonstrate the ability to achieve passive stability and can be an invaluable tool for consideration in the design of parachute-less Earth-return vehicles.

Rohan Deshmukh