Search NASASearch

SEARCH · Search NASA

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

TraceContract: A Scala DSL for Trace Analysis

In this paper we describe TRACECONTRACT, an API for trace analysis, implemented in the SCALA programming language. We argue that for certain forms of trace analysis the best weapon is a high level programming language augmented with constructs for temporal reasoning. A trace is a sequence of events, which may for example be generated by a running program, instrumented appropriately to generate events. The API supports writing properties in a notation that combines an advanced form of data parameterized state machines with temporal logic. The implementation utilizes SCALA's support for defining internal Domain Specific Languages (DSLs). Furthermore SCALA's combination of object oriented and functional programming features, including partial functions and pattern matching, makes it an ideal host language for such an API.

log file analysis

Time-series metagenomics reveals changing protistan ecology of a temperate dimictic lake

Abstract Background Protists, single-celled eukaryotic organisms, are critical to food web ecology, contributing to primary productivity and connecting small bacteria and archaea to higher trophic levels. Lake Mendota is a large, eutrophic natural lake that is a Long-Term Ecological Research site and among the world’s best-studied freshwater systems. Metagenomic samples have been collected and shotgun sequenced from Lake Mendota for the last 20 years. Here, we analyze this comprehensive time series to infer changes to the structure and function of the protistan community and to hypothesize about their interactions with bacteria. Results Based on small subunit rRNA genes extracted from the metagenomes and metagenome-assembled genomes of microeukaryotes, we identify shifts in the eukaryotic phytoplankton community over time, which we predict to be a consequence of reduced zooplankton grazing pressures after the invasion of a invasive predator (the spiny water flea) to the lake. The metagenomic data also reveal the presence of the spiny water flea and the zebra mussel, a second invasive species to Lake Mendota, prior to their visual identification during routine monitoring. Furthermore, we use species co-occurrence and co-abundance analysis to connect the protistan community with bacterial taxa. Correlation analysis suggests that protists and bacteria may interact or respond similarly to environmental conditions. Cryptophytes declined in the second decade of the timeseries, while many alveolate groups (e.g., ciliates and dinoflagellates) and diatoms increased in abundance, changes that have implications for food web efficiency in Lake Mendota. Conclusions We demonstrate that metagenomic sequence-based community analysis can complement existing efforts to monitor protists in Lake Mendota based on microscopy-based count surveys. We observed patterns of seasonal abundance in microeukaryotes in Lake Mendota that corroborated expectations from other systems, including high abundance of cryptophytes in winter and diatoms in fall and spring, but with much higher resolution than previous surveys. Our study identified long-term changes in the abundance of eukaryotic microbes and provided context for the known establishment of an invasive species that catalyzes a trophic cascade involving protists. Our findings are important for decoding potential long-term consequences of human interventions, including invasive species introduction.

59 BASIC BIOLOGICAL SCIENCES

Designing a training tool for imaging mental models

The training process can be conceptualized as the student acquiring an evolutionary sequence of classification-problem solving mental models. For example a physician learns (1) classification systems for patient symptoms, diagnostic procedures, diseases, and therapeutic interventions and (2) interrelationships among these classifications (e.g., how to use diagnostic procedures to collect data about a patient's symptoms in order to identify the disease so that therapeutic measures can be taken. This project developed functional specifications for a computer-based tool, Mental Link, that allows the evaluative imaging of such mental models. The fundamental design approach underlying this representational medium is traversal of virtual cognition space. Typically intangible cognitive entities and links among them are visible as a three-dimensional web that represents a knowledge structure. The tool has a high degree of flexibility and customizability to allow extension to other types of uses, such a front-end to an intelligent tutoring system, knowledge base, hypermedia system, or semantic network.

Dede, Christopher J.

Nocturnal Observations of the Semidiurnal Tide at a Midlatitude Site

Fabry-Perot interferometer observations of the mesospheric hydroxyl emission and the lower thermospheric OI (5577A) emission have been conducted from an airglow observatory at a dark field site in southeastern Michigan for the past several years. The primary functions of the observatory are to provide a database for correlative observations with the UARS satellite and to provide a synoptic measurement program for the coupling energetics and dynamics of atmospheric regions effort, An intensive operational effort between May 1993 and July 1994 has resulted in a substantial data set from which neutral winds have been determined from the bifilter acquisition sequence. A 'best fit' analysis in the least squares sense of the simultaneous measurements of the neutral winds to a 12-hour periodicity has provided amplitude and phase parameters for the semidiurnal tide as well as a measure of the mean wind. The measured tidal amplitude is greater at the higher altitude, though the seasonal behavior at both altitudes is similar with greater amplitudes during August/September and April/May. Both meridional and zonal wind components are consistent with a semidiurnal tidal description during the entire observational sequence except for the May to July 1993 period. The mean winds show annual variation in the meridional flow, being equatorward from May to October and poleward during the winter. The zonal flow is primarily eastward during the entire observational window with higher speed flows during May/June at the higher attitude and June/July at the lower altitude. A comparison with a semidiurnal tidal model indicates that the measured tidal amplitudes are a factor of 2 times greater, while the phases show similar equinoctial transitions.

Niciejewski, R. J.

Voyager flight engineering - Preparing for Uranus

Two Voyager spacecraft are currently engaged in exploration of the outer solar system with Voyager 2 scheduled to conduct the first close-up investigation of the planet Uranus during the period November 4, 1985 through March 3, 1986. Flight engineering for the Voyager project has the objectives of delivering a functioning spacecraft containing observing sequences to the right places at the right times. Due to the changing environment as the mission has progressed outward from Jupiter to Saturn to Uranus (and on to Neptune), this engineering task has included the development of significant new capabilities. The paper utilizes the case-study method to examine some new spacecraft capabilities in three subsystems: data, attitude and articulation control, and power. The implementation of a new navigational data-type, delta DOR, is also reviewed. An overview is given of the Voyager sequencing process for the cruise and encounter phases with a case study focusing on late updating of part of the near encounter sequence. The prospective mission to Neptune is previewed.

Mclaughlin, W. I.

Numerical Investigation and Optimization of a Flushwall Injector for Scramjet Applications at Hypervelocity Flow Conditions

An investigation utilizing Reynolds-averaged simulations (RAS) was performed in order to demonstrate the use of design and analysis of computer experiments (DACE) methods in Sandia’s DAKOTA software package for surrogate modeling and optimization. These methods were applied to a flow- path fueled with an interdigitated flushwall injector suitable for scramjet applications at hyper- velocity conditions and ascending along a constant dynamic pressure flight trajectory. The flight Mach number, duct height, spanwise width, and injection angle were the design variables selected to maximize two objective functions: the thrust potential and combustion efficiency. Because the RAS of this case are computationally expensive, surrogate models are used for optimization. To build a surrogate model a RAS database is created. The sequence of the design variables comprising the database were generated using a Latin hypercube sampling (LHS) method. A methodology was also developed to automatically build geometries and generate structured grids for each design point. The ensuing RAS analysis generated the simulation database from which the two objective functions were computed using a one-dimensionalization (1D) of the three-dimensional simulation data. The data were fitted using four surrogate models: an artificial neural network (ANN), a cubic polynomial, a quadratic polynomial, and a Kriging model. Variance-based decomposition showed that both objective functions were primarily driven by changes in the duct height. Multiobjective design optimization was performed for all four surrogate models via a genetic algorithm method. Optimal solutions were obtained at the upper and lower bounds of the flight Mach number range. The Kriging model predicted an optimal solution set that exhibited high values for both objective functions. Additionally, three challenge points were selected to assess the designs on the Pareto fronts. Further sampling among the designs of the Pareto fronts may be required to lower the surrogate model errors and perform more accurate surrogate-model-based optimization.

Shenoy, Rajiv R.

Nulling at the Keck Interferometer

The nulling mode of the Keck Interferometer is being commissioned at the Mauna Kea summit. The nuller combines the two Keck telescope apertures in a split-pupil mode to both cancel the on-axis starlight and to coherently detect the residual signal. The nuller, working at 10 um, is tightly integrated with the other interferometer subsystems including the fringe and angle trackers, the delay lines and laser metrology, and the real-time control system. Since first 10 um light in August 2004, the system integration is proceeding with increasing functionality and performance, leading to demonstration of a 100:1 on-sky null in 2005. That level of performance has now been extended to observations with longer coherent integration times. An overview of the overall system is presented, with emphasis on the observing sequence, phasing system, and differences with respect to the V2 system, along with a presentation of some recent engineering data.

interferometry

Irreducible Tests for Space Mission Sequencing Software

As missions extend further into space, the modeling and simulation of their every action and instruction becomes critical. The greater the distance between Earth and the spacecraft, the smaller the window for communication becomes. Therefore, through modeling and simulating the planned operations, the most efficient sequence of commands can be sent to the spacecraft. The Space Mission Sequencing Software is being developed as the next generation of sequencing software to ensure the most efficient communication to interplanetary and deep space mission spacecraft. Aside from efficiency, the software also checks to make sure that communication during a specified time is even possible, meaning that there is not a planet or moon preventing reception of a signal from Earth or that two opposing commands are being given simultaneously. In this way, the software not only models the proposed instructions to the spacecraft, but also validates the commands as well.To ensure that all spacecraft communications are sequenced properly, a timeline is used to structure the data. The created timelines are immutable and once data is as-signed to a timeline, it shall never be deleted nor renamed. This is to prevent the need for storing and filing the timelines for use by other programs. Several types of timelines can be created to accommodate different types of communications (activities, measurements, commands, states, events). Each of these timeline types requires specific parameters and all have options for additional parameters if needed. With so many combinations of parameters available, the robustness and stability of the software is a necessity. Therefore a baseline must be established to ensure the full functionality of the software and it is here where the irreducible tests come into use.

sequencing test

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

Multimodal Approaches for Leveraging Domain Knowledge with State-of-the-Art Machine Learning to Engineer Biocatalysts

This grant aimed to accelerate the development of specialized enzymes—biological catalysts essential for sustainable manufacturing and medicine—by integrating traditional laboratory evolution with cutting-edge artificial intelligence. To achieve this, we developed a suite of high-throughput sequencing tools and a centralized database to bridge the gap between a protein’s genetic "code" and its physical function. By training machine learning models on large datasets, we also demonstrated the ability to move beyond slow, trial-and-error testing to a "generative" approach, where AI can independently design new, versatile enzymes like tryptophan synthases. Ultimately, these findings demonstrate that combining laboratory data with computer-guided design enables the engineering of highly efficient biological tools with unprecedented speed and precision.

59 BASIC BIOLOGICAL SCIENCES

Design of Neural Networks for Fast Convergence and Accuracy: Dynamics and Control

A procedure for the design and training of artificial neural networks, used for rapid and efficient controls and dynamics design and analysis for flexible space systems, has been developed. Artificial neural networks are employed, such that once properly trained, they provide a means of evaluating the impact of design changes rapidly. Specifically, two-layer feedforward neural networks are designed to approximate the functional relationship between the component/spacecraft design changes and measures of its performance or nonlinear dynamics of the system/components. A training algorithm, based on statistical sampling theory, is presented, which guarantees that the trained networks provide a designer-specified degree of accuracy in mapping the functional relationship. Within each iteration of this statistical-based algorithm, a sequential design algorithm is used for the design and training of the feedforward network to provide rapid convergence to the network goals. Here, at each sequence a new network is trained to minimize the error of previous network. The proposed method should work for applications wherein an arbitrary large source of training data can be generated. Two numerical examples are performed on a spacecraft application in order to demonstrate the feasibility of the proposed approach.

Maghami, Peiman G.

Analysis of the myosins encoded in the recently completed Arabidopsis thaliana genome sequence

BACKGROUND: Three types of molecular motors play an important role in the organization, dynamics and transport processes associated with the cytoskeleton. The myosin family of molecular motors move cargo on actin filaments, whereas kinesin and dynein motors move cargo along microtubules. These motors have been highly characterized in non-plant systems and information is becoming available about plant motors. The actin cytoskeleton in plants has been shown to be involved in processes such as transportation, signaling, cell division, cytoplasmic streaming and morphogenesis. The role of myosin in these processes has been established in a few cases but many questions remain to be answered about the number, types and roles of myosins in plants. RESULTS: Using the motor domain of an Arabidopsis myosin we identified 17 myosin sequences in the Arabidopsis genome. Phylogenetic analysis of the Arabidopsis myosins with non-plant and plant myosins revealed that all the Arabidopsis myosins and other plant myosins fall into two groups - class VIII and class XI. These groups contain exclusively plant or algal myosins with no animal or fungal myosins. Exon/intron data suggest that the myosins are highly conserved and that some may be a result of gene duplication. CONCLUSIONS: Plant myosins are unlike myosins from any other organisms except algae. As a percentage of the total gene number, the number of myosins is small overall in Arabidopsis compared with the other sequenced eukaryotic genomes. There are, however, a large number of class XI myosins. The function of each myosin has yet to be determined.

NASA Discipline Plant Biology

Mesh generation/refinement using fractal concepts and iterated function systems

A novel method of mesh generation is proposed which is based on the use of fractal concepts to derive contractive, affine transformations. The transformations are constructed in such a manner that the attractors of the resulting maps are a union of the points, lines and surfaces in the domain. In particular, the mesh nodes may be generated recursively as a sequence of points which are obtained by applying the transformations to a coarse background mesh constructed from the given boundary data. A Delaunay triangulation or similar edge connection approach can then be performed on the resulting set of nodes in order to generate the mesh. Local refinement of an existing mesh can also be performed using the procedure. The method is easily extended to three dimensions, in which case the Delaunay triangulation is replaced by an analogous 3D tesselation.

Bova, S. W.

(abstract) Pluto Integrated Camera-Spectrometer (PICS): a Low Mass, Low Power Instrument for Planetary Exploration

The concept we describe is an integrated instrument (a Pluto Integrated Camera-Spectrometer -- PICS) that will perform the functions of all three optical instruments required by the Pluto Fast Flyby Mission: the near-IR spectrometer, the camera, and the UV spectrometer. This integrated approach minimizes mass and power use. It also forced us early in the conceptual design to consider integrated observational sequences and integrated power management, thus ensuring compatible duty cycles (i.e., exposure times, readout rates) to meet the composite requirements for data collection, compression, and storage. Based on flight mission experience we believe that this integrated approach will result in substantial cost savings, both in reworking instrument designs during accommodation, as well as in sequence planning and integration. Finally, this integrated payload automatically yields a cohesive mission data set, optimized for correlative analysis. The presentation will provide details of the PICS instrument design and describe the fabrication and testing of the integrated SiC structure and optics at SSG Inc. Final integration and test plans for the prototype will also be described.

Pluto

Learning genetic perturbation effects with variational causal inference

Advances in sequencing technologies have enhanced the understanding of gene regulation in cells. In particular, Perturb-seq has enabled high-resolution profiling of the transcriptomic response to genetic perturbations at the single-cell level. This understanding has implications in functional genomics and potentially for identifying therapeutic targets. Various computational models have been developed to predict perturbational effects. While deep learning models excel at interpolating observed perturbational data, they tend to overfit in the lack of enough data and may not generalize well to unseen perturbations. In contrast, mechanistic models, such as linear causal models based on gene regulatory networks, hold greater potential for extrapolation, as they encapsulate regulatory information that can predict responses to unseen perturbations. However, their application has been limited to small studies due to overly simplistic assumptions, making them less effective in handling noisy, large-scale single-cell data. We propose a hybrid approach that combines a mechanistic causal model with variational deep learning, termed Single Cell Causal Variational Autoencoder (SCCVAE). The mechanistic model employs a learned regulatory network to represent perturbational changes as shift interventions that propagate through the learned network. SCCVAE integrates this mechanistic causal model into a variational autoencoder, generating rich, comprehensive transcriptomic responses. Our results indicate that SCCVAE exhibits superior performance over current state-of-the-art baselines for extrapolating to predict unseen perturbational responses. Additionally, for the observed perturbations, the latent space learned by SCCVAE allows for the identification of functional perturbation modules and simulation of single-gene knockdown experiments of varying penetrance, presenting a robust tool for interpreting and interpolating perturbational responses at the single-cell level.

59 BASIC BIOLOGICAL SCIENCES

Controlled cellular energy conversion in brown adipose tissue thermogenesis

Brown adipose tissue serves as a model system for nonshivering thermogenesis (NST) since a) it has as a primary physiological function the conversion of chemical energy to heat; and b) preliminary data from other tissues involved in NST (e.g., muscle) indicate that parallel mechanisms may be involved. Now that biochemical pathways have been proposed for brown fat thermogenesis, cellular models consistent with a thermodynamic representation can be formulated. Stated concisely, the thermogenic mechanism in a brown fat cell can be considered as an energy converter involving a sequence of cellular events controlled by signals over the autonomic nervous system. A thermodynamic description for NST is developed in terms of a nonisothermal system under steady-state conditions using network thermodynamics. Pathways simulated include mitochondrial ATP synthesis, a Na+/K+ membrane pump, and ionic diffusion through the adipocyte membrane.

Horowitz, J. M.

A Galerkin method for the estimation of parameters in hybrid systems governing the vibration of flexible beams with tip bodies

An approximation scheme is developed for the identification of hybrid systems describing the transverse vibrations of flexible beams with attached tip bodies. In particular, problems involving the estimation of functional parameters are considered. The identification problem is formulated as a least squares fit to data subject to the coupled system of partial and ordinary differential equations describing the transverse displacement of the beam and the motion of the tip bodies respectively. A cubic spline-based Galerkin method applied to the state equations in weak form and the discretization of the admissible parameter space yield a sequence of approximating finite dimensional identification problems. It is shown that each of the approximating problems admits a solution and that from the resulting sequence of optimal solutions a convergent subsequence can be extracted, the limit of which is a solution to the original identification problem. The approximating identification problems can be solved using standard techniques and readily available software.

Banks, H. T.