Search NASASearch

SEARCH · Search NASA

Results for “Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES

MISIP: a data standard for the reuse and reproducibility of any stable isotope probing-derived nucleic acid sequence and experiment

DNA/RNA-stable isotope probing (SIP) is a powerful tool to link in situ microbial activity to sequencing data. Every SIP dataset captures distinct information about microbial community metabolism, process rates, and population dynamics, offering valuable insights for a wide range of research questions. Data reuse maximizes the information derived from the labor and resource-intensive SIP approaches. Yet, a review of publicly available SIP sequencing metadata showed that critical information necessary for reproducibility and reuse was often missing. Here, we outline the Minimum Information for any Stable Isotope Probing Sequence (MISIP) according to the Minimum Information for any (x) Sequence (MIxS) framework and include examples of MISIP reporting for common SIP experiments. Our objectives are to expand the capacity of MIxS to accommodate SIP-specific metadata and guide SIP users in metadata collection when planning and reporting an experiment. The MISIP standard requires 5 metadata fields—isotope, isotopolog, isotopolog label, labeling approach, and gradient position—and recommends several fields that represent best practices in acquiring and reporting SIP sequencing data (e.g., gradient density and nucleic acid amount). The standard is intended to be used in concert with other MIxS checklists to comprehensively describe the origin of sequence data, such as for marker genes (MISIP-MIMARKS) or metagenomes (MISIP-MIMS), in combination with metadata required by an environmental extension (e.g., soil). The adoption of the proposed data standard will improve the reuse of any sequence derived from a SIP experiment and, by extension, deepen understanding of in situ biogeochemical processes and microbial ecology.

Simpson, Abigayle

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

This dataset accompanies the publication "ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation". This paper introduces a new AI model for protein sequence generation. This dataset contains data related to experiments discussed in the publication. This includes generated sequences and evaluation metrics supporting all unconditional and bias-controlled experiments in the ProtNHF paper.

60 APPLIED LIFE SCIENCES

Rapid wavefield forecasting for earthquake early warning via deep sequence to sequence learning

We propose a deep learning model, WaveCastNet, to forecast high-dimensional wavefields. WaveCastNet integrates a convolutional long expressive memory architecture into a sequence-to-sequence forecasting framework, enabling it to model long-term dependencies and multiscale patterns in both space and time. By sharing weights across spatial and temporal dimensions, WaveCastNet requires significantly fewer parameters than more resource-intensive models such as transformers, resulting in faster inference times. Crucially, WaveCastNet also generalizes better than transformers to rare and critical seismic scenarios, such as high-magnitude earthquakes. Here, we show the ability of the model to predict the intensity and timing of destructive ground motions in real time, using simulated data from the San Francisco Bay Area. Furthermore, we demonstrate its zero-shot capabilities by evaluating WaveCastNet on real earthquake data. Our approach does not require estimating earthquake magnitudes and epicenters, steps that are prone to error in conventional methods, nor does it rely on empirical ground-motion models, which often fail to capture strongly heterogeneous wave propagation effects.

Geophysics

Signal sequences target enzymes and structural proteins to bacterial microcompartments and are critical for microcompartment formation

ABSTRACT Spatial organization of pathway enzymes has emerged as a promising tool to address several challenges in metabolic engineering, such as flux imbalances and off-target product formation. Bacterial microcompartments (MCPs) are a spatial organization strategy used natively by many bacteria to encapsulate metabolic pathways that produce toxic, volatile intermediates. Several recent studies have focused on engineering MCPs to encapsulate heterologous pathways of interest, but how this engineering affects MCP assembly and function is poorly understood. In this study, we investigated the role of signal sequences, short domains that target proteins to the MCP core, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterized two novel Pdu signal sequences on the structural proteins PduM and PduB, which constitute the first report of metabolosome signal sequences on structural proteins rather than enzymes. We then explored the role of enzymatic and structural Pdu signal sequences on MCP assembly by deleting their encoding sequences from the genome alone and in combination. Deleting enzymatic signal sequences decreased the MCP formation, but this defect could be recovered in some cases by overexpressing genes encoding the knocked-out signal sequence fused to a heterologous protein. By contrast, deleting structural signal sequences caused similar defects to knocking out the genes encoding the full-length PduM and PduB proteins. Our results contribute to a growing understanding of how MCPs form and function in bacteria and provide strategies to mitigate assembly disruption when encapsulating heterologous pathways in MCPs. IMPORTANCE Spatially organizing biosynthetic pathway enzymes is a promising strategy to increase pathway throughput and yield. Bacterial microcompartments (MCPs) are proteinaceous organelles that many bacteria natively use as a spatial organization strategy to encapsulate niche metabolic pathways, providing significant metabolic benefits. Encapsulating heterologous pathways of interest in MCPs could confer these benefits to industrially relevant pathways. Here, we investigate the role of signal sequences, short domains that target proteins for encapsulation in MCPs, in the assembly of 1,2-propanediol utilization (Pdu) MCPs. We characterize two novel signal sequences on structural proteins, constituting the first Pdu signal sequences found on structural proteins rather than enzymes, and perform knockout studies to compare the impacts of enzymatic and structural signal sequences on MCP assembly. Our results demonstrate that enzymatic and structural signal sequences play critical but distinct roles in Pdu MCP assembly and provide design rules for engineering MCPs while minimizing disruption to MCP assembly.

Johnson, Elizabeth R. (ORCID:0000000179236881)

Circularity in Sequence-Controlled Copolyamides Enabled by Regioselective Enzymatic Hydrolysis

Sequence-controlled polymers enable precise control over macromolecular structures and function, but both their synthesis and end-of-life management remain fundamental challenges. Achieving high sequence fidelity is synthetically demanding, and conventional depolymerization methods lack regioselectivity, leading to irreversible loss of encoded molecular information and limiting polymer circularity. Enzymatic catalysis offers a potential solution by combining substrate specificity with selective bond cleavage. Here, we report the synthesis, characterization, and regioselective enzymatic depolymerization of poly- (X,AMA), a sequence-controlled copolyamide composed of alternating hexamethylenediamine−adipic acid (MA) and pxylylenediamine− adipic acid (XA) repeat units. Poly(X,AMA) was synthesized via solid-state polycondensation (SSP) of sequence-defined oligomers, enabling precise control over repeat-unit order. Polymer microstructure and sequence fidelity were confirmed by 13 C NMR spectroscopy and MALDI−TOF mass spectrometry. Comparison with a statistical copolymer analogue and Nylon-66 demonstrated pronounced differences in crystallinity, morphology, and thermal behavior arising from sequence control. Screening of 96 Nylon hydrolase homologues against poly(X,AMA) revealed strongly enzyme-dependent depolymerization profiles. While tetrad formation was generally favored, enzymes displayed pronounced sequence selectivity, preferentially releasing distinct sequence-defined tetrads XAMA or MAXA. SSP of sequence-defined tetrad MAXA produced a copolyamide with near identical monomer ordering as poly(X,AMA). Computational modeling of enzyme−substrate complexes identified structural features consistent with the observed regioselectivity. Together, these results establish selective enzymatic depolymerization as a viable strategy for the circular recycling of sequence-controlled polymers and provide a foundation for the rational engineering of enzymes for programmable polymer deconstruction.

Amides

Development of near-optimal advanced control sequences for chiller plants with water-side economizers in U.S. Climates (ASHRAE RP-1661)

Various advanced control sequences for chiller plants with water-side economizers (WSE) have been proposed in literature, but the evaluation and optimization of those controls is limited. It is possible to maximize energy savings by selecting different sequences and related parameters based on the plant configuration, load, and climate. This paper addresses this gap by developing near-optimal advanced control sequences for chiller plants with WSEs. First, advanced control sequences for chiller plants with WSEs are categorized into condenser water, chilled water, and hybrid controls and representative sequences from each category are identified. Next, 504 different scenarios are optimized. These scenarios represent all possible combinations of two plant configurations, a constant or variable load profile, three advanced control sequences, and seven optimization parameter combinations in six climate zones. The results show the recommended near-optimal sequences can reduce energy consumption by up to 15% relative to the baseline depending on the configuration, load profile, and climate. Specifically, the CW-CHW sequence is recommended for the majority of systems because it is often the most energy efficient and/or reduces the runtime of chillers. The methodology in this paper provides practical guidance for achieving energy savings through near-optimal control of chiller plants with WSEs.

42 ENGINEERING

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Sequence-defined structural transitions by calcium-responsive proteins

Biopolymer sequences dictate their functions, and protein-based polymers are a promising platform to establish sequence–function relationships for novel biopolymers. To efficiently explore vast sequence spaces of natural proteins, sequence repetition is a common strategy to tune and amplify specific functions. This strategy is applied to repeats-in-toxin (RTX) proteins with calcium-responsive folding behavior, which stems from tandem repeats of the nonapeptide GGXGXDXUX in which X can be any amino acid and U is a hydrophobic amino acid. To determine the functional range of this nonapeptide, we modified a naturally occurring RTX protein that forms β-roll structures in the presence of calcium. Sequence modifications focused on calcium-binding turns within the repetitive region, including either global substitution of nonconserved residues or complete replacement with tandem repeats of a consensus nonapeptide GGAGXDTLY. Some sequence modifications disrupted the typical transition from intrinsically disordered random coils to folded β rolls, despite conservation of the underlying nonapeptide sequence. Proteins enriched with smaller, hydrophobic amino acids adopted secondary structures in the absence of calcium and underwent structural rearrangements in calcium-rich environments. In contrast, proteins with bulkier, hydrophilic amino acids maintained intrinsic disorder in the absence of calcium. In conclusion, these results indicate a significant role of nonconserved amino acids in calcium-responsive folding, thereby revealing a strategy to leverage sequences in the design of tunable, calcium-responsive biopolymers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

Sequence length scaling in vision transformers for scientific images on frontier

Vision Transformers (ViTs) are pivotal for foundational models in scientific imagery, including Earth science applications, due to their capability to process large sequence lengths. While transformers for text have inspired scaling sequence lengths in ViTs, adapting these for ViTs introduces unique challenges. We develop distributed sequence parallelism for ViTs, enabling them to handle up to 1M tokens. Our approach, leveraging DeepSpeed-Ulysses and Long-Sequence-Segmentation with model sharding, is the first to apply sequence parallelism in ViT training, achieving a 94% batch scaling efficiency on 2,048 AMD-MI250X GPUs. Evaluating sequence parallelism in ViTs, particularly in models up to 10B parameters, highlighted substantial bottlenecks. We countered these with hybrid sequence, pipeline, and flash attention strategies, to scale beyond single GPU memory limits. Our method significantly enhances climate modeling accuracy by 20% in temperature predictions, marking the first training of a vision transformer model to convergence with a sequence length of 188K tokens, using full self-attention.

Tsaris, Aristeidis (aris) [ORNL] (ORCID:0000000277

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis

Structure–Function Relationships in Sequence-Controlled Copolymers for Rare Earth Element Chelation

The ability to tune material function through primary sequence is a defining feature of biological macromolecules, allowing precise control over structure and target interactions in complex aqueous environments. However, translating sequence–structure–function relationships to synthetic macromolecules is challenging due to their dispersity in sequence, conformation, and composition. Here, we report systematic studies of amphiphilic polymer chelators designed to probe how composition and patterning influence binding affinity and selectivity for rare earth elements (REEs), a series of technologically relevant metals with challenging separation profiles. A library of copolymers varying hydrophobic monomer composition and patterning was synthesized via reversible addition–fragmentation chain transfer (RAFT) polymerization, spanning statistical, gradient, and block architectures. REE binding was quantified using a high-throughput colorimetric assay, and reconstruction of polymer ensembles using kinetic stochastic simulations enabled quantitative comparisons of sequence heterogeneity, linking local monomer colocalization to emergent REE binding. Further, we investigated the role of different hydrophobic comonomers in tuning metal coordination, with binding trends linked to structural features that influence binding site desolvation. Complementary dynamic light scattering (DLS) and small-angle X-ray scattering (SAXS) measurements showed that both polymer and monomer architecture modulate metal-induced conformational changes, and that multichain assembly behavior emerges beyond critical hydrophobic thresholds. Sequence control also altered REE selectivity, with nonmonotonic differences observed across compositionally identical polymers with different sequence architectures. Together, these findings establish design principles that connect polymer sequence and structure to binding performance, guiding the design of macromolecular chelators with enhanced affinity and selectivity for applications in separations, sensing, and catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Local Chain Dynamics in Sequence-Controlled Polymers as a Tunable Handle for Rare Earth Sequestration

Chain dynamics govern the intricate behaviors of proteins, underpinning functions such as catalysis, recognition, and stimulus response, and are an increasingly appreciated aspect of structure–function relationships. Analogously, manipulating chain dynamics and structure in abiotic polymers via sequence control is an exciting, yet underexplored, strategy for improving material functions. In this work, we report a systematic study relating the sequence of polymeric sequestrants to their structure and dynamics, as well as to their binding affinity and selectivity for model substrates, rare earth elements (REEs). A series of sequence-controlled polymers with metal chelating, solubilizing, and structure forming monomers was synthesized via multiblock polymerization, yielding compositionally identical polymers with spectroscopically resolved domains and distinct morphologies. Using a combination of small-angle X-ray scattering and 19 F NMR relaxometry measurements, we connected differences in polymer structure and dynamics to polymer sequence variables such as the patchiness (density) of the structure forming monomer and the location of the chelating monomer. Furthermore, we found that, relative to calcium, all polymers in the series collapse more and have slower dynamics when binding REEs (lanthanum and lutetium) , though the extent of these effects were sequence-dependent and localized to specific domains within the polymer. Notably, sequence-controlled polymers that exhibited the largest conformational and dynamic changes upon binding REEs also bound REEs with the greatest affinity and modest selectivity. Collectively, these results correlate monomer patterning with dynamics, morphology, and REE binding performance en route to the development of efficient and selective macromolecular chelators.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH