Search NASASearch

SEARCH · Search NASA

Results for “sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES

Local Chain Dynamics in Sequence-Controlled Polymers as a Tunable Handle for Rare Earth Sequestration

Chain dynamics govern the intricate behaviors of proteins, underpinning functions such as catalysis, recognition, and stimulus response, and are an increasingly appreciated aspect of structure–function relationships. Analogously, manipulating chain dynamics and structure in abiotic polymers via sequence control is an exciting, yet underexplored, strategy for improving material functions. In this work, we report a systematic study relating the sequence of polymeric sequestrants to their structure and dynamics, as well as to their binding affinity and selectivity for model substrates, rare earth elements (REEs). A series of sequence-controlled polymers with metal chelating, solubilizing, and structure forming monomers was synthesized via multiblock polymerization, yielding compositionally identical polymers with spectroscopically resolved domains and distinct morphologies. Using a combination of small-angle X-ray scattering and 19 F NMR relaxometry measurements, we connected differences in polymer structure and dynamics to polymer sequence variables such as the patchiness (density) of the structure forming monomer and the location of the chelating monomer. Furthermore, we found that, relative to calcium, all polymers in the series collapse more and have slower dynamics when binding REEs (lanthanum and lutetium) , though the extent of these effects were sequence-dependent and localized to specific domains within the polymer. Notably, sequence-controlled polymers that exhibited the largest conformational and dynamic changes upon binding REEs also bound REEs with the greatest affinity and modest selectivity. Collectively, these results correlate monomer patterning with dynamics, morphology, and REE binding performance en route to the development of efficient and selective macromolecular chelators.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

wavess 1.2: presenting an HLA-aware within-host virus sequence simulation framework

Motivation Understanding how virus sequences are shaped by selection can inform vaccine design and transmission inference. Modeling within-host evolution to interrogate these questions requires a detailed mechanistic framework that accurately captures sequence diversification. The CD8 + cytotoxic T-lymphocyte (CTL) response plays an important role in immune-mediated selection and can leave strong signatures in virus sequences; however, existing sequence-based within-host virus modeling frameworks do not explicitly include a human leukocyte antigen (HLA)-aware CTL response. Results We extended our previously published within-host sequence evolution simulator, wavess, to include an explicit CTL response, and share a method for identifying HLA-specific CTL epitopes given a founder virus sequence. We also updated the model to permit a variable recombination rate, which allows for modeling non-adjacent genes, segmented genomes, and recombination hotspots. These extensions to wavess allow for more accurate simulation of viruses and virus genes, particularly in regions of the genome where the immune response is dominated by CTLs (rather than antibodies). It also provides the foundation for investigations of how these newly-added biological mechanisms influence within-host evolution. Availability and implementation The core of wavess is written in Python 3, with helper functions written in R. It is available at https://github.com/MolEvolEpid/wavess.

60 APPLIED LIFE SCIENCES

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology

Beyond sequence similarity: toward function-based screening of nucleic acid synthesis

Synthetic nucleic acids are a key input to modern biotechnology, yet they represent dual-use materials that require robust screening to mitigate biosecurity risks. The prevailing screening paradigm, which identifies sequences of concern (SoCs) through sequence similarity to controlled pathogens and toxins, may not fully capture risks posed by AI tools that can decouple biomolecular function from reliance on known sequences. Rapidly advancing biodesign capabilities enable the generation of genes and proteins that might evade sequence-based detection. We highlight the critical need for function-based screening approaches that can detect sequences capable of hazardous biological functions, regardless of similarity to known SoCs. We examine the feasibility of function-based screening with an initial focus on proteins, arguing that, while protein sequence space is vast, biologically functional proteins are significantly constrained by biophysical and biochemical requirements that can be learned and modeled. We propose a concrete implementation framework organized along a continuum of complexity, starting with toxins as the most tractable targets before expanding to more complex pathogenic functions. We then discuss open challenges and describe a research and development strategy to address them.

59 BASIC BIOLOGICAL SCIENCES

Diversity of Sordariales Fungi: Identification of Seven New Species of Naviculisporaceae Through Morphological Analyses and Genome Sequencing

Thanks to next-generation sequencing (NGS) technologies, the diversity of fungi can now be investigated through the analysis of their genome sequences. Naviculisporaceae is a family within the Sordariales, whose diversity is not well-known, with only one genome sequence published for this family. Here, we report on the isolation and cultivation of 20 new strains of Naviculisporaceae. Their genome sequences, as well as those of the five commercially available strains, were determined, thus providing complete genome sequences for 25 new Naviculisporaceae strains. Species delimitation was conducted using a combination of (1) ITS + LSU phylogenetic analysis of the new isolates along with other known species of the family, (2) comparisons between DNA barcode sequences of the new strains with those of the known species, and (3) average genome-wide nucleotide identity calculation. We built a phylogenomic tree and studied the organization of the mating-type locus. In vitro fruiting was obtained for 16 strains, enabling the definition of seven new species, namely Pseudorhypophila gallica, Pseudorhypophila guyanensis Rhypophila alpibus, Rhypophila brasiliensis, Rhypophila camarguensis, Rhypophila reunionensis and Rhypophila thailandica, as well as two new combinations, namely Pseudorhypophila latipes and Pseudorhypophila oryzae. Eight strains for which in vitro fruiting was not obtained may belong to additional new species. These results expand the known diversity of the Naviculisporaceae and greatly enlarge the genomic data available for the family.

Naviculisporaceae

Comparison of the spatial statistics of random and defined-sequence photoresist films

The resolution-line edge roughness-sensitivity tradeoff has motivated the exploration of potential improvements using defined sequence polymers and polymer-bound photoacid generators and quenchers. We characterize the internal structures of positive tone photoresist polymer films formed from defined sequence polymers and compare them with random copolymers of the same composition. We model their imaging to connect initially to developable film structures. We use a polymer packing algorithm to simulate films of diverse compositions and locations of photoacid generators and quenchers, using the composition of an ESCAP photoresist. We use a simple extreme ultraviolet exposure-deprotection algorithm to model developable image formation within them. In all cases, the spatial distribution of chemical moieties in the film for defined sequence polymers is nearly indistinguishable from random copolymers. We evaluate several exposure-deprotection scenarios and find that a defined sequence copolymer has a distinctive developable image under certain circumstances. The use of defined sequence polymers within a photoresist layer does not automatically result in improved imaging; however, they do have some characteristics different from random polymers of the same composition. Further study of these characteristics may provide a route to improved control over the nanoscale imaging process.

36 MATERIALS SCIENCE

Challenges in Pulsed-Field Gradient Nuclear Magnetic Resonance on Magnetically Heterogeneous Interfaces: Sequence and Field-Dependent Apparent Diffusion Coefficients

It is well known that the internal gradient (gi) that exists within pores haunts the diffusion coefficient (D) as measured by the pulsed-field gradient (PFG) nuclear magnetic resonance (NMR). Several PFG-NMR methods developed to determine an accurate D were not successful. Then, the steady-state diffusion coefficient (Dapp,8) for the cation [C4mim]+ of [C4mim][Tf2N]; [1-butyl-3-methylimidazolium][bis(trifluoromethylsulfonyl)imde] ionic liquid confined in ordered mesoporous carbon (OMC) were determined by comparing Dapp,8 obtained from 1H PFG-NMR performed with three different stimulated echo sequences: STE, APFG, and MPFG under the two external magnetic field strength, B0 = 9.4 and 14.1 Tesla. The measured Dapp,8 which is an order of magnitude smaller than D of bulk [C4mim][Tf2N], is in good agreement between APFG and MPFG both in B0 = 9.4 and 14.1 Tesla. However, the strong gi artifact, which caused apparent diffusion coefficient (Dapp) depending strongly and weakly on B0 and temperature, respectively, in diffusion-time dependent Dapp, Dapp(?) obtained from a sequence with monopolar gradients (STE) was suppressed by using sequences employing bipolar gradients (APFG and MPFG) in the region of steady-state diffusion. But incompletely suppressed gi artifact resulting in the different behaviors of the early part of Dapp(?) between the sequences leads a ˜ 0.6 and 0.9 in MPFG and APFG, respectively, in the relationship between mean squared displacement and diffusion time: = 2Dta, where a = 0.5 and 1 for 1-dimensional single file diffusion and 3-dimensional bulk diffusion, respectively. The above observations clearly show that the diffusion behavior of ions/molecules within the pores and pore structure, such as the surface-to-volume ratio? (D?_app (?)=D_0 [1-4/(9vp) S/V v(D_0 ?)]) and tortuosity (T = D0/Dapp,8), are possible to be misunderstood, especially in the systems with a non-negligible gi. This work demonstrates that it may be necessary to test several PFG sequences under multiple external magnetic fields for the correct determination of the diffusion behavior of ions/molecules in the pores with a larger internal gradient, gi.

Han, Kee Sung

Sequence-Structure–Property Relationships in Short-Chain Polyesters: How Primary Structure Governs Macroscopic Performance

While polymer properties are fundamentally linked to their nanostructure, the influence of monomer sequence remains less understood than stereochemical factors like tacticity. This study examines how sequence distribution affects the thermal behavior and morphology of homo- and copolyesters, specifically comparing polymers derived from constitutionally identical monomers but with varying degrees of sequence regularity depending on monomer structure or polymerization selectivity. Our findings show that increasing sequence defects progressively diminish thermal stability, crystallinity, melting temperatures, and morphological order. As new materials become more compositionally complex, this work underscores the importance of sequence control in the design of advanced polymers for emerging applications.

Bocharova, Vera [Oak Ridge National Laboratory (OR

Amino Acid Sequence Controls Enhanced Electron Transport in Heme-Binding Peptide Monolayers

Metal-binding proteins have the exceptional ability to facilitate long-range electron transport in nature. Despite recent progress, the sequence-structure–function relationships governing electron transport in heme-binding peptides and protein assemblies are not yet fully understood. In this work, the electronic properties of a series of heme-binding peptides inspired by cytochrome bc1 are studied using a combination of molecular electronics experiments, molecular modeling, and simulation. Self-assembled monolayers (SAMs) are prepared using sequence-defined heme-binding peptides capable of forming helical secondary structures. Following monolayer formation, the structural properties and chemical composition of assembled peptides are determined using atomic force microscopy and X-ray photoelectron spectroscopy, and the electronic properties (current density–voltage response) are characterized using a soft contact liquid metal electrode method based on eutectic gallium–indium alloys (EGaIn). Our results show a substantial 1000-fold increase in current density across SAM junctions upon addition of heme compared to identical peptide sequences in the absence of heme, while maintaining a constant junction thickness. These findings show that amino acid composition and sequence directly control enhancements in electron transport in heme-binding peptides. Overall, this study demonstrates the potential of using sequence-defined synthetic peptides inspired by nature as functional bioelectronic materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Sequence Programmable Order–Disorder Transitions in Supramolecular Assembly of Peptide Nanofibers

Protein–protein interactions determine the assembly of complexes that are responsible for numerous key biological processes. The assembly of many natural protein complexes is mediated by post-translational structural changes and environmental stimuli. In this study, we show that incorporation of adjacent lysine residues results in the pH-tunable stability of peptide secondary structure and assembly, allowing for the incorporation of complementary order-inducing motifs. The strategic placement of cysteine pairs in the same peptide sequence results in redox-dependent disulfide staple formation, inducing a transition from random coil to β-sheet conformation and subsequent supramolecular nanofiber assembly from otherwise disordered peptide monomers. Spectroscopic, imaging, molecular dynamics, and kinetic studies highlight the critical role of sequence motif location, oligomerization, and the competitive interplay between intra- and interpeptide disulfide bonding in determining assembly outcomes. We extend this approach to demonstrate phosphorylation-dependent assembly from the design of the same parent peptide sequence, suggesting a general approach to the design of diverse stimulus-responsive peptide sequences for supramolecular assembly. Furthermore, these findings also provide a framework for investigating sequence-dependent pathways in amyloid fiber formation with potential implications for neurodegenerative disease research.

Disulfides

Sequence-specific dynamic DNA bending explains mitochondrial TFAM’s dual role in DNA packaging and transcription initiation

Abstract Mitochondrial transcription factor A (TFAM) employs DNA bending to package mitochondrial DNA (mtDNA) into nucleoids and recruit mitochondrial RNA polymerase (POLRMT) at specific promoter sites, light strand promoter (LSP) and heavy strand promoter (HSP). Herein, we characterize the conformational dynamics of TFAM on promoter and non-promoter sequences using single-molecule fluorescence resonance energy transfer (smFRET) and single-molecule protein-induced fluorescence enhancement (smPIFE) methods. The DNA-TFAM complexes dynamically transition between partially and fully bent DNA conformational states. The bending/unbending transition rates and bending stability are DNA sequence-dependent—LSP forms the most stable fully bent complex and the non-specific sequence the least, which correlates with the lifetimes and affinities of TFAM with these DNA sequences. By quantifying the dynamic nature of the DNA-TFAM complexes, our study provides insights into how TFAM acts as a multifunctional protein through the DNA bending states to achieve sequence specificity and fidelity in mitochondrial transcription while performing mtDNA packaging.

59 BASIC BIOLOGICAL SCIENCES

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W

Higher-order Zeno sequences

The quantum Zeno effect typically refers to freezing the dynamics of a quantum system through frequent observations. In general, quantum Zeno dynamics is obtained with an error of order 𝒪⁢(1/𝑁), where 𝑁 is the number of projective measurements performed within a fixed evolution time. In this work, we develop higher-order Zeno sequences that achieve faster convergence to Zeno dynamics, yielding an improved error scaling of 𝒪⁢(1/𝑁 2⁢𝑘 ), where 𝑘 describes the order of the Zeno sequence. This is achieved by relating higher-order Zeno sequences to higher-order Trotter formulas that achieve similar convergence behavior. We leverage this relation to develop higher-order Zeno sequences for different manifestations of the quantum Zeno effect, including frequent projective measurements and unitary kicks. We go on to discuss achieving quantum Zeno dynamics through periodic control fields of high frequency. We explicitly develop control fields that yield a second-order type improvement in the Zeno error scaling and present shorter Zeno sequences. Finally, we discuss the connection to randomized and Uhrig dynamical decoupling to develop more efficient implementations in the weak-coupling regime.

Quantum Zeno dynamics

Next-Generation Sequencing Data from a CUT&RUN Study of R. toruloides IFO0880 Cse4 and Orc1 Binding Sites

Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

Genome Engineering

Sequencing and analysis of 131 SARS-CoV-2 isolates in previously sampled and unsampled regions of Jordan from 2020 to 2023

The Hashemite Kingdom of Jordan remains an understudied country for next generation sequencing analysis of SARS-CoV-2 genomes collected during the 2019 pandemic. Here we provide 131 additional reference genomes collected between 2020–2023 from SARS-CoV-2-positive patients across Jordan. Phylogenetic analysis supports existing pandemic narratives of changing clade dominance over time and adds genomes in novel Jordanian locations and timepoints to make Jordan SARS-CoV-2 databases more comprehensive. Samples from the less-sequenced cities of Ajloun, Jaresh, Karak, and Madaba identified previously unreported lineages while Amman, Irbid, and Zarqa have existing sequencing efforts bolstered. Despite many incomplete patient records and a relatively small sample size, we observe interesting symptom patterns that support existing global and Jordanian pandemic narratives. We note how in-country COVID-19 pandemic genomic studies showcase Jordan’s efforts to expand next generation sequencing capabilities, especially through the leveraging of EDGE COVID-19, a bioinformatics platform for performing rapid, batched analysis of SARS-CoV-2 sequencing that streamlines sample processing prepared from a network of hospital locations.

60 APPLIED LIFE SCIENCES