Search NASA⌕ Search

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

Algorithm 1049: The Delaunay Density Diagnostic

Accurate approximation of a real-valued function depends on two aspects of the available data: the density of inputs within the domain of interest and the variation of the outputs over that domain. There are few methods for assessing whether the density of inputs is sufficient to identify the relevant variations in outputs—i.e., the “geometric scale” of the function—despite the fact that sampling density is closely tied to the success or failure of an approximation method. In this article, we introduce a general purpose, computational approach to detecting the geometric scale of real-valued functions over a fixed domain using a deterministic interpolation technique from computational geometry. The algorithm is intended to work on scalar data in moderate dimensions (2–10). Our algorithm is based on the observation that a sequence of piecewise linear interpolants will converge to a continuous function at a quadratic rate (in L 2 norm) if and only if the data are sampled densely enough to distinguish the feature from noise (assuming sufficiently regular sampling). We present numerical experiments demonstrating how our method can identify feature scale, estimate uncertainty in feature scale, and assess the sampling density for fixed (i.e., static) datasets of input–output pairs. Finally, we include analytical results in support of our numerical findings and have released lightweight code that can be adapted for use in a variety of data science settings.

97 MATHEMATICS AND COMPUTING↗

DIVA/DeviceEditor v6.1.2

DIVA is an end-to-end DNA design and construction management platform that streamlines how researchers design, build, and receive sequence-verified DNA constructs. Through a web-based BioCAD interface (DeviceEditor), researchers independently design DNA constructs and submit them to a centralized queue with a single action. Designs progress transparently through standardized states which allow researchers to track status and access finished constructs via a central DNA repository. Submitted designs are reviewed by dedicated staff for feasibility and optimization, reducing costly failures and improving downstream execution. Automated DNA assembly software optimizes construction strategies by reusing existing parts where possible and sourcing synthetic DNA only when needed. Standardized, sequence-agnostic assembly methods enable many independent constructs to be built in parallel using lab automation, dramatically increasing throughput. High-throughput next-generation sequencing is used to verify construct accuracy, with flexible platforms selected based on task requirements. Throughout the process, detailed success and failure data are captured and analyzed, enabling continuous improvement of assembly protocols. Compared to traditional, manual DNA construction workflows, DIVA offers higher scalability, transparency, reproducibility, and data-driven optimization.

Plahar, Hector [Lawrence Berkeley National Laborat↗

From Ensemble Climate to Ensemble Impacts

Many climate-risk tools rely on ensemble mean projections or endpoint climate snapshots to characterize future hazards. Although convenient for communication, these representations remove the statistical, temporal, and physical information that real infrastructure systems respond to. Infrastructure degradation and failure arise from extremes, sequences, cumulative stress, compound hazards, and nonlinear fragility relationships, none of which survive ensemble averaging or temporal compression. Power-system failure statistics and cascading failure models further show that infrastructure risk is dominated by tail events and path-dependent dynamics rather than by mean conditions. This paper demonstrates why ensemble mean or endpoint-only climate representations are mathematically and physically inconsistent with engineering-grade risk analysis. We outline a model-resolved, time-series-based workflow that preserves extremes, variability, and sequencing by propagating each climate-model realization independently through hazard formation, exposure, fragility, and cascading failure mechanisms. Taking the ensemble of impacts—rather than the ensemble of climate—provides a defensible, physically coherent foundation for infrastructure resilience planning, regulatory compliance, and long-term investment decisions.

54 - ENVIRONMENTAL SCIENCES/GLOBAL CLIMATE CHANGE ↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Behavior of Single-Line-Ground Faults in Inverter-Based Resource Dominated Grids Explained

It has been observed by protection engineers that it is difficult for a protective relay to identify the faulted phase during a single-line-ground (SLG) fault in a power system with a high ingression of inverter-based resources (IBR) using currents (phase or sequence). Further studies using electromagnetic transient (EMT) simulation show that the initial operating conditions of the IBRs influence the response of phase currents during an SLG fault. In this letter, we conduct a quantitative analysis using sequence components. We find that the pre-fault condition determines the relative position of the current contributed by the grid versus that from the IBR, and further dictates which phase has the largest magnitude during an SLG condition. Finally, this finding is further verified by the EMT simulation results.

electromagnetic transient simulation↗

The promising role of proteomes and metabolomes in defining the single-cell landscapes of plants

The plant community has a strong track-record of RNA sequencing technology deployment, which combined with the recent advent of spatial platforms (e.g., 10x genomics), has resulted in an explosion of outstanding single cell and nuclei datasets that can be put in an in situ context within tissues (e.g., a cell atlas)1. In the genomics era, application of proteomics technologies in the plant sciences has always trailed behind that of RNA sequencing technologies, largely due to accessibility, ease-of-use and access to expertise along with depth of analysis benefits. On the other hand, the use of early analytical tools for characterizing small molecules (metabolites) from plant systems predates nucleic acid sequencing and proteomics analysis2, as the search for plant-based natural products has played a significant role in improving human health throughout history. However, the employment of proteomics and metabolomics assays for characterizing plant cell processes now remains significantly behind transcriptional approaches, even though both provide a direct functional readout of cell states and phenotypes.

Anderton, Christopher R. [BATTELLE (PACIFIC NW LAB↗

Divergent viral phosphodiesterases for immune signaling evasion

Cyclic dinucleotides (CDNs) and other short oligonucleotides play fundamental roles in immune system activation in organisms ranging from bacteria to humans. In response, viruses use phosphodiesterase (PDE)-mediated oligonucleotide cleavage for immune evasion, a strategy whose diversity has not yet been explored. Here, we use a canonical 2H PDE (2H PDE) structure-based search of prokaryotic and eukaryotic viral sequences to identify an exceptional diversity of 2H PDEs across the virome, including enzymes not detectable with sequence search methods alone. Despite active site conservation, biochemical experiments reveal remarkable substrate specificity of these PDEs that corresponds to variations in the core 2H fold. This nuanced specificity allows 2H PDEs to selectively degrade oligonucleotide messengers to avoid interfering with host nucleotide signaling. Together, these findings nominate viral 2H PDEs as key regulators of CDN signaling across the tree of life.

CBASS↗

Biogenic carbon capture at pulp mills via sodium spiking and oxy-fuel calcination

Over 13 million metric tons of biogenic carbon dioxide (CO 2 ) are mineralized yearly in United States (US) pulp mill recovery boilers as molten sodium carbonate. The mineralized CO 2 is released downstream in the rotary lime kiln where it can be captured at a relatively low cost for permanent sequestration. The use of biocarbon and bioenergy at pulp mills enables the possibility for atmospheric carbon removal when biogenic CO2 is captured and sequestered. We demonstrate the feasibility of capturing biogenic CO 2 via sodium spiking coupled with oxy-fuel calcination in the rotary lime kiln at existing kraft pulp mills. Sodium (Na) spiking elevates the Na ion loading in the kraft chemical looping process by replacing chlorine-based bleaching with a highly alkaline bleaching sequence. Chemical pulping processes are modeled and simulated to understand the technical limits of implementing sodium spiking in existing pulp mills. Each 1 % of sodium added to the bleaching operations increases the rate of CO 2 mineralization by 4 %, and the maximum increase in sodium content in the kraft process is 9 %. Without sodium spiking, estimated CO 2 capture costs are $131 and $107/mt CO 2 for air and oxy-fuel combustion, respectively. With the implementation of sodium spiking, the cost of CO 2 captured decreases by 27 % and 31 %, showing costs of $96 and $74/mt CO 2 for air and oxy-fuel combustion, respectively. Oxy-fuel combustion reduces the costs of CO 2 capture compared to air combustion, but merged deployment with sodium spiking seems to be the most cost-effective pathway to integrate carbon capture.

60 APPLIED LIFE SCIENCES↗

Distinct evolutionary lineages of Schistocephalus parasites infecting co-occurring sculpin and stickleback fishes in Alaska

Sculpins (coastrange and slimy) and sticklebacks (ninespine and threespine) are widely distributed fishes cohabiting 2 south-central Alaskan lakes (Aleknagik and Iliamna), and all these species are parasitized by cryptic diphyllobothriidean cestodes in the genus Schistocephalus. The goal of this investigation was to test for host-specific parasitic relationships between sculpins and sticklebacks based upon morphological traits (segment counts) and sequence variation across the NADH1 gene. A total of 446 plerocercoids was examined. Large, significant differences in mean segment counts were found between cestodes in sculpin (mean = 112; standard deviation [S.D.] = 15) and stickleback (mean = 86; S.D. = 9) hosts within and between lakes. Nucleotide sequence divergence between parasites from sculpin and stickleback hosts was 20.5%, and Bayesian phylogenetic analysis recovered 2 well-supported clades of cestodes reflecting intermediate host family (i.e. sculpin, Cottidae vs stickleback, Gasterosteidae). Our findings point to the presence of a distinct lineage of cryptic Schistocephalus in sculpins from Aleknagik and Iliamna lakes that warrants further investigation to determine appropriate evolutionary and taxonomic recognition.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput Single-Cell Proteomics and Transcriptomics from the Same Cells with a Nanoliter-Scale Spin-Transfer Approach

Single-cell multiomic platforms provide a comprehensive snapshot of cellular states and cell types by offering critical insights into the spatiotemporal regulation of biomolecular networks at a systems level, thereby defining the basis of multicellularity. Here, we introduce nanoSPINS, an advanced platform that enables high-throughput profiling and integrative analysis of the transcriptome and proteome from the same single cells using RNA sequencing and isobaric labeling LC-MS-based proteomics, respectively. NanoSPINS can efficiently transfer mRNA-containing droplets across two microarrays via a centrifugation-based approach, while proteins are retained on the initial platform. Benchmarking of nanoSPINS on two cell lines demonstrates its ability to generate global proteomic and transcriptomic profiles that align well with previously established methodologies/platforms. The incorporation of isobaric TMTpro labeling into this single-cell multiomics platform significantly enhances the throughput of single-cell proteomic analyses. Through the high-throughput quantification of the proteome and transcriptome, nanoSPINS not only facilitates the identification of molecular features at both mRNA and protein level but also provides larger sample sizes for improved statistical power in clustering and differential abundance. Given the broad applicability of single-cell multiomics in biological research and clinical settings, we believe nanoSPINS represents a powerful platform for the characterization of heterogeneous cell populations.

multi 'omics↗

Hamiltonian simulation in Zeno subspaces

Here, we investigate the quantum Zeno effect as a framework for designing and analyzing quantum algorithms for Hamiltonian simulation. We show that frequent projective measurements of an ancilla qubit register can be used to simulate quantum dynamics on a target qubit register with a circuit complexity similar to randomized approaches. The classical sampling overhead in the latter approaches is traded for ancilla qubit overhead in Zeno-based approaches. A second-order Zeno sequence is developed to improve scaling and implementations through unitary kicks are discussed. We derive rigorous error bounds that allow for identifying the associated circuit complexities for the first- and second-order Zeno sequences. We show that the circuits over the combined register can be identified as a subroutine commonly used in post-Trotter Hamiltonian simulation methods. We build on this observation to reveal connections between different Hamiltonian simulation algorithms.

Hamiltonian simulation↗

Theory of topological defects and textures in two-dimensional quantum orders with spontaneous symmetry breaking

In this article, we consider two-dimensional (2d) quantum many-body systems with long-range orders, where the only gapless excitations in the spectrum are Goldstone modes of spontaneously broken continuous symmetries. To understand the interplay between classical long-range order of local order parameters and quantum order of long-range entanglement in the ground states, we study the topological point defects and textures of order parameters in such systems. We show that the universal properties of point defects and textures are determined by the remnant symmetry enriched topological order in the symmetry-breaking ground states with a nonfluctuating order parameter, and provide a classification for their properties based on the inflation-restriction exact sequence. We highlight a few phenomena revealed by our theory framework. First, in the absence of intrinsic topological orders, we show a connection between the symmetry properties of point defects and textures to deconfined quantum criticality. Second, when the symmetry-breaking ground state has intrinsic topological orders, we show that the point defects can permute different anyons when braided around. They can also obey projective fusion rules in the sense that multiple vortices can fuse into an Abelian anyon, a phenomenon for which we coin “defect fractionalization.” Finally, we provide a formula to compute the fractional statistics and fractional quantum numbers carried by textures (skyrmions) in Abelian topological orders.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

High-Power Clock Laser Spectrally Tailored for High-Fidelity Quantum State Engineering

Highly frequency-stable lasers are ubiquitous tools for optical-frequency metrology, precision interferometry, and quantum information science. While making a universally applicable laser is unrealistic, spectral noise can be tailored for specific applications. Here we report a high-power 698-nm clock laser with a maximum output of 4W and minimized frequency noise up to a few kHz Fourier frequency, together with long-term instability of 3.5 × 10 −17 at one to thousands of seconds. The laser-frequency noise is precisely characterized with atom-based spectral analysis that employs a pulse sequence designed to suppress sensitivity to intensity noise. This method provides universally applicable tunability of the spectral response and analysis of quantum sensors over a wide frequency range. With the optimized laser system characterized by this technique, we achieve an average single-qubit Clifford gate fidelity of up to 𝐹$^2_1$ = 0.999⁢64⁢(3) when simultaneously driving 3000 optical qubits with a homogeneous Rabi frequency ranging from 10 Hz to 1 kHz. This result represents the highest single optical-qubit-gate fidelity for a large number of atoms.

atomic gases↗

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES↗

Surfactant-like peptide gels are based on cross-β amyloid fibrils

Surfactant-like peptides, in which hydrophilic and hydrophobic residues are encoded within different domains in the peptide sequence, undergo facile self-assembly in aqueous solution to form supramolecular hydrogels. These peptides have been explored extensively as substrates for the creation of functional materials since a wide variety of amphipathic sequences can be prepared from commonly available amino acid precursors. The self-assembly behavior of surfactant-like peptides has been compared to that observed for small molecule amphiphiles in which nanoscale phase separation of the hydrophobic domains drives the self-assembly of supramolecular structures. Here, we investigate the relationship between sequence and supramolecular structure for a pair of bola-amphiphilic peptides, Ac-KLIIIK-NH 2 (L2) and Ac-KIIILK-NH 2 (L5). Despite similar length, composition, and polar sequence pattern, L2 and L5 form morphologically distinct assemblies, nanosheets and nanotubes, respectively. Cryo-EM helical reconstruction was employed to determine the structure of the L5 nanotube at near-atomic resolution. Rather than displaying self-assembly behavior analogous to conventional amphiphiles, the packing arrangement of peptides in the L5 nanotube displayed steric zipper interfaces that resembled those observed in the structures of β-amyloid fibrils. Like amyloids, the supramolecular structures of the L2 and L5 assemblies were sensitive to conservative amino acid substitutions within an otherwise identical amphipathic sequence pattern. This study highlights the need to better understand the relationship between sequence and supramolecular structure to facilitate the development of functional peptide-based materials for biomaterials applications.

Das, Abhinaba [Emory University, Atlanta, GA (Unit↗