Search NASASearch

SEARCH · Search NASA

Results for “Base Sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Performance Evaluation of a Novel Sequence-Based Directional Detection Strategy for Protection of Active Distribution Networks

Directional elements are relied on to achieve selectivity in fault detection in power systems. Although such elements have been deployed successfully for many years, there is an increased need for novel methods to deal with the unique challenges of directional protection in modern distribution networks. This article analyzes the impact of inverter-based resources (IBRs) on existing directional protection methods in distribution systems. It identifies parts of such elements that pose a risk of misoperation when IBRs are used in distribution networks. The authors have developed a new directional detection method for unbalanced faults in such networks using superimposed symmetrical sequence quantities. The phase angle of the superimposed negative sequence admittance is used to determine fault direction. The paper also presents a real-time co-simulation platform between a simulated distribution system and physical protection relay, using OPAL-RT. An SEL-411L relay is used to program the detection algorithm. This hardware-in-the-loop (HIL) setup is used to verify the performance of the method and the results are compared with existing directional methods

24 POWER TRANSMISSION AND DISTRIBUTION

Sequence-based generative AI design of versatile tryptophan synthases

Enzymes are powerful and sustainable catalysts, but their widespread application is limited by the difficulty of identifying functional starting points for optimization, creating a major bottleneck in early- stage biocatalyst discovery. Designing libraries of such starting enzymes remains particularly challenging. Here, we use the GenSLM protein language model to generate novel β-subunit of tryptophan synthase (TrpB) enzymes that express in Escherichia coli and are both stable and catalytically active. Many generated TrpBs also display significant substrate promiscuity, outperforming their natural counterparts on non-native substrates. Some even surpass laboratory-evolved TrpBs. Comparison of the most-active and most-promiscuous generated TrpB to its closest natural homolog confirms that the enhanced versatility is absent from the natural enzyme, highlighting the creative potential of generative models. These results demonstrate that the generated TrpBs not only preserve natural structure and function but also acquire non-natural properties, establishing generative models as powerful tools for biocatalyst discovery and engineering.

biocatalysis

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan

Comparison of Sequence Component-Based Fault Detection and Relay Coordination Algorithms in Inverter-Based Networks

Protection of inverter-based microgrids using sequence component-based relaying schemes is a promising solution. These methods offer several advantages, including lower computational requirements, compatibility with commercial relay systems, and cost-effectiveness compared to communication-based approaches. This article investigate the performance of various sequence component based schemes with the objective of identifying the algorithms that provide the best fault detection and relay coordination, solely relying on local voltages and current at relay terminals. Positive, negative and zero sequence impedance, admittance and power detection algorithms were tested on modified IEEE 13 bus test network for various shunt faults (LG, LL, LLG, LLL). Hardware-in-the-loop validation was achieved using the Typhoon real-time simulator, interfacing with a SEL 751 relay. This research demonstrates that while several algorithms are capable of detecting faults with sufficient accuracy, only a few are effective in achieving proper coordination. Validation results indicate that the negative sequence power approach provides the best performance in both fault detection and coordination.

Patel, Deepika [ORNL] (ORCID:0000000341099994)

A 1-year study on SARS-CoV-2 variant shifts in wastewater using dPCR: comparison with clinical and GISAID data

Wastewater testing can be used to monitor SARS-CoV-2 infections in communities. Data from PCR-based wastewater testing are usually available to public health authorities within 5–7 days after excreta and other body fluids enter the sewer. While PCR-based methods can accurately detect and quantify SARS-CoV-2, sequencing-based methods are usually required to distinguish between variants, delaying the results and adding cost to the process. We developed and assessed a novel, customizable digital PCR (dPCR)-based genotyping method for SARS-CoV-2 variant detection in wastewater, which is more cost-effective, faster, and more accessible than sequencing. This approach was applied to more than 1,400 wastewater samples

Wilton, Rose

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES

Viromics approaches for the study of viral diversity and ecology in microbiomes

Viruses are found across all ecosystems and infect every type of organism on Earth. Traditional culture-based methods have proven insufficient to explore this viral diversity at scale, driving the development of viromics, the sequence-based analysis of uncultivated viruses. Viromics approaches have been particularly useful for studying viruses of microorganisms, which can act as crucial regulators of microbiomes across ecosystems. They have already revealed the broad geographic distribution of viral communities and are progressively uncovering the expansive genetic and functional diversity of the global virome. Moving forward, large-scale viral ecogenomics studies combined with new experimental and computational approaches to identify virus activity and host interactions will enable a more complete characterization of global viral diversity and its effects.

Ecology

wavess 1.2: presenting an HLA-aware within-host virus sequence simulation framework

Motivation Understanding how virus sequences are shaped by selection can inform vaccine design and transmission inference. Modeling within-host evolution to interrogate these questions requires a detailed mechanistic framework that accurately captures sequence diversification. The CD8 + cytotoxic T-lymphocyte (CTL) response plays an important role in immune-mediated selection and can leave strong signatures in virus sequences; however, existing sequence-based within-host virus modeling frameworks do not explicitly include a human leukocyte antigen (HLA)-aware CTL response. Results We extended our previously published within-host sequence evolution simulator, wavess, to include an explicit CTL response, and share a method for identifying HLA-specific CTL epitopes given a founder virus sequence. We also updated the model to permit a variable recombination rate, which allows for modeling non-adjacent genes, segmented genomes, and recombination hotspots. These extensions to wavess allow for more accurate simulation of viruses and virus genes, particularly in regions of the genome where the immune response is dominated by CTLs (rather than antibodies). It also provides the foundation for investigations of how these newly-added biological mechanisms influence within-host evolution. Availability and implementation The core of wavess is written in Python 3, with helper functions written in R. It is available at https://github.com/MolEvolEpid/wavess.

60 APPLIED LIFE SCIENCES

Geologic Characterization of the South Georgia Rift Basin for Source Proximal CO2 Storage

The project Geologic Characterization of the South Georgia Rift Basin for Source Proximal CO2 Storage is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The South Carolina Research Foundation and partners evaluated the feasibility of CCS in the Jurassic/ Triassic (J / TR) saline formations of the buried Mesozoic South Georgia Rift (SGR) Basin that extends from South Carolina into Georgia. The J / TR sequence, based on preliminary assessment of limited geologic and geophysical data, appears to have both the appropriate areal extent and multiple horizons to permanently and safely store CO2 The presence of several igneous rock layers within the sequence may potentially provide adequate seals to prevent upward CO2 migration into the Coastal Plain aquifer systems. Approximately 81 kilometers of 2-D seismic reflection data were collected by Bay Geophysical, Inc. to explore a portion of the SGR located in southern Georgia. The 81 kilometers were divided into two lines approximately 40.5 kilometers each, with Line 1 intersecting Georgia well GGS 3457. Line 2 intersects Line 1 at the southern portion of Line 1 to maximize the extent of coverage away from GGS-3457 (a deep well drilled in the 1980s for oil and gas exploration). This well had a set of usable logs, including gamma and neutron logs that provided promising results related to CO2 storage. Results showed sandstone with porosity values greater than 10 percent and a thickness of 120 meters. The design of the seismic shot was to extrapolate information away from the well and to better define the extent of the SGR and the necessary reservoir and caprock for a successful CO2 injection. A numerical simulation model of CO2 Injection and migration was developed based on the geology log for the GGS-3457 well. The simulation model was used to investigate the feasibility of injecting 30 million metric tons of CO2 into SGR J / TA sediments and integrity of the diabase layers as seals to prevent CO2 migration.

2-D seismic

Prediction of Specificity of α-Conotoxins to Subtypes of Human Nicotinic Acetylcholine Receptors with Semi-supervised Machine Learning

Conotoxins are a family of highly toxic neurotoxins composed of cysteine-rich peptides produced by marine cone snails. The most lethal cone snail species to humans is Conus geographus, with fatality rates of up to ∼65% from a single sting, which is caused mostly by the activity of α-conotoxins against human nicotinic acetylcholine receptors (nAChRs). While sequence-based machine learning (ML) classifiers have been trained to identify targets of conotoxins binding voltage-gated ion channels, no ML model has been built to predict the subtype-specific nAChR targets of α-conotoxins. Here, we trained an ML model in a semi-supervised manner to predict the specificity of α-conotoxin binding toward different human nAChR subtypes to overcome the challenge of limited data in subtype-specific nAChR targets of α-conotoxins and the issue that one α-conotoxin can bind multiple nAChR subtypes with high selectivity. We considered additional features of sequences of α-conotoxins in training our ML model, including the secondary structure propensities and electrostatic properties, which resulted in better prediction capability for the ML model. Notably, we identify that most α-conotoxins bind to α3β2, α1γδ, and α7 subtypes of human nAChRs. Our findings from this study provide a framework for predicting targets of various kinds of toxins.

59 BASIC BIOLOGICAL SCIENCES

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State

Extracellular DNA Alters Detection of Subtle Bacterial Responses to Soil Rewetting

Microbial communities are often characterized using DNA-based sequencing, but these approaches also capture extracellular DNA (exDNA) released from dead cells, potentially altering inference about microbial responses to environmental change. This may be especially important during pulse disturbances, such as soil drying–rewetting, which can increase microbial mortality and transient necromass pools. We assessed whether exDNA altered inference about bacterial responses to drying–rewetting (an 80 mm simulated rainfall event following a 28-day drought) in conventionally tilled corn and perennial switchgrass soils. We quantified bacterial abundance (16 S rRNA gene copies), alpha diversity, and community composition in paired soil samples with exDNA included (+ exDNA) and in samples treated with propidium monoazide (PMAxx) to reduce amplification of exDNA (− exDNA). At our level of replication (n = 4), PMAxx treatment did not significantly alter overall temporal response patterns (i.e., no significant main effect of DNA treatment or DNA × time interaction). However, PMAxx treatment increased sensitivity to detect some pairwise temporal changes in bacterial abundance and community composition in corn soils following rewetting. exDNA pools were proportionally highest immediately after rewetting in corn soils, suggesting transient extracellular DNA may contribute to masking during disturbance recovery. In contrast, PMAxx treatment had comparatively small effects in switchgrass soils, which exhibited weaker temporal responses overall. Inclusion of exDNA also changed which taxa appeared most responsive to rewetting. Together, our results suggest that exDNA does not uniformly bias soil microbial inference, but may reduce detectability of subtle disturbance-driven shifts in certain soils. Future studies should advance knowledge of microbial turnover and necromass dynamics, particularly using multiple complementary methods, to help predict when exDNA is most likely to influence ecological inference.

drying-rewetting

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES

The SRG/eROSITA All-Sky Survey: Optical identification and properties of galaxy clusters and groups in the western galactic hemisphere

The first SRG/eROSITA All-Sky Survey (eRASS1) provides the largest intracluster medium-selected galaxy cluster and group catalog covering the western Galactic hemisphere. Compared to samples selected purely on X-ray extent, the sample purity can be enhanced by identifying cluster candidates using optical and near-infrared data from the DESI Legacy Imaging Surveys. Using the red-sequence-based cluster findereROMaPPer, we measured individual photometric properties (redshiftz λ , richnessλ, optical center, and BCG position) for 12000 eRASS1 clusters over a sky area of 13 116 deg 2 , augmented by 247 cases identified by matching the candidates with known clusters from the literature. The median redshift of the identified eRASS1 sample isz= 0.31, with 10% of the clusters atz> 0.72. The photometric redshifts have an accuracy ofδz/(1 +z) ≲ 0.005 for 0.05 specand velocity dispersionσ) were measured a posteriori for a subsample of 3210 and 1499 eRASS1 clusters, respectively, using an extensive compilation of spectroscopic redshifts of galaxies from the literature. We infer that the primary eRASS1 sample has a purity of 86% and optical completeness >95% forz> 0.05. For these and further quality assessments of the eRASS1 identified catalog, we applied our identification method to a collection of galaxy cluster catalogs in the literature, as well as blindly on the full Legacy Surveys covering 24069 deg 2 . Using a combination of these cluster samples, we investigated the velocity dispersion-richness relation, finding that it scales with richness as log(λ norm ) = 2.401 × log(σ) − 5.074 with an intrinsic scatter ofδ in = 0.10 ± 0.01 dex. The primary product of our work is the identified eRASS1 cluster catalog with high purity and a well-defined X-ray selection process, opening the path for precise cosmological analyses presented in companion papers.

Astronomy & Astrophysics

Lanthanide binding peptide surfactants at air–aqueous interfaces for interfacial separation of rare earth elements

Rare earth elements (REEs) are critical materials to modern technologies. They are obtained by selective separation from mining feedstocks consisting of mixtures of their trivalent cation. We are developing an all-aqueous, bioinspired, interfacial separation using peptides as amphiphilic molecular extractants. Lanthanide binding tags (LBTs) are amphiphilic peptide sequences based on the EF-hand metal binding loops of calcium-binding proteins which complex selectively REEs. We study LBTs optimized for coordination to Tb 3+ using luminescence spectroscopy, surface tensiometry, X-ray reflectivity, and X-ray fluorescence near total reflection, and find that these LBTs capture Tb 3+ in bulk and adsorb the complex to the interface. Molecular dynamics show that the binding pocket remains intact upon adsorption. We find that, if the net negative charge on the peptide results in a negatively charged complex, excess cations are recruited to the interface by nonselective Coulombic interactions that compromise selective REE capture. If, however, the net negative charge on the peptide is −3, resulting in a neutral complex, a 1:1 surface ratio of cation to peptide is achieved. Surface adsorption of the neutral peptide complexes from an equimolar mixture of Tb 3+ and La 3+ demonstrates a switchable platform dictated by bulk and interfacial effects. The adsorption layer becomes enriched in the favored Tb 3+ when the bulk peptide is saturated, but selective to La 3+ for undersaturation due to a higher surface activity of the La 3+ complex.

Ortuno Macias, Luis E. (ORCID:0000000284342192)

Association between optically identified galaxy clusters and the underlying dark matter halos

Clusters of galaxies trace massive dark matter halos in the Universe, but they can include multiple halos projected along lines of sight. Here, we study the halos contributing to clusters using the Cardinal simulation, which mimics the Dark Energy Survey data. We use the red-sequence-based cluster finding algorithm redMaPPer as a case study. For each cluster, we identify the halos hosting its member galaxies, and we define the main halo as the one contributing the most to the cluster's richness ($λ$, the estimated number of member galaxies). At $z=0.3$, for clusters with $λ> 60$, the main halo typically contributes to $92\%$ of the richness, and this fraction drops to $67\%$ for $λ\approx 20$. Defining "clean" clusters as those with $\geq50\%$ of the richness contributed by the main halo, we find that $100\%$ of the $λ> 60$ clusters are clean, while $73\%$ of the $λ\approx 20$ clusters are clean. Three halos can usually account for more than $80\%$ of the richness of a cluster. The main halos associated with redMaPPer clusters have a completeness ranging from $98\%$ at virial mass $10^{14.6}~h^{-1}M_{\odot}$ to $64\%$ at $10^{14}~h^{-1}M_{\odot}$. In addition, we compare the inferred cluster centers with true halo centers, finding that $30\%$ of the clusters are miscentered with a mean offset $40\%$ of the cluster radii, in agreement with recent X-ray studies. These systematics worsen as redshift increases, but we expect that upcoming surveys extending to longer wavelengths will improve the cluster finding at high redshifts. Our results affirm the robustness of the redMaPPer algorithm and provide a framework for benchmarking other cluster-finding strategies.

79 ASTRONOMY AND ASTROPHYSICS

SPADES

Sequence-based Pathogen-Agnostic Diagnostics/Detection Solution

Li, Po-E [Los Alamos National Laboratory]