Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Ni-catalysed dicarbofunctionalization for the synthesis of sequence-encoded cyclooctene monomers

The properties of polymeric materials can be modulated by factors such as sequence control or functional group modifications. However, the synthesis of new macromolecular scaffolds is limited by the accessibility of structurally diverse monomers. This work describes a one-step, nickel-catalysed synthesis of 5,6-diaryl cyclooctene monomers from the feedstock chemical 1,5-cyclooctadiene. The reaction proceeds in a modular, regio- and diastereoselective fashion, granting access to both homo- and hetero-diaryl cyclooctene monomers that smoothly undergo ring-opening metathesis polymerization (ROMP). The resulting 1,2-diaryl-substituted polymers possess sequences with head-to-head styrene dyads that have not been previously explored, giving rise to unique and tunable properties. Density functional theory calculations highlight mechanistic aspects of the nickel-catalysed diarylation reaction and the ruthenium-catalysed ROMP process, revealing a previously unappreciated role of the boronic ester in promoting migratory insertion, which was leveraged to provide enantioinduction.

Catalytic mechanisms↗

ALS mutations disrupt self-association between the ubiquilin STI1 hydrophobic groove and internal placeholder sequences

Ubiquilins are molecular chaperones that play multifaceted roles in proteostasis, with point mutations in UBQLN2 leading to altered phase-separation properties and amyotrophic lateral sclerosis (ALS). Our mechanistic understanding of this essential process has been hindered by a lack of structural information on the STI1 domain, which is essential for ubiquilin chaperone activity and phase separation. Here, we present the first crystal structure of a ubiquilin-family STI1 domain bound to a transmembrane domain (TMD), and show that ALS mutations disrupt the STI1-TMD interaction. We further demonstrate that ubiquilins contain multiple conserved internal sequences that bind to the STI1 domain, including the PXX-repeat region that is a hotspot for ALS mutations. We propose that these placeholder sequences prevent solvent exposure of the STI1 hydrophobic groove and contribute to the multivalency that drives ubiquilin phase-separation. Together, this work provides a new paradigm for understanding how STI1 domains modulate ubiquilin chaperone activity and phase separation, and offers insights into the molecular basis of ALS pathogenesis.

Onwunma, Joan [Univ. of Toledo, OH (United States)↗

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES↗

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA↗

Performance Evaluation of a Novel Sequence-Based Directional Detection Strategy for Protection of Active Distribution Networks

Directional elements are relied on to achieve selectivity in fault detection in power systems. Although such elements have been deployed successfully for many years, there is an increased need for novel methods to deal with the unique challenges of directional protection in modern distribution networks. This article analyzes the impact of inverter-based resources (IBRs) on existing directional protection methods in distribution systems. It identifies parts of such elements that pose a risk of misoperation when IBRs are used in distribution networks. The authors have developed a new directional detection method for unbalanced faults in such networks using superimposed symmetrical sequence quantities. The phase angle of the superimposed negative sequence admittance is used to determine fault direction. The paper also presents a real-time co-simulation platform between a simulated distribution system and physical protection relay, using OPAL-RT. An SEL-411L relay is used to program the detection algorithm. This hardware-in-the-loop (HIL) setup is used to verify the performance of the method and the results are compared with existing directional methods

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Digital Three Level Space Vector Modulator for High Frequency Vector Sequence Generation

This letter proposes a digital high-speed three-level space vector pulse width modulator (3L-SVPWM). A conventional 3L-SVPWM is typically computation-based, involving a sequential execution of sub-tasks on a digital signal processor (DSP) based controller. The resulting high computation time of 5.4 μs limits the implementation of additional control blocks for switching frequencies greater than 100 kHz. This is overcome by transforming sub-tasks into digital blocks with 1-0 decisions and simpler arithmetic operations. The sub-task blocks are executed concurrently on a programmable logic device (PLD). Hence, a fast 3L-SVPWM execution in 140 ns is achieved. The proposed digital 3L-SVPWM enables high switching frequency operation of wide bandgap (WBG) device-based 3 L inverters to generate high fundamental frequency waveforms. A finite state machine is an integral part of the proposed implementation with the ability to generate any vector sequence, maximizing the usage of redundant vector states in 3L-SVPWM. Here, the proposed digital 3L-SVPWM operation is demonstrated with a GaN-based 3 L active neutral point clamped (3L-ANPC) inverter. Experimental results are presented at 250 kHz switching frequency to generate vector sequences for center-aligned SVPWM (CA-SVPWM) and common mode voltage reduced SVPWM (CMVR-SVPWM). The results also showcase a high fundamental frequency generation capability of 10 kHz.

active neutral point clamped inverter↗

Complete genome sequence of Luteolibacter sp. strain Populi, a member of phylum Verrucomicrobiota isolated from the Populus trichocarpa rhizosphere

Luteolibacter sp. strain Populi is a bacterium from the phylum Verrucomicrobiota, isolated from the rhizosphere of a black cottonwood tree, Populus trichocarpa, from the Cascade mountains in Washington. Its 6.6-Mb chromosome was completely sequenced using Oxford Nanopore long-read sequencing and is predicted to encode 5,301 proteins and 60 RNAs.

59 BASIC BIOLOGICAL SCIENCES↗

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Design of Diverse, Functional Mitochondrial Targeting Sequences Across Eukaryotic Organisms Using Variational Autoencoder"

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

AI/ML↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Overview of the SCEC/USGS Community Stress Drop Validation Study Using the 2019 Ridgecrest Earthquake Sequence

We present initial findings from the ongoing Community Stress Drop Validation Study to compare spectral stress-drop estimates for earthquakes in the 2019 Ridgecrest, California, sequence. This study uses a unified dataset to independently estimate earthquake source parameters through various methods. Stress drop, which denotes the change in average shear stress along a fault during earthquake rupture, is a critical parameter in earthquake science, impacting ground motion, rupture simulation, and source physics. Spectral stress drop is commonly derived by fitting the amplitude-spectrum shape, but estimates can vary substantially across studies for individual earthquakes. Sponsored jointly by the U.S. Geological Survey and the Statewide (previously, Southern) California Earthquake Center our community study aims to elucidate sources of variability and uncertainty in earthquake spectral stress-drop estimates through quantitative comparison of submitted results from independent analyses. The dataset includes nearly 13,000 earthquakes ranging from M 1 to 7 during a two-week period of the 2019 Ridgecrest sequence, recorded within a 1° radius. Here, in this article, we report on 56 unique submissions received from 20 different groups, detailing spectral corner frequencies (or source durations), moment magnitudes, and estimated spectral stress drops. Methods employed encompass spectral ratio analysis, spectral decomposition and inversion, finite-fault modeling, ground-motion-based approaches, and combined methods. Initial analysis reveals significant scatter across submitted spectral stress drops spanning over six orders of magnitude. However, we can identify between-method trends and offsets within the data to mitigate this variability. Averaging submissions for a prioritized subset of 56 events shows reduced variability of spectral stress drop, indicating overall consistency in recovered spectral stress-drop values.

58 GEOSCIENCES↗

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab↗

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso, ↗

KBase Narrative - Genome sequences of four bacterial strains isolated from organofluorine enrichment cultures

These genome announcement narratives are associated with the manuscript "Genome sequences of four bacterial strains isolated from organofluorine enrichment cultures" by Jennifer L. Goff, Chris Hahn, Rebecca A. Ingrassia, Chanistha Tiyapun, Gray G. Waldschmidt, Emma R. Smith, Miranda N. Marini, Olga Shevchenko, and Julia A. Maresca Assembeled genome sequences for all four genomes are available on this page and in the individual genome announcement narratives listed below. All methodological details can be found within the individual narratives.

Goff, Jennifer↗

KBase Narrative - Complete genome sequence of a novel Microbacterium sp. strain Clip185.

We have isolated a new species of Microbacterium, an Actinobacterium. We have temporarily named this bacterium as Microbacterium sp. strain Clip185 (hereafter called strain Clip185) from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii strain LMJ.RY0402.185141 (a Chlamydomonas Library project CLiP strain). We sequenced the whole genome of strain Clip185 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of this new Microbacterium species that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

Mitra, Mautusi↗

Complete genome sequence of Sphingobium yanoikuyae strain CC4533

We have isolated a new strain of Sphingobium yanoikuyae , which belongs to the class Alphaproteobacteria, order Sphingomonadales, and family Sphingomonadaceae. This carotenoid-producing strain is capable of degrading xenobiotics and is tolerant to toxic levels of six heavy metals. We have designated the newly isolated strain of S. yanoikuyae as S. yanoikuyae strain CC4533 (hereafter called strain CC4533) because it was isolated from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii wild type strain CC4533. We sequenced the whole genome of strain CC4533 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of S. yanoikuyae strain CC4533 that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

59 BASIC BIOLOGICAL SCIENCES↗

Next-generation sequencing dataset of genome-scale CRISPRi in Synechococcus sp. PCC 7002 across seven conditions

A 33,298-member sgRNA library developed for Synechococus sp. PCC 7002 was screened with two replicates across seven growth conditions and sequenced with Illumina NextSeq (paired end, 2x150 bp) for a total of ~700M reads. The original plasmid library and the library after transformation into a dCas9-containing and dCas9-absent strain were also sequenced as a reference for initial sgRNA abundance.

genome- wide screens environmental acclimation spe↗