Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

From soil to sequence: filling the critical gap in genome-resolved metagenomics is essential to the future of soil microbial ecology

Abstract Soil microbiomes are heterogeneous, complex microbial communities. Metagenomic analysis is generating vast amounts of data, creating immense challenges in sequence assembly and analysis. Although advances in technology have resulted in the ability to easily collect large amounts of sequence data, soil samples containing thousands of unique taxa are often poorly characterized. These challenges reduce the usefulness of genome-resolved metagenomic (GRM) analysis seen in other fields of microbiology, such as the creation of high quality metagenomic assembled genomes and the adoption of genome scale modeling approaches. The absence of these resources restricts the scale of future research, limiting hypothesis generation and the predictive modeling of microbial communities. Creating publicly available databases of soil MAGs, similar to databases produced for other microbiomes, has the potential to transform scientific insights about soil microbiomes without requiring the computational resources and domain expertise for assembly and binning.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Design of Diverse, Functional Mitochondrial Targeting Sequences Across Eukaryotic Organisms Using Variational Autoencoder"

Mitochondria play a key role in energy production and metabolism, making them a promising target for metabolic engineering and disease treatment. However, despite the known influence of passenger proteins on localization efficiency, only a few protein-localization tags have been characterized for mitochondrial targeting. To address this limitation, we leverage a Variational Autoencoder to design novel mitochondrial targeting sequences. In silico analysis reveals that a high fraction of the generated peptides (90.14%) are functional and possess features important for mitochondrial targeting. We characterize artificial peptides in four eukaryotic organisms and, as a proof-of-concept, demonstrate their utility in increasing 3-hydroxypropionic acid titers through pathway compartmentalization and improving 5-aminolevulinate synthase delivery by 1.62-fold and 4.76-fold, respectively. Moreover, we employ latent space interpolation to shed light on the evolutionary origins of dual-targeting sequences. Overall, our work demonstrates the potential of generative artificial intelligence for both fundamental research and practical applications in mitochondrial biology.

AI/ML↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Overview of the SCEC/USGS Community Stress Drop Validation Study Using the 2019 Ridgecrest Earthquake Sequence

We present initial findings from the ongoing Community Stress Drop Validation Study to compare spectral stress-drop estimates for earthquakes in the 2019 Ridgecrest, California, sequence. This study uses a unified dataset to independently estimate earthquake source parameters through various methods. Stress drop, which denotes the change in average shear stress along a fault during earthquake rupture, is a critical parameter in earthquake science, impacting ground motion, rupture simulation, and source physics. Spectral stress drop is commonly derived by fitting the amplitude-spectrum shape, but estimates can vary substantially across studies for individual earthquakes. Sponsored jointly by the U.S. Geological Survey and the Statewide (previously, Southern) California Earthquake Center our community study aims to elucidate sources of variability and uncertainty in earthquake spectral stress-drop estimates through quantitative comparison of submitted results from independent analyses. The dataset includes nearly 13,000 earthquakes ranging from M 1 to 7 during a two-week period of the 2019 Ridgecrest sequence, recorded within a 1° radius. Here, in this article, we report on 56 unique submissions received from 20 different groups, detailing spectral corner frequencies (or source durations), moment magnitudes, and estimated spectral stress drops. Methods employed encompass spectral ratio analysis, spectral decomposition and inversion, finite-fault modeling, ground-motion-based approaches, and combined methods. Initial analysis reveals significant scatter across submitted spectral stress drops spanning over six orders of magnitude. However, we can identify between-method trends and offsets within the data to mitigate this variability. Averaging submissions for a prioritized subset of 56 events shows reduced variability of spectral stress drop, indicating overall consistency in recovered spectral stress-drop values.

58 GEOSCIENCES↗

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab↗

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso, ↗

KBase Narrative - Genome sequences of four bacterial strains isolated from organofluorine enrichment cultures

These genome announcement narratives are associated with the manuscript "Genome sequences of four bacterial strains isolated from organofluorine enrichment cultures" by Jennifer L. Goff, Chris Hahn, Rebecca A. Ingrassia, Chanistha Tiyapun, Gray G. Waldschmidt, Emma R. Smith, Miranda N. Marini, Olga Shevchenko, and Julia A. Maresca Assembeled genome sequences for all four genomes are available on this page and in the individual genome announcement narratives listed below. All methodological details can be found within the individual narratives.

Goff, Jennifer↗

KBase Narrative - Complete genome sequence of a novel Microbacterium sp. strain Clip185.

We have isolated a new species of Microbacterium, an Actinobacterium. We have temporarily named this bacterium as Microbacterium sp. strain Clip185 (hereafter called strain Clip185) from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii strain LMJ.RY0402.185141 (a Chlamydomonas Library project CLiP strain). We sequenced the whole genome of strain Clip185 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of this new Microbacterium species that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

Mitra, Mautusi↗

Complete genome sequence of Sphingobium yanoikuyae strain CC4533

We have isolated a new strain of Sphingobium yanoikuyae , which belongs to the class Alphaproteobacteria, order Sphingomonadales, and family Sphingomonadaceae. This carotenoid-producing strain is capable of degrading xenobiotics and is tolerant to toxic levels of six heavy metals. We have designated the newly isolated strain of S. yanoikuyae as S. yanoikuyae strain CC4533 (hereafter called strain CC4533) because it was isolated from a contaminated Tris-Acetate-Phosphate (TAP) medium culture plate of a green micro-alga Chlamydomonas reinhardtii wild type strain CC4533. We sequenced the whole genome of strain CC4533 using the PacBio Sequel II Continuous Long Read technology and have submitted it to NCBI along with the SRA and PacBio methylation motif data. Additionally, we have submitted the PacBio methylome to REBASE, Ref#35996. We present the whole genome sequence of S. yanoikuyae strain CC4533 that offers insights into its coding and non-coding genes and its nearest taxonomic neighbors.

59 BASIC BIOLOGICAL SCIENCES↗

Next-generation sequencing dataset of genome-scale CRISPRi in Synechococcus sp. PCC 7002 across seven conditions

A 33,298-member sgRNA library developed for Synechococus sp. PCC 7002 was screened with two replicates across seven growth conditions and sequenced with Illumina NextSeq (paired end, 2x150 bp) for a total of ~700M reads. The original plasmid library and the library after transformation into a dCas9-containing and dCas9-absent strain were also sequenced as a reference for initial sgRNA abundance.

genome- wide screens environmental acclimation spe↗

Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti.

The genus Aspergillus is diverse, including species of industrial importance, human pathogens, plant pests, and model organisms. Aspergillus includes species from sections Usti and Cavernicolus, which until recently were joined in section Usti, but have now been proposed to be non-monophyletic and were split by section Nidulantes, Aenei and Raperi. To learn more about these sections, we have sequenced the genomes of 13 Aspergillus species from section Cavernicolus (A. cavernicola, A. californicus, and A. egyptiacus), section Usti (A. carlsbadensis, A. germanicus, A. granulosus, A. heterothallicus, A. insuetus, A. keveii, A. lucknowensis, A. pseudodeflectus and A. pseudoustus), and section Nidulantes (A. quadrilineatus, previously A. tetrazonus). We compared these genomes with 16 additional species from Aspergillus to explore their genetic diversity, based on their genome content, repeat-induced point mutations (RIPs), transposable elements, carbohydrate-active enzyme (CAZyme) profile, growth on plant polysaccharides, and secondary metabolite gene clusters (SMGCs). All analyses support the split of section Usti and provide additional insights: Analyses of genes found only in single species show that these constitute genes which appear to be involved in adaptation to new carbon sources, regulation to fit new niches, and bioactive compounds for competitive advantages, suggesting that these support species differentiation in Aspergillus species. Sections Usti and Cavernicolus have mainly unique SMGCs. Section Usti contains very large and information-rich genomes, an expansion partially driven by CAZymes, as section Usti contains the most CAZyme-rich species seen in genus Aspergillus. Section Usti is clearly an underutilized source of plant biomass degraders and shows great potential as industrial enzyme producers. Citation: Nybo JL, Vesth TC, Theobald S, Frisvad JC, Larsen TO, Kjaerboelling I, Rothschild-Mancinelli K, Lyhne EK, Barry K, Clum A, Yoshinaga Y, Ledsgaard L, Daum C, Lipzen A, Kuo A, Riley R, Mondo S, LaButti K, Haridas S, Pangalinan J, Salamov AA, Simmons BA, Magnuson JK, Chen J, Drula E, Henrissat B, Wiebenga A, Lubbers RJM, Müller A, dos Santos Gomes AC, Mäkelä MR, Stajich JE, Grigoriev IV, Mortensen UH, de Vries RP, Baker SE, Andersen MR (2025). Section-level genome sequencing and comparative genomics of Aspergillus sections Cavernicolus and Usti. Studies in Mycology 111: 101-114. doi: 10.3114/sim.2025.111.03.

59 BASIC BIOLOGICAL SCIENCES↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Long-term effects of cycle time and volume exchange ratio on poly(3-hydroxybutyrate-co-3-hydroxyvalerate) production from food waste digestate by Haloferax mediterranei cultivated in sequencing batch reactors for 450 days

Food waste digestate was fed into a sequencing batch reactor (SBR) for Haloferax mediterranei (HM) to produce poly(3-hydroxybutyrate-co-3-hydroxyvalerate) (PHBV). This SBR was operated uninterruptedly for 450 days to test its stability, during which the cycle time and volume exchange ratio were varied to understand their impacts on the PHBV fermentation performance under ranged organic loading rates (OLR). Results showed that 1) PHBV productivity was proportional to OLR of food waste digestate; 2) substrate and product inhibitions were two limiting factors constraining substrate utilization and PHBV yields; 3) PHBV titer was dependent on the hydraulic retention time of the SBR while a volume exchange ratio lower than 0.5 is unfavorable due to the product inhibitor accumulation. Furthermore, this study for the first time demonstrated that the long-term stability of food waste-fed PHBV production by HM and revealed that inhibition effects could be barriers in SBR limiting the full-scale application of the technology.

09 BIOMASS FUELS↗

Single cell RNA sequencing reveals shifts in cell maturity and function of endogenous and infiltrating cell types in response to acute intervertebral disc injury

Intervertebral disc (IVD) degeneration contributes to disabling back pain. Degeneration can be initiated by injury and progressively leads to an irreversible loss of cells and function. IVD function restoration through cell replacement therapies have had limited success due to knowledge gaps in the critical cell populations important for repair. Here, in this study, we used single cell RNA sequencing to identify the transcriptional changes of IVD resident and infiltrating cell populations from Control and Injured coccygeal IVDs extracted from 12-week-old female C57BL/6J mice 7 days post injury. Clustering, gene ontology, and pseudotime trajectory analyses determined transcriptomic divergences with injury, flow cytometry identified they types of infiltrating immune cells, and immunofluorescence was utilized to define mesenchymal stem cell (MSC) localization. We identified 11 distinct clusters that included IVD, immune, vascular cells, and MSCs. Differential gene expression analysis determined that Outer Annulus Fibrosus, Neutrophils, Saa2-High MSCs, Macrophages, and Krt18 + Nucleus Pulposus (NP) cells were the major drivers of transcriptomic differences between Control and Injured cells. Gene ontology revealed that the most upregulated biological pathways were angiogenesis and T cell-related while wound healing and ECM regulation were downregulated. Pseudotime trajectory analyses revealed that IVD injury directed cells towards increased differentiation in all clusters, except for Krt18 + NP cells which remained in a less mature cell state. Saa2-High and Grem1-High MSCs populations shifted towards more differentiated IVD cells profiles with injury and localized distinctly within the IVD. This study revealed novel MSC populations with the potential to be leveraged for future IVD repair studies.

Cartilage↗

Frameshifting Stimulatory Sequence Induces Large Structural Change of Ribosomal Proteins When Bound to E. coli Ribosomes

Biological macromolecular machines occupy a continuum of structural conformations to perform cellular tasks. Mapping this conformational space provides an insight into its functionality. While the cryo-electron microscopy resolution revolution has expanded our ability to characterize the conformational continuums, there are obstacles in structurally characterizing regions of high flexibility. These technical barriers have impeded characterization of flexible ribosomal proteins when the ribosome is interacting with mRNA stem-loop structures such as a frameshifting stimulatory sequence (FSS). Small-angle neutron/X-ray scattering and electron microscopy were used to study ribosomal samples and compared structural differences between a ribosome that is bound to an FSS stem-loop compared to a ribosome bound to linear mRNA. This comparison shows that a large protein stalk elongates by 22% when the 70S interacts with an mRNA stem-loop. Finally, our results suggest that ribosomal proteins have extensive flexibility and may influence important ribosomal mechanisms, such as those that involve FSS.

36 MATERIALS SCIENCE↗

Optimized Substrate Positioning Enables Switches in the C–H Cleavage Site and Reaction Outcome in the Hydroxylation–Epoxidation Sequence Catalyzed by Hyoscyamine 6β-Hydroxylase

Hyoscyamine 6β-hydroxylase (H6H) is an Fe(II)- and 2-oxoglutarate-dependent (Fe/2OG) oxygenase that catalyzes the last two steps in the biosynthesis of scopolamine, a prolifically administered anti-nausea drug. After its namesake first reaction, H6H couples the newly installed C6-bonded oxygen to C7 to form the epoxide of scopolamine. Oxoiron(IV) (ferryl) intermediates initiate both reactions by cleaving C–H bonds, but it remains unclear how the enzyme switches target site and promotes (C6)O–C7 coupling in preference to C7 hydroxylation in the second step. In one possible epoxidation mechanism, the C6 oxygen would – analogously to mechanisms proposed for the Fe/2OG halogenases and, in the preceding paper, N-acetylnorloline synthase (LolO) – coordinate as alkoxide to the C7–H-cleaving ferryl intermediate to enable alkoxyl coupling to the ensuing C7 radical. Here we provide structural and kinetic evidence that H6H instead exploits the distinct spatial dependencies of competitive C–H-cleavage (C6 vs C7) and C–O-coupling (oxygen rebound vs cyclization) steps to promote the two-step sequence without substrate coordination or repositioning for the epoxidation step. Structural comparisons of ferryl-mimicking vanadyl complexes of wild-type H6H and a variant that preferentially hydroxylates C7 of 6-hydroxyhyoscyamine suggest that only a modest (~ 10°) shift in the Fe–O–H(C7) approach angle is sufficient to change the outcome. Finally, the observation that, in wild-type H6H, 2 H 2 O solvent also increases the C7-hydroxylation:epoxidation ratio by ~ 8-fold implies that the latter outcome requires cleavage of the alcohol O-H bond, which, unlike in the LolO oxacyclization, is not accomplished in advance of C–H cleavage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Seismic Features Predict Ground Motions During Repeating Caldera Collapse Sequence

Abstract Applying machine learning to continuous acoustic emissions, signals previously deemed noise, from laboratory faults and slowly slipping subduction‐zone faults, demonstrates hidden signatures are emitted that describe physical details, including fault displacement and friction. However, no evidence currently exists to demonstrate that similar hidden signals occur during seismogenic stick‐slip on earthquake faults—the damaging earthquakes of most societal interest. We show that continuous seismic emissions emitted during the 2018 multi‐month caldera collapse sequence at the Kı̄lauea volcano in Hawai'i contain hidden signatures characterizing the earthquake cycle. Multi‐spectral data features extracted from 30 s intervals of the continuous seismic emission are used to train a gradient boosted tree regression model to predict the GNSS‐derived contemporaneous surface displacement and time‐to‐failure of the upcoming collapse event. This striking result suggests that at least some faults emit such signals and provide a potential path to characterizing the instantaneous and future behavior of earthquake faults.

58 GEOSCIENCES↗