Search NASASearch

SEARCH · Search NASA

Results for “Sequence Alignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Standardized Residue Numbering and Secondary Structure Nomenclature in the Class D β-Lactamases

Over 1370 class D β-lactamases are currently known, and they pose a serious threat to the effective treatment of many infectious diseases, particularly in some pathogenic bacteria where evolving carbapenemase activity has been reported. Detailed understanding of their molecular biology, enzymology, and structural biology are critically important, but the lack of a standardized residue numbering scheme and inconsistent secondary structure annotation has made comparative analyses sometimes difficult and cumbersome. Compounding this, in the post-AlphaFold world where we currently find ourselves, an extraordinary wealth of detailed structural information on these enzymes is literally at our fingertips; therefore it is vitally important that a standard numbering system is in place to facilitate the accurate and straightforward analysis of their structures. In conclusion, here we present a residue numbering and secondary structure scheme for the class D enzymes based on the sequence and structure of OXA-48 and apply it to test targets to demonstrate the ease with which it can be used.

59 BASIC BIOLOGICAL SCIENCES

Contact-dependent growth inhibition (CDI) systems deploy a large family of polymorphic ionophoric toxins for inter-bacterial competition

Contact-dependent growth inhibition (CDI) is a widespread form of inter-bacterial competition mediated by CdiA effector proteins. CdiA is presented on the inhibitor cell surface and delivers its toxic C-terminal region (CdiA-CT) into neighboring bacteria upon contact. Inhibitor cells also produce CdiI immunity proteins, which neutralize CdiA-CT toxins to prevent auto-inhibition. Here, we describe a diverse group of CDI ionophore toxins that dissipate the transmembrane potential in target bacteria. These CdiA-CT toxins are composed of two distinct domains based on AlphaFold2 modeling. The C-terminal ionophore domains are all predicted to form five-helix bundles capable of spanning the cell membrane. The N-terminal "entry" domains are variable in structure and appear to hijack different integral membrane proteins to promote toxin assembly into the lipid bilayer. The CDI ionophores deployed by E. coli isolates partition into six major groups based on their entry domain structures. Comparative sequence analyses led to the identification of receptor proteins for ionophore toxins from groups 1 & 3 (AcrB), group 2 (SecY) and groups 4 (YciB). Using forward genetic approaches, we identify novel receptors for the group 5 and 6 ionophores. Group 5 exploits homologous putrescine import proteins encoded by puuP and plaP, and group 6 toxins recognize di/tripeptide transporters encoded by paralogous dtpA and dtpB genes. Finally, we find that the ionophore domains exhibit significant intra-group sequence variation, particularly at positions that are predicted to interact with CdiI. Accordingly, the corresponding immunity proteins are also highly polymorphic, typically sharing only ~30% sequence identity with members of the same group. Competition experiments confirm that the immunity proteins are specific for their cognate ionophores and provide no protection against other toxins from the same group. The specificity of this protein interaction network provides a mechanism for self/nonself discrimination between E. coli isolates.

59 BASIC BIOLOGICAL SCIENCES

CoverM: read alignment statistics for metagenomics

SUMMARY: Genome-centric analysis of metagenomic samples is a powerful method for understanding the function of microbial communities. Calculating read coverage is a central part of analysis, enabling differential coverage binning for recovery of genomes and estimation of microbial community composition. Coverage is determined by processing read alignments to reference sequences of either contigs or genomes. Per-reference coverage is typically calculated in an ad-hoc manner, with each software package providing its own implementation and specific definition of coverage. Here we present a unified software package CoverM which calculates several coverage statistics for contigs and genomes in an ergonomic and flexible manner. It uses "Mosdepth arrays" for computational efficiency and avoids unnecessary I/O overhead by calculating coverage statistics from streamed read alignment results. AVAILABILITY AND IMPLEMENTATION: CoverM is free software available at https://github.com/wwood/coverm. CoverM is implemented in Rust, with Python (https://github.com/apcamargo/pycoverm) and Julia (https://github.com/JuliaBinaryWrappers/CoverM_jll.jl) interfaces.

Aroney, Samuel T N

A Digital Three Level Space Vector Modulator for High Frequency Vector Sequence Generation

This letter proposes a digital high-speed three-level space vector pulse width modulator (3L-SVPWM). A conventional 3L-SVPWM is typically computation-based, involving a sequential execution of sub-tasks on a digital signal processor (DSP) based controller. The resulting high computation time of 5.4 μs limits the implementation of additional control blocks for switching frequencies greater than 100 kHz. This is overcome by transforming sub-tasks into digital blocks with 1-0 decisions and simpler arithmetic operations. The sub-task blocks are executed concurrently on a programmable logic device (PLD). Hence, a fast 3L-SVPWM execution in 140 ns is achieved. The proposed digital 3L-SVPWM enables high switching frequency operation of wide bandgap (WBG) device-based 3 L inverters to generate high fundamental frequency waveforms. A finite state machine is an integral part of the proposed implementation with the ability to generate any vector sequence, maximizing the usage of redundant vector states in 3L-SVPWM. Here, the proposed digital 3L-SVPWM operation is demonstrated with a GaN-based 3 L active neutral point clamped (3L-ANPC) inverter. Experimental results are presented at 250 kHz switching frequency to generate vector sequences for center-aligned SVPWM (CA-SVPWM) and common mode voltage reduced SVPWM (CMVR-SVPWM). The results also showcase a high fundamental frequency generation capability of 10 kHz.

active neutral point clamped inverter

Single-nuclei transcriptome analysis of IgM+ cells isolated from channel catfish (Ictalurus punctatus) spleen

Catfish production is the primary aquaculture sector in the United States, and the key cultured species is channel catfish (Ictalurus punctatus). The major causes of production losses are pathogenic diseases, and the spleen, an important site of adaptive immunity, is implicated in these diseases. To examine the channel catfish immune system, single-nuclei transcriptomes of sorted and captured IgM + cells were produced from adult channel catfish. Three channel catfish (~1 kg) were euthanized, the spleen dissected, and the tissue dissociated. The lymphocytes were isolated using a Ficoll gradient and IgM + cells were then sorted with flow cytometry. The IgM + cells were lysed and single-nuclei libraries generated using a Chromium Next GEM Single Cell 3’ GEM Kit and the Chromium X Instrument (10x Genomics) and sequenced with the Illumina NovaSeq X Plus sequencer. The reads were aligned to theI. punctatusreference assembly (Coco_2.0) using Cell Ranger, and normalization, cluster analysis, and differential gene expression analysis were carried out with Seurat. Across the three samples, approximately 753.5 million reads were generated for 18,686 cells. After filtering, 10,637 cells remained for the cluster analysis. The cluster analysis identified 16 clusters which were classified as B cells (10,276), natural killer-like (NK-like) cells (178), T cells or natural killer cells (45), hematopoietic stem and progenitor cells (HSPC)/megakaryocytes (MK) (66), myeloid/epithelial cells (40), and plasma cells (32). The B cell clusters were further defined as different populations of mature B cells, cycling B cells, and plasma cells. The plasma cells highly expressedighmand we demonstrated that the secreted form of the transcript was largely being expressed by these cells. This atlas provides insight into the gene expression of IgM + immune cells in channel catfish. The atlas is publicly available and could be used garner more important information regarding the gene expression of splenic immune cells.

Immunology

Single-nuclei transcriptome analysis of channel catfish spleen provides insight into the immunome of an aquaculture-relevant species

The catfish industry is the largest sector of U.S. aquaculture production. Given its role in food production, the catfish immune response to industry-relevant pathogens has been extensively studied and has provided crucial information on innate and adaptive immune function during disease progression. To further examine the channel catfish immune system, we performed single-cell RNA sequencing on nuclei isolated from whole spleens, a major lymphoid organ in teleost fish. Libraries were prepared using the 10X Genomics Chromium X with the Next GEM Single Cell 3’ reagents and sequenced on an Illumina sequencer. Each demultiplexed sample was aligned to the Coco_2.0 channel catfish reference assembly, filtered, and counted to generate feature-barcode matrices. From whole spleen samples, outputs were analyzed both individually and as an integrated dataset. The three splenic transcriptome libraries generated an average of 278,717,872 reads from a mean 8,157 cells. The integrated data included 19,613 cells, counts for 20,121 genes, with a median 665 genes/cell. Cluster analysis of all cells identified 17 clusters which were classified as erythroid, hematopoietic stem cells, B cells, T cells, myeloid cells, and endothelial cells. Subcluster analysis was carried out on the immune cell populations. Here, distinct subclusters such as immature B cells, mature B cells, plasma cells, γδ T cells, dendritic cells, and macrophages were further identified. Differential gene expression analyses allowed for the identification of the most highly expressed genes for each cluster and subcluster. This dataset is a rich cellular gene expression resource for investigation of the channel catfish and teleost splenic immunome.

Science & Technology - Other Topics

Primordial Black Hole Triggered Type Ia Supernovae. I. Impact on Explosion Dynamics and Light Curves

Primordial black holes (PBHs) in the asteroid-mass window are compelling dark matter candidates, made plausible by the existence of black holes and by the variety of mechanisms of their production in the early Universe. If a PBH falls into a white dwarf (WD), the strong tidal forces can generate enough heat to trigger a thermonuclear runaway explosion, depending on the WD’s mass and the PBH’s orbital parameters. In this work, we investigate the WD explosion triggered by the passage of a PBH. We perform 2D simulations of the WD undergoing thermonuclear explosion in this scenario, with the predicted ignition site as a parameter assuming the deflagration–detonation transition model. We study the explosion dynamics, and predict the associated light curves and nucleosynthesis. We find that the model sequence predicts light curves which align with the Phillips relation (B max versus ΔM 15 ). Our models hint at a unifying approach in triggering Type Ia supernovae without involving two distinctive evolutionary tracks.

Dark matter

A Systematic Framework for Tuning Open-Source Multifunctional IBR Models To Emulate OEM Black-Box Fault Dynamics

This paper presents a systematic framework to tune a generic IBR EMT model to match with an OEM provided balckbox inverter model based on the fault current responses. The key learnings and findings are summarized as follows: The tunable key parameters include inner control loops and current limiters to align the fault current magnitude, sequence content, and phase trajectories with the OEM models across diverse fault type and locations. The tuned model's fidelity is validated through comparative analysis with an OEM blackbox model, assessing both the fault current response and the responses of multiple relay elements. The results demonstrate the tuned generic model can trigger relay decision logic that is identical or near identical to that of the OEM model, thus generating very good match model for fault studies.

24 POWER TRANSMISSION AND DISTRIBUTION

Canted antiferromagnetism and spin reorientation in corner-shared single chain quasi-one-dimensional Ba 2 ⁢FeSe 3

Here, we report the canted antiferromagnetic (AFM) structure together with a spin reorientation in a single chain quasi-one-dimensional (Q-1D) iron chalcogenide Ba 2⁢ FeSe 3 . Ba 2 ⁢FeSe 3 crystallizes in Pnma (No. 62) orthorhombic structure with linear single iron chains consisting of corner-shared distorted FeSe 4 tetrahedra along the 𝑏 axis. Ba 2 ⁢FeSe 3 is a narrow-gap semiconductor and orders AFM below 60 K. Modeling of neutron powder diffraction data reveals a canted AFM ground state of magnetic space group 𝑃⁢𝑎⁢21/𝑐 (BNS No. 14.80) with commensurate propagation vector 𝐤 =(0, $\frac{1}{2}$, 0), where the Fe ion spins are AFM aligned with up-down-up-down (↑−↓−↑−↓) sequence along the Q-1D chain direction of the 𝑏 axis. In the magnetically ordered state, the canting of magnetic moments reorients from the 𝑎⁢𝑐 plane to the 𝑎⁢𝑏 plane below 30 K, with a 10° tilting angle toward the 𝑎 axis, and the magnetic moment does not induce a net moment in either orientation. The density functional theory results indicate that an ↑−↓−↑−↓ AFM state is stabilized along the chain direction. In this work, we elucidate the unique canted AFM of the iron chalcogenide and pave the way for searching exotic physics in Q-1D Ba 2⁢ FeSe 3 .

Gao, Fei [Univ. of Texas at Dallas, Richardson, TX

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS

Sensitive and error-tolerant annotation of protein-coding DNA with BATH

We present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.

59 BASIC BIOLOGICAL SCIENCES

Design and calculations for HWR Strongback

The PIP-II Superconductive Linac is a linear accelerator that provides the first stage of beam acceleration. The velocity increase is achieved through a sequence of different cryomodules, where all the cavities are aligned to form a continuous string. The cryomodule team is currently developing a spare for the HWR cryomodule. This spare will not be an exact replica of the existing one, but rather will be based on the SSR f650 cryomodule design created for PIP-II . Each cavity , magnet , and coupler form a unit, with a total of eight units per string. These components, together with additional equipment (such as piping , the thermal shield , and beamline interconnections ), are supported by the Strongback . The aim of this project is to develop a brand-new Strongback design for HWR cavities, and to perform the related analyses and calculations in order to meet the required specifications.

43 PARTICLE ACCELERATORS

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora

Spatially Aligned Binary Single-Site Catalyst on Defective SiO 2 for Cascading Reactions

Capitalizing on the success of single-atom catalysts (SACs), dual-atom catalysts (DACs) have emerged as a new frontier in heterogeneous catalysis. However, most SACs and DACs studies seek to uniformly distribute the catalytic sites on the support material, which can hinder their effectiveness in intricate multistep cascading reactions. Particularly, it is a grand challenge to precisely control the spatial distribution of two different single sites forming binary sites so that reactants and intermediates contact the catalytic sites in the exact sequence required by the reaction steps. Here, in this work, we report a new type of binary single-site catalyst, Cu 1 –Zr 1 @SiO 2 , with Cu 1 and Zr 1 sites spatially aligned with the reaction sequence of the cascade reactions. The catalyst is synthesized by a modified reverse microemulsion approach, with single Cu sites anchored by nonbridging oxygen hole centers, which were induced by doping single Zr sites into SiO 2 . Low-energy ion scattering spectroscopy (LEIS) reveals that the outermost surface of the catalyst contains only Cu single sites, while the Zr sites are dispersed in the bulk. The catalytic performance is demonstrated in ethanol conversion to butenes, a model cascade reaction which includes ethanol dehydrogenation and aldol condensation steps. The precisely spatially controlled binary sites enable ethanol to first undergo dehydrogenation to acetaldehyde on Cu sites, followed by aldol condensation of acetaldehyde on Zr sites. As a result, C 3+ olefins selectivity as high as 77.0% (56.0% selectivity of butenes) is achieved by suppressing ethylene formation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Profiling expression strategies for a type III polyketide synthase in a lysate-based, cell-free system

Abstract Some of the most metabolically diverse species of bacteria (e.g., Actinobacteria) have higher GC content in their DNA, differ substantially in codon usage, and have distinct protein folding environments compared to tractable expression hosts like Escherichia coli . Consequentially, expressing biosynthetic gene clusters (BGCs) from these bacteria in E. coli often results in a myriad of unpredictable issues with regard to protein expression and folding, delaying the biochemical characterization of new natural products. Current strategies to achieve soluble, active expression of these enzymes in tractable hosts can be a lengthy trial-and-error process. Cell-free expression (CFE) has emerged as a valuable expression platform as a testbed for rapid prototyping expression parameters. Here, we use a type III polyketide synthase from Streptomyces griseus , RppA, which catalyzes the formation of the red pigment flaviolin, as a reporter to investigate BGC refactoring techniques. We applied a library of constructs with different combinations of promoters and rppA coding sequences to investigate the synergies between promoter and codon usage. Subsequently, we assess the utility of cell-free systems for prototyping these refactoring tactics prior to their implementation in cells. Overall, codon harmonization improves natural product synthesis more than traditional codon optimization across cell-free and cellular environments. More importantly, the choice of coding sequences and promoters impact protein expression synergistically, which should be considered for future efforts to use CFE for high-yield protein expression. The promoter strategy when applied to RppA was not completely correlated with that observed with GFP, indicating that different promoter strategies should be applied for different proteins. In vivo experiments suggest that there is correlation, but not complete alignment between expressing in cell free and in vivo. Refactoring promoters and/or coding sequences via CFE can be a valuable strategy to rapidly screen for catalytically functional production of enzymes from BCGs, which advances CFE as a tool for natural product research.

59 BASIC BIOLOGICAL SCIENCES

Automation of Laser Plasma Focused Ion Beam Microscopy for Next-Gen Energy Materials

Automation can revolutionize the use of ultrafast laser ablation and plasma-focused ion beam (PFIB) techniques for high-throughput, reproducible cross-sectioning and various sample preparation in materials characterization. As these methods become essential for analyzing complex energy materials and next-generation devices, efficient, standardized workflows are needed to minimize variability and enhance precision. This work highlights our advancements in developing automated processes for sample preparation that integrates machine learning, workflow optimization, and large-scale data acquisition to improve efficiency and scalability in applications such as electrolyzers, photovoltaic cells, and microelectronics. To streamline cross-sectioning and lamella fabrication, we have implemented fully automated workflows that standardize laser ablation and PFIB milling sequences. These workflows incorporate pre-programmed protocols for material removal, alignment, and thinning, reducing user intervention and ensuring consistency across different sample types. Machine learning algorithms further enhance automation by predicting optimal milling strategies and adapting parameters based on material properties and sectioning requirements. This approach significantly improves throughput while maintaining the structural integrity of prepared samples for high-resolution imaging and analysis, including transmission electron microscopy. Beyond sample preparation, our automation platform enables the acquisition of large, high-resolution datasets through serial sectioning, image alignment, and 3D reconstruction. These automated routines facilitate multi-scale characterization, capturing structural and compositional details from the nanoscale to the device level. By reducing variability and increasing efficiency, our automated approach enhances defect analysis, failure diagnostics, and process optimization, accelerating advancements in materials research and device engineering.

36 MATERIALS SCIENCE

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES

Constructing a High‐Resolution Aftershock Catalog for the 2017 Mw 8.2 Tehuantepec Earthquake Sequence Using a Machine Learning–Based Workflow

The 8 September 2017 Mw 8.2 Tehuantepec earthquake was the largest instrumentally recorded normal‐faulting earthquake in Mexico. The mainshock occurred offshore within the Tehuantepec seismic gap, generating >30,000 aftershocks in the following year. We applied an open‐source, machine learning (ML)–assisted workflow to construct a high‐resolution aftershock catalog using data from temporary and permanent seismic networks in southern Mexico. The workflow integrates PhaseNet for phase detection; GaMMA for phase association; and VELEST, HypoInverse, and HypoDD for velocity modeling and relocation. We processed seven months of continuous waveform data from 29 broadband stations, including a temporary rapid‐response deployment that improved station coverage of the offshore rupture zone. To evaluate performance, we compared our results against analyst‐reviewed picks and event locations from the Servicio Sismológico Nacional catalog. The resulting catalog contains 11,374 relocated earthquakes and represents the most comprehensive published dataset for this sequence, incorporating the first full use of the temporary network. Relocated hypocenters show improved depth control and align well with the Slab2.0 subduction geometry, revealing clearer separation between offshore slab events and onshore crustal seismicity. This study demonstrates that combining ML‐based detection with established methods provides a scalable and reproducible approach for constructing high‐quality earthquake catalogs in tectonically complex environments and offers practical guidance for adapting similar workflows to other earthquake sequences.

Garcia, Marc [The University of Texas at El Paso,