Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequence Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Microbiome Comparison and Pathogen Identification for Three Migrating Passerines Captured During Spring Season in Jordan Using 16S rRNA Sequencing

Jordan is located on an important spot along the Mediterranean and Black Sea Flyway. Hundreds of migratory bird species have been identified stopping over in Jordan during spring and autumn migratory seasons. Compared to mammals and economically important birds, the microbiomes of wild bird species are severely understudied. Gut microbial composition is a valuable source of information that reflects food preferences, foraging behavior, and the risk of pathogen transmission to humans and other animals. In this study, we assessed the microbiome composition of three species of migrating passerines (willow warblers, lesser whitethroats, and common reed warblers) captured during the spring migration stopover in Jordan in 2023. A total of 59 fecal samples were selected evenly from the three species and subjected to 16S sequencing and microbiome analysis. Our objectives were to determine the diversity of bacteria in these three species, assess the amount of intra- and inter-specific variation, and detect pathogenic genera and species that could pose health risks to humans, domestic animals, and wildlife. Bacteria mainly belonged to the phyla Proteobacteria (62%), Actinobacteriota (18%), Firmicutes (13%), Cyanobacteria (5%), and Bacteroidota (1%). The results reveal that lesser whitethroats had the greatest variation in bacterial genus richness, Shannon diversity, and microbial composition compared to willow warblers and common reed warblers. The three bird species harbored several pathogenic genera and species, including Campylobacter, Enterococcus, Escherichia-Shigella, Mycoplasma, Rickettsia, Clostridium perfringens, and Vibrio cholerae. We suggest further investigation to understand the relationship between migratory behavior and their gut microbiome. We advocate for the use of advanced molecular techniques to characterize the pathogens found in migratory birds that might have public and environmental health impacts in addition to economic loss.

59 BASIC BIOLOGICAL SCIENCES↗

Genomes OnLine Database (GOLD) v.10: new features and updates

The Genomes OnLine Database (GOLD; https://gold.jgi.doe.gov/) at the Department of Energy Joint Genome Institute is a comprehensive online metadata repository designed to catalog and manage information related to (meta)genomic sequence projects. GOLD provides a centralized platform where researchers can access a wide array of metadata from its four organization levels namely Study, Organism/Biosample, Sequencing Project and Analysis Project. GOLD continues to serve as a valuable resource and has seen significant growth and expansion since its inception in 1997. With its expanded role as a collaborative platform, it not only actively imports data from other primary repositories like National Center for Biotechnology Information but also supports contributions from researchers worldwide. This collaborative approach has enriched the database with diverse datasets, creating a more integrated resource to enhance scientific insights. As genomic research becomes increasingly integral to various scientific disciplines, more researchers and institutions are turning to GOLD for their metadata needs. To meet this growing demand, GOLD has expanded by adding diverse metadata fields, intuitive features, advanced search capabilities and enhanced data visualization tools, making it easier for users to find and interpret relevant information. This manuscript provides an update and highlights the new features introduced over the last 2 years.

59 BASIC BIOLOGICAL SCIENCES↗

Unique Structural Features Relate to Evolutionary Adaptation of Cytochrome P450 in the Abyssal Zone

Cytochromes P450 (CYPs) form one of the largest enzyme superfamilies, with similar structural folds yet biological functions varying from synthesis of physiologically essential compounds to metabolism of myriad xenobiotics. Sterol 14α-demethylases (CYP51s) represent a very special P450 family, regarded as a possible evolutionary progenitor for all currently existing P450s. In metazoans CYP51 is critical for the biosynthesis of sterols including cholesterol. Here we determined the crystal structures of ligand-free CYP51s from the abyssal fish Coryphaenoides armatus and human-. Comparative sequence–structure–function analysis revealed specific structural elements that imply elevated conformational flexibility, uncovering a molecular basis for faster catalytic rates, lower substrate selectivity, and intrinsic resistance to inhibition. In addition, the C. armatus structure displayed a large-scale repositioning of structural segments that, in vivo, are immersed in the endoplasmic reticulum membrane and border the substrate entrance (the FG arm, >20 Å, and the β4 hairpin, >15 Å). The structural distinction of C. armatus CYP51, which is the first structurally characterized deep sea P450, suggests stronger involvement of the membrane environment in regulation of the enzyme function. We interpret this as a co-adaptation of the membrane protein structure with membrane lipid composition during evolutionary incursion to life in the deep sea.

Biochemistry & Molecular Biology↗

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Visualizing metagenomic and metatranscriptomic data: A comprehensive review

The fields of Metagenomics and Metatranscriptomics involve the examination of complete nucleotide sequences, gene identification, and analysis of potential biological functions within diverse organisms or environmental samples. Despite the vast opportunities for discovery in metagenomics, the sheer volume and complexity of sequence data often present challenges in processing analysis and visualization. This article highlights the critical role of advanced visualization tools in enabling effective exploration, querying, and analysis of these complex datasets. Emphasizing the importance of accessibility, the article categorizes various visualizers based on their intended applications and highlights their utility in empowering bioinformaticians and non-bioinformaticians to interpret and derive insights from meta-omics data effectively.

59 BASIC BIOLOGICAL SCIENCES↗

Host analysis-guided selection and targeted engineering (HASTE) of Lipomyces tetrasporus for the conversion of CO2-derived feedstocks

Efficient and cost-competitive bioproduction calls for utilizing CO2-derived feedstocks, such as products from electro-reduction of CO2 and hydrolysate from lignocellulosic biomass. However, efficiently using all their carbon components, including acetate, glucose, and xylose, remains a challenge. Here, we characterize Lipomyces tetrasporus, a novel, robust yeast strain capable of effectively assimilating these carbon sources. We used an integrated systems biology approach combining ¹³C metabolic flux analysis, dynamic labeling experiments, and RNA sequencing. We conducted the first metabolic flux analysis for glucose, xylose, and acetate catabolism in this species. Dynamic labeling revealed a highly active TCA cycle during acetate metabolism, evidenced by rapid citrate and malate accumulation. The strain demonstrated strong NADH/NADPH production and acetyl-CoA synthase activity. Using insights and gene targets from this analysis, we engineered L. tetrasporus for malate production. The engineered strain produced 7.5 g/L malic acid (0.25 g/g yield) in shake flasks with glucose-acetate media and 28.8 g/L malic acid at a yield of 0.20 g/g in fed-batch mode with corn-stover hydrolysate. Together, these insights and rational strain engineering establish L. tetrasporus as a versatile, Crabtree-negative platform that is an energy-CO2-bioproduction nexus for channeling CO2 carbon into value-added bioproducts.

Xiao, Zhengyang↗

Genomic factors shaping codon usage across the Saccharomycotina subphylum

Codon usage bias, or the unequal use of synonymous codons, is observed across genes, genomes, and between species. It has been implicated in many cellular functions, such as translation dynamics and transcript stability, but can also be shaped by neutral forces. We characterized codon usage across 1,154 strains from 1,051 species from the fungal subphylum Saccharomycotina to gain insight into the biases, molecular mechanisms, evolution, and genomic features contributing to codon usage patterns. We found a general preference for A/T-ending codons and correlations between codon usage bias, GC content, and tRNA-ome size. Codon usage bias is distinct between the 12 orders to such a degree that yeasts can be classified with an accuracy >90% using a machine learning algorithm. We also characterized the degree to which codon usage bias is impacted by translational selection. We found it was influenced by a combination of features, including the number of coding sequences, BUSCO count, and genome length. Our analysis also revealed an extreme bias in codon usage in the Saccharomycodales associated with a lack of predicted arginine tRNAs that decode CGN codons, leaving only the AGN codons to encode arginine. Analysis of Saccharomycodales gene expression, tRNA sequences, and codon evolution suggests that avoidance of the CGN codons is associated with a decline in arginine tRNA function. Consistent with previous findings, codon usage bias within the Saccharomycotina is shaped by genomic features and GC bias. However, we find cases of extreme codon usage preference and avoidance along yeast lineages, suggesting additional forces may be shaping the evolution of specific codons.

59 BASIC BIOLOGICAL SCIENCES↗

Hybridization capture sequencing for Vibrio spp. and associated virulence factors

ABSTRACT Proliferation ofVibriospp. in aquatic ecosystems is associated with climate change and, concomitantly, increased incidence of vibriosis. They are autochthonous to aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing (HCS) was employed to profile low-abundanceVibriospp. in environmental samples. The HCS panel targeted a family of molecular chaperones (CPN60) specific to 69Vibriospp. and 162Vibrio-specific virulence factors. This approach was evaluated in parallel with traditional whole-community shotgun sequencing in a metagenomic analysis of water and oyster samples collected from the Chesapeake Bay. In addition,Vibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples were subjected to whole-genome sequencing to determine the genetic characteristics of pathogenicVibriospp. circulating in an aquatic environment. HCS, employed to determine the incidence and characterization of specificVibriospp., yielded significantly greater metagenomic insight, notably a variety of otherVibriospp., including detection ofVibrio cholerae,Vibrio fluvialis, andVibrio aestuarianus, in addition toVibrio parahaemolyticusandVibrio vulnificus, and also important virulence factors not detectable using traditional molecular methods. Thus, pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood. It is concluded that environmental surveillance should include HCS, a valuable tool for the detection and characterization of pathogenic agents in aquatic ecosystems, notably vibrios. IMPORTANCE The increasing prevalence of pathogenicVibriospp. in aquatic ecosystems, driven by climate change, is closely linked to a rise in cholera and vibriosis cases, emphasizing the need for improved environmental surveillance. Vibrios are naturally occurring in aquatic environments globally, but traditional metagenomic methods for detecting and typing pathogenicVibriospp. are challenged by their presence in relatively low abundance and ability to persist in a viable but nonculturable state. In the study reported here, hybridization capture sequencing was employed to profile low-abundanceVibriospp. in metagenomic samples, namely water and oysters collected from the Chesapeake Bay. This approach was evaluated in parallel with traditional whole-community shotgun sequencing and whole-genome sequencing ofVibrio parahaemolyticusandVibrio vulnificusstrains isolated from the samples. Results suggest pathogenicVibriospp. in aquatic ecosystems may be far more common than currently understood, when multiple methods are considered for environmental surveillance.

Microbiology↗

Metagenome-assembled genomes provide insight into the metabolic potential during early production of Hydraulic Fracturing Test Site 2 in the Delaware Basin

Demand for natural gas continues to climb in the United States, having reached a record monthly high of 104.9 billion cubic feet per day (Bcf/d) in November 2023. Hydraulic fracturing, a technique used to extract natural gas and oil from deep underground reservoirs, involves injecting large volumes of fluid, proppant, and chemical additives into shale units. This is followed by a “shut-in” period, during which the fracture fluid remains pressurized in the well for several weeks. The microbial processes that occur within the reservoir during this shut-in period are not well understood; yet, these reactions may significantly impact the structural integrity and overall recovery of oil and gas from the well. To shed light on this critical phase, we conducted an analysis of both pre-shut-in material alongside production fluid collected throughout the initial production phase at the Hydraulic Fracturing Test Site 2 (HFTS 2) located in the prolific Wolfcamp formation within the Permian Delaware Basin of west Texas, USA. Specifically, we aimed to assess the microbial ecology and functional potential of the microbial community during this crucial time frame. Prior analysis of 16S rRNA sequencing data through the first 35 days of production revealed a strong selection for a Clostridia species corresponding to a significant decrease in microbial diversity. Here, we performed a metagenomic analysis of produced water sampled on Day 33 of production. This analysis yielded three high-quality metagenome-assembled genomes (MAGs), one of which was a Clostridia draft genome closely related to the recently classified Petromonas tenebris. This draft genome likely represents the dominant Clostridia species observed in our 16S rRNA profile. Annotation of the MAGs revealed the presence of genes involved in critical metabolic processes, including thiosulfate reduction, mixed acid fermentation, and biofilm formation. These findings suggest that this microbial community has the potential to contribute to well souring, biocorrosion, and biofouling within the reservoir. Our research provides unique insights into the early stages of production in one of the most prolific unconventional plays in the United States, with important implications for well management and energy recovery.

natural gas↗

A Chemoselective and Stereodivergent Platform of Heme‐Nitrene Transferases to Access Chiral Aryl‐β‐Amino Esters and An Investigation of the Sequence‐Activity Landscape

Engineered biocatalysts can utilize nitrene precursors to access enantioenriched amination products, yet they have not been applied to produce valuable, enantiomerically enriched noncanonical β-amino esters. Current approaches to synthesizing β-amino acids rely on pre-oxidized precursors and multistep synthetic approaches involving various protecting groups. We engineered a platform of heme enzymes for stereoselective C–H bond amination of readily available carboxylic ester derivatives to install primary amines. A directed evolution campaign coupled with sequencing of over 1000 variants enabled us to develop engineered variants that use either O-pivaloylhydroxylamine triflic acid (PONT) or hydroxylamine hydrochloride (H 2 NOH∙HCl) as aminating reagents. An analysis of the resulting sequence–activity dataset revealed additional improvements that could be made to the final variant, highlighting the utility of sequencing data to guide future steps in directed evolution campaigns. Furthermore, the evolved nitrene transferases expand the scope of accessible chiral β-amino acid building blocks for peptidomimetic applications and provide new starting points for the design and synthesis of enantioenriched β-amino acid motifs.

amino ester building blocks↗

Rhizosphere Microbiome Diversity Potentially Supports Robust Nature of Field Pennycress ( Thlaspi arvense L.) in Dryland Cropping Systems of Eastern Washington

ABSTRACT Field pennycress ( Thlaspi arvense L.) is an annual in the Brassicaceae family and is currently being developed as an oilseed intermediate crop suitable for renewable biodiesel and jet fuel. It displays many desirable characteristics for this role including cold tolerance, a rapid life cycle, and a seed fatty acid profile conducive to bioenergy generation. These traits make field pennycress favorable for winter oilseed cultivation in the inland Pacific Northwest (iPNW). Simultaneously, intermediate crops are an increasingly recognized component of both agronomic sustainability and soil health management. Intermediate crops enhance soil microbial diversity, which benefits both soil and plant health. To understand the impact of field pennycress on soil microbial diversity, two natural accessions and seven experimental accessions were grown at three sites in Eastern Washington. Aboveground biomass and rhizosphere soil were then collected. Soil genomic DNA was extracted from rhizosphere samples and used to generate an amplicon library for bacterial (16S) and fungal (ITS) rRNA sequences. The resulting libraries were analyzed in QIIME2, which revealed that not only did the fad2 deficient line from the Spring32‐10 background have significantly increased aboveground biomass production compared to other pennycress genotypes, but also displayed significantly higher β‐diversity in the rhizosphere community specifically at the site experiencing the driest conditions. ANCOM analysis showed that multiple sequences similar to beneficial plant and soil health enhancing organisms such as Trichoderma spirale , Pseudomonas spp., and Methylobacterium goesingense were found to be enriched in the microbiome of the fad2 Spring32‐10 background also at that site. To add additional context to rhizosphere community data, root exudates from two pennycress genotypes were captured in magenta boxes and analyzed using HPLC. Future work will expand our understanding of the mechanisms by which field pennycress creates diversity in the rhizosphere, thus expanding our ability to cultivate this crop in the iPNW.

54 ENVIRONMENTAL SCIENCES↗

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

High-Power Clock Laser Spectrally Tailored for High-Fidelity Quantum State Engineering

Highly frequency-stable lasers are ubiquitous tools for optical-frequency metrology, precision interferometry, and quantum information science. While making a universally applicable laser is unrealistic, spectral noise can be tailored for specific applications. Here we report a high-power 698-nm clock laser with a maximum output of 4W and minimized frequency noise up to a few kHz Fourier frequency, together with long-term instability of 3.5 × 10 −17 at one to thousands of seconds. The laser-frequency noise is precisely characterized with atom-based spectral analysis that employs a pulse sequence designed to suppress sensitivity to intensity noise. This method provides universally applicable tunability of the spectral response and analysis of quantum sensors over a wide frequency range. With the optimized laser system characterized by this technique, we achieve an average single-qubit Clifford gate fidelity of up to 𝐹$^2_1$ = 0.999⁢64⁢(3) when simultaneously driving 3000 optical qubits with a homogeneous Rabi frequency ranging from 10 Hz to 1 kHz. This result represents the highest single optical-qubit-gate fidelity for a large number of atoms.

atomic gases↗

Pressure–Temperature–Magnetic Field Phase Diagram of Multiferroic (NH 4 ) 2 FeCl 5 ·H 2 O

We combined synchrotron-based infrared absorbance and Raman scattering spectroscopies with diamond anvil cell techniques and a symmetry analysis to explore the properties of multiferroic (NH 4 ) 2 FeCl 5 ·H 2 O under extreme pressure–temperature conditions. Compression-induced splitting of the Fe–Cl stretching, Cl–Fe–Cl and Cl–Fe–O bending, and NH 4 + librational modes defines two structural phase transitions, and a group–subgroup analysis reveals space group sequences that vary depending upon proximity to the unexpectedly wide order–disorder transition. Here, we bring these findings together with prior high-field work to develop the pressure–temperature–magnetic field phase diagram uncovering competing polar, chiral, and magnetic phases in this system.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES↗

Deficiency in transmitter release triggers homeostatic transcriptional changes that increase presynaptic excitability

Weakening of synaptic transmission at theDrosophilalarval neuromuscular junction triggers two forms of homeostatic compensation, one that increases the probability of glutamate release per action potential (P r ) and another that increases motoneuron (MN) activity. We investigated the molecular changes in MNs that underlie the increase in MN activity. RNA sequencing (RNA-seq) analysis on MNs whose glutamate release is weakened by knockdown of components of the MN transmitter release machinery reveals a reduction in expression of a group of genes that encode potassium channels and their positive modulators. These results identify a mechanism of compensation for weakened synaptic transmission by MNs, which engages a transcriptional program in those cells to increase firing and, thereby, ensure sufficient locomotory drive.

Science & Technology - Other Topics↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗