Search NASASearch

SEARCH · Search NASA

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Investigating Electron Conductivity Regimes in the Bacterial Cytochrome Wire OmcS

The anaerobic bacterium Geobacter sulfurreducens produces extracellular, electronically conductive cytochrome polymer wires that are conductive over micron length scales. Structure models from cryo-electron microscopy data show OmcS wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. However, the mechanism by which this heme arrangement supports long-range electronic conduction remains unknown. Structure models from cryo-electron microscopy data show these wires form a linear chain of hemes along the protein wire axis, which is proposed as the structural basis supporting their electronic properties. Existing computational models using static heme redox potentials and coupling energies fail to explain experimental observations, predicting conductances 10,000 to 100,000 times lower than measured values. Here, we investigate how dynamic disorder affects site energies, interheme coupling, and long-range electronic conductivity within these cytochrome wires. We introduce an approach to extract charge carrier site information directly from Kohn–Sham density functional theory, without employing projector schemes, and show that site and coupling energies are highly sensitive to changes in interheme geometry and the surrounding electrostatic environment. Unlike models that incorporate dynamic disorder as a thermally averaged quantity, our quantum charge carrier model incorporates proxies for dynamic disorder through decoherence corrections, yielding predicted diffusion coefficient closer to what is expected from experiment and comparable with other organic-based electronic materials. Based on these simulations, we propose that the instantaneous fluctuations of the local electrostatic environment can transiently lift energy degeneracies and delocalize charge carriers. Furthermore, these studies reveal how incorporating dynamic fluctuations associated with the environment resolves the discrepancy between theory and experiment in microbial cytochrome wires and highlight design principles for bioinspired, heme-based conductive materials.

Bioinorganic chemistry

The Exoproteome and Surfaceome of Toxigenic Corynebacterium diphtheriae 1737 and Its Response to Iron Restriction and Growth on Human Hemoglobin

Toxin-producing Corynebacterium diphtheriae strains are the etiological agents of the severe upper respiratory disease, diphtheria. A global phylogenetic analysis revealed that biotype gravis is particularly lethal as it produces diphtheria toxin and a range of other virulence factors, particularly when it encounters low levels of iron at sites of infection. Here, to gain insight into how it colonizes its host, we have identified iron-dependent changes in the exoproteome and surfaceome of C. diphtheriae strain 1737 using a combination of whole-cell fractionation, intact cell surface proteolysis, and quantitative proteomics. In total, we identified 1414 of the predicted 2265 proteins (62%) encoded by its reference genome. For each protein, we quantified its degree of secretion and surface exposure, revealing that exoproteases and hydrolases predominate in the exoproteome, while the surfaceome is enriched with adhesins, particularly DIP2093. Our analysis provides insight into how components in the heme-acquisition system are positioned, showing pronounced surface exposure of the strain-specific ChtA/ChtC paralogues and high secretion of the species-conserved heme-binding HtaA protein, suggesting it functions as a hemophore. Profiling the response of the exoproteome and surfaceome after microbial exposure to human hemoglobin and iron limitation reveals potential virulence factors that may be expressed at sites of infection. Data are available via ProteomeXchange with identifier PXD051674.

cell envelope

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Carbon source–driven metabolic and regulatory remodeling defines phenomic states in Lipomyces starkeyi

Lipomyces is a genus of oleaginous yeasts with potential for contributing to reliable biomanufacturing supply chains. However, progress in advanced strain designs and engineering efforts are still constrained by a lack of understanding of the underlying molecular drivers of Lipomyces phenotypes. To address this gap, we collected a suite of multi-omic data to dissect how carbon source availability reshapes the metabolic network, lipid allocation, and regulatory architecture of Lipomyces starkeyi. We observed that glucose promotes biosynthetic and proliferative processes supported by abundant energy and carbon intermediates, xylose enhances redox-balancing mechanisms centered on the pentose phosphate pathway, and glycerol activates respiratory metabolism, ß-oxidation, and the glyoxylate cycle. Lipid species distributions remained consistent in both nitrogen replete and depleted conditions across the carbon sources, indicating robust production mechanisms. Regulatory protein identification and network analysis revealed glycerol-driven respiratory growth favors regulatory programs integrating stress tolerance, redox balance, and lipid-associated metabolism, whereas xylose growth activates compensatory transcriptional responses aimed at maintaining mitochondrial function. Nitrogen limitation modulates the strength of these responses but does not fundamentally alter their direction, reinforcing carbon source as the dominant driver of regulatory architecture. Taken together, this data enhances the understanding of Lipomyces molecular rearrangements and provides a foundation for further development of predictive phenotypic tools in this genus.

Biotechnology

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES

Coupling Microdroplet-Based Sample Preparation, Multiplexed Isobaric Labeling, and Nanoflow Peptide Fractionation for Deep Proteome Profiling of the Tissue Microenvironment

There is increasing interest in developing in-depth proteomic approaches for mapping tissue heterogeneity in a cell-type-specific manner to better understand and predict the function of complex biological systems such as human organs. Existing spatially resolved proteomics technologies cannot provide deep proteome coverage due to limited sensitivity and poor sample recovery. Herein, we seamlessly combined laser capture microdissection with a low-volume sample processing technology that includes a microfluidic device named microPOTS (microdroplet processing in one pot for trace samples), multiplexed isobaric labeling, and a nanoflow peptide fractionation approach. The integrated workflow allowed us to maximize proteome coverage of laser-isolated tissue samples containing nanogram levels of proteins. We demonstrated that the deep spatial proteomics platform can quantify more than 5000 unique proteins from a small-sized human pancreatic tissue pixel (∼60,000 μm2) and differentiate unique protein abundance patterns in pancreas. Furthermore, the use of the microPOTS chip eliminated the requirement for advanced microfabrication capabilities and specialized nanoliter liquid handling equipment, making it more accessible to proteomic laboratories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Modeling Clustered DNA Damage by Ionizing Radiation Using Multinomial Damage Probabilities and Energy Imparted Spectra

Simple and complex clustered DNA damage represent the critical initial damage caused by radiation. In this paper, a multinomial probability model of clustered damage is developed with probabilities dependent on the energy imparted to DNA and surrounding water molecules. The model consists of four probabilities: (A) direct damage of sugar-phosphate moieties leading to SSB, (B) OH− radical formation with subsequent SSB and BD formation, (C) direct damage to DNA bases, and (D) energy imparted to histone proteins and other molecules in a volume not leading to SSB or BD. These probabilities are augmented by introducing probabilities for the relative location of SSB using a ≤10 bp criteria for a double-strand break (DSB) and for the possible success of a radical attack that leads to SSB or BD. Model predictions for electrons, 4He, and 12C ions are compared to the experimental data and show good agreement. Thus, the developed model allows an accurate and rapid computational method to predict simple and complex clustered DNA damage as a function of radiation quality and to explore the resulting challenges to DNA repair.

Biochemistry & Molecular Biology

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES

Structural Insights into Mechanisms Underlying Mitochondrial and Bacterial Cytochrome c Synthases

Mitochondrial holocytochrome c synthase (HCCS) is an essential protein in assembling cytochrome c (cyt c) of the electron transport system. HCCS binds heme and covalently attaches the two vinyls of heme to two cysteine thiols of the cyt c CXXCH motif. Human HCCS recognizes both cyt c and cytochrome c1 of complex III (cytochrome bc1). HCCS is mutated in some human diseases and it has been investigated recombinantly by mutational, biochemical, and reconstitution studies in the past decade. Here, we employ structural prediction programs (e.g., AlphaFold 3) on HCCS and its two substrates, heme and cytochrome c. The results, when combined with spectroscopic and functional analyses of HCCS and variants, provide insights into the structural basis for heme binding, apocyt c binding, covalent attachment, and release of the holocyt c product. Results from in vitro reconstitution of purified human HCCS using cyt c and cyt c1 peptides as acceptors are consistent with the structural modeling of substrate binding. Reconstitution of HCCS and cyt c1 provides an approach to studying cyt c1 assembly, which has been refractile to recombinant in vivo reconstitution (unlike HCCS and cyt c). We propose a structural basis for release of the holocyt c product from HCCS based on in vitro studies and on cryoEM structures of the bacterial cyt c synthase (CcsBA) active site. We analyze the kinetoplastid mitochondrial synthase (KCCS), and hypothesize a molecular evolutionary path from mitochondrial endosymbiosis to the current HCCS.

Biochemistry & Molecular Biology

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics

Knocking out the carboxyltransferase interactor 1 (CTI1) in Chlamydomonas boosted oil content by fivefold without affecting cell growth

Summary The first step in chloroplast de novo fatty acid synthesis is catalysed by acetyl‐CoA carboxylase (ACCase). As the rate‐limiting step for this pathway, ACCase is subject to both positive and negative regulation. In this study, we identify a Chlamydomonas homologue of the plant carboxyltransferase interactor 1 (CrCTI1) and show that this protein interacts with the Chlamydomonas α‐carboxyltransferase (Crα‐CT) subunit of the ACCase by yeast two‐hybrid protein–protein interaction assay. Three independent CRISPR‐Cas9 mediated knockout mutants for CrCTI1 each produced an ‘enhanced oil’ phenotype, accumulating 25% more total fatty acids and storing up to fivefold more triacylglycerols (TAGs) in lipid droplets. The TAG phenotype of the crcti1 mutants was not influenced by light but was affected by trophic growth conditions. By growing cells under heterotrophic conditions, we observed a crucial function of CrCTI1 in balancing lipid accumulation and cell growth. Mutating a previously mapped in vivo phosphorylation site (CrCTI1 Ser108 to either Ala or to Asp), did not affect the interaction with Crα‐CT. However, mutating all six predicted phosphorylation sites within Crα‐CT to create a phosphomimetic mutant reduced this pairwise interaction significantly. Comparative proteomic analyses of the crcti1 mutants and WT suggested a role for CrCTI1 in regulating carbon flux by coordinating carbon metabolism, antioxidant and fatty acid β‐oxidation pathways, to enable cells to adapt to carbon availability. Taken together, this study identifies CrCTI1 as a negative regulator of fatty acid synthesis in algae and provides a new molecular brick for the genetic engineering of microalgae for biotechnology purposes.

Li, Zhongze [Aix‐Marseille Université, CEA, CNRS,

Uncovering Sequence and Structural Characteristics of Fungal Expansin‐Related Proteins With Potential to Drive Substrate Targeting

Expansins loosen plant cell wall networks through disrupting non-covalent bonds between cellulose microfibrils and matrix polysaccharides. Whereas expansins were first discovered in plants, expansin-related proteins have since been identified in bacteria and fungi. The biological function of microbial expansins remains unclear; however, several studies have shown distinct binding preferences toward different structural polysaccharides. Earlier studies of bacterial expansin-related proteins uncovered sequence and structural features that correlate to substrate binding. Herein, 20 fungal expansin-related sequences were recombinantly produced in Komagataella phaffii, and the purified proteins were compared in terms of substrate binding to cellulosic and chitinous substrates. The impact of pH on the zeta potential of prioritized substrates was also measured, and Principal Component Analysis was performed to uncover correlations between protein characteristics (e.g., pI, hydrophobicity, surface charge distribution) and measured substrate binding preferences. Whereas acidic proteins with a predicted pI less than 5.0 preferentially bound to chitin, basic proteins with pI greater than 8.0 preferentially bound to xylan and xylan-containing fiber. Similar to many cellulases, binding to cellulose was correlated to relatively high aromatic amino acid content in the protein sequence and presence of a carbohydrate binding module (CBM), which in the case of expansins is a C-terminal CBM63. Whereas overall sequence characteristics could be correlated to substrate binding preference, the identity of amino acids occupying conserved positions that impact protein activity was better correlated with loosenin versus expansin classifications.

chitin

Integrative Modeling and Analysis of Fungal Central Carbon Metabolism

Over a thousand fungal genomes have been sequenced, yet manually curated genome-scale metabolic models (GEMs) are available for only a limited number of species. Moreover, these models have often been developed independently, leading to inconsistencies in namespaces, compartment definitions, and pathway representations that hinder comparative analysis, the systematic reuse of prior curation efforts, and the integration of consolidated metabolic knowledge. Here, we present the Consolidated Fungal Core Metabolism Model (CFCMM), constructed by integrating thirteen published fungal models spanning Ascomycota, Mucoromycota, and both Crabtree-positive and Crabtree-negative yeasts. We harmonized metabolites and reactions into a non-redundant shared ModelSEED ontological space, standardized compartmentalization, and refined gene–protein–reaction (GPR) rules. Using pathway-level visualization and systematic gap detection, we further improved the integrated network through literature-guided curation to correct stoichiometry, stereospecificity, and pathway architecture. Orthologous protein family reconstruction and functional annotation workflows were used to validate and inform GPR associations, with particular emphasis on ambiguous enzyme superfamilies and membrane-associated components. Using the resulting CFCMM, we built high-quality central carbon core models for each fungus and performed flux balance analysis to quantify ATP-yield variation under aerobic and anaerobic conditions, explicitly evaluating scenarios driven by differences in electron transport chain (ETC) composition. Simulations reproduced the expected fermentative yield of approximately 2 mmol ATP per mmol glucose under anaerobic conditions and separated the thirteen fungi into two bioenergetic groups under aerobic respiration based on Complex I status, with predicted yields of approximately 30 versus 22 mmol ATP per mmol glucose. Forcing flux through the alternative oxidase bypass further reduced ATP yields to approximately 12 and 4 mmol ATP per mmol glucose in Complex I-containing and Complex I-lacking fungi, respectively. Collectively, this work provides a manually curated, ModelSEED-consistent, and extensible fungal core metabolic template, deployed in DOE KBase as a resource for automated reconstruction of central carbon core models from any sequenced fungal genome. In addition, the CFCMM provides modular components for developing GEMs with more accurate energy predictions and enables robust comparative analyses of fungal bioenergetics and core metabolic diversity

59 BASIC BIOLOGICAL SCIENCES

Structural Heterogeneity and Hydrodynamics of an Intrinsically Disordered Protein Condensate

Biology demonstrates precise control over the free-energy landscape through the selective partitioning of biomacromolecules into membraneless organelles, enabling essential functions such as biochemical transformations, signaling cascades, and mechanical reinforcement. Although the function of these condensates depends on their underlying structure and hydrodynamics, molecular-scale information on these systems remains sparse. Here, in this study, neutron scattering is used to probe the organization and dynamics of the intrinsically disordered N-terminal domain of Galectin-3, an extracellular lectin responsible for facilitating liquid–liquid phase separation on the cellular surface, in both dilute and condensed phases. Dilute solutions contain isolated protein chains in equilibrium with mesoscopic clusters, whereas the condensed phase adopts a bicontinuous, microemulsion-like morphology. The dilute phase behavior is quantitatively described by coarse-grained polymer models from soft-matter physics, demonstrating their predictive power for complex biological proteins. At elevated concentrations, the proteins self-assemble akin to block copolymers, microphase separating through the aggregation of hydrophobic domains along the protein contour. The resulting condensate remains fluid-like despite a 25-fold increase in concentration; its internal hydrodynamics slow by only a factor of 3 relative to dilute protein chains. These results provide a molecular-level framework for how disordered proteins achieve both the structural complexity and dynamic fluidity of biomolecular condensates.

Carrick, Brian R. [Massachusetts Inst. of Technolo

Characterization of a widespread sugar phosphate-processing bacterial microcompartment

Many prokaryotes form Bacterial Microcompartments (BMCs) that encapsulate segments of specialized metabolic pathways to enhance catalysis. The various functions of metabolosomes, catabolic BMCs, are dictated by the signature enzyme that processes initial substrates of the confined pathway. The components and native functions of several metabolosomes have been experimentally characterized; however one of the most prevalent across all bacteria has yet to be studied. Sugar Phosphate Utilizing (SPU) BMC loci encode enzymes predicted to be involved in sugar phosphate metabolism. The SPU genetic loci are found in organisms occupying habitats ranging from soils to hot springs, highlighting the ubiquity of the SPU BMC. We bioinformatically characterized seven SPU subtypes, all which contain an enzyme unique to SPU BMCs, a deoxyribose 5-phosphate aldolase (DERA). Here, we define the fundamental characteristics of SPU BMCs and have expressed, purified, and characterized a set of SPU core enzymes. These include a protein-protein complex formed between a SPU BMC DERA and a predicted ribose 5-phosphate isomerase. Further, we show that the SPU BMC DERA is catalytically active and propose that it acts as the universal signature enzyme for the SPU BMC, with implications for fundamental understanding and biotechnological applications of SPU BMCs.

59 BASIC BIOLOGICAL SCIENCES

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES