Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Covalent Drug Binding in Live Cells Monitored by Mid-Infrared Quantum Cascade Laser Spectroscopy: Photoactive Yellow Protein as a Model System

The detection of drug-target interactions in live cells enables analysis of therapeutic compounds in a native cellular environment. Recent advances in spectroscopy and molecular biology have facilitated the development of genetically encoded vibrational probes like nitriles that can sensitively report on molecular interactions. Nitriles are powerful tools for measuring electrostatic environments within condensed media like proteins, but such measurements in live cells have been hindered by low signal-to-noise ratios. In this study, we design a spectrometer based on a double-beam quantum cascade laser (QCL)-based transmission infrared (IR) source with balanced detection that can significantly enhance sensitivity to nitrile vibrational probes embedded in proteins within cells compared to a conventional FTIR spectrometer. Here, using this approach, we detect small-molecule binding in Escherichia coli, with particular focus on the interaction between para-Coumaric acid (pCA) and nitrile-incorporated photoactive yellow protein (PYP). This system effectively serves as a model for investigating covalent drug binding in a cellular environment. Notably, we observe large spectral shifts of up to 15 cm –1 for nitriles embedded in PYP between the unbound and drug-bound states directly within bacteria, in agreement with observations for purified proteins. Such large spectral shifts are ascribed to the changes in the hydrogen-bonding environment around the local environment of nitriles, accurately modeled through high-level molecular dynamics simulations using the AMOEBA force field. Our findings underscore the QCL spectrometer’s ability to enhance sensitivity for monitoring drug–protein interactions, offering new opportunities for advanced methodologies in drug development and biochemical research.

chromophores↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology↗

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES↗

Unfolding of the Villin Headpiece Domain: Revealing Structural Heterogeneity with Time‐Resolved X‐Ray Solution Scattering and Markov State Modeling

Understanding protein folding pathways is crucial to deciphering the principles of protein structure and function. Here, the unfolding dynamics of the 35‐residue villin headpiece (HP35) and a norleucine‐substituted variant (2F4K) using a combination of experimental and computational techniques is investigated. Time‐resolved X‐ray solution scattering coupled with equilibrium molecular dynamics simulations and Markov state modeling reveals distinct unfolding mechanisms between the two variants: HP35 and 2F4K. Specifically, HP35 exhibits a two‐state unfolding process, whereas an intermediate state is identified for the 2F4K mutant. A Markov state model constructed from simulations is used to map atomic‐level transitions to experimental observations, providing insights into the role of sequence variations in modulating folding pathways. The findings underscore the importance of integrating experimental and computational approaches to unravel protein unfolding mechanisms between heterogenous structural ensembles.

Nijhawan, Adam K. [Department of Chemistry Northwe↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Physical models reveal indirect reader protein interactions that facilitate epigenetic crosstalk

The spatial organization of chromatin is governed by epigenetic factors, including epigenetic marks and the reader proteins that bind them. By dictating the accessibility of genomic loci, epigenetic factors contribute to the physical regulation of gene expression, enabling diverse cellular phenotypes to be encoded by a shared genome in an individual. Epigenetic dysregulation can lead to aberrations in chromatin architecture, contributing to diseases such as neurological disorders and cancers. Despite the known importance of chromatin organization for human health, the physical mechanisms governing chromatin folding remain underspecified. In this work, we develop a physical model of chromatin organization based on contributions from multiple epigenetic factors. Using our model, we evaluate how conditions in the nuclear environment and crosstalk between epigenetic marks affect the compartmentalization of chromatin into dense heterochromatin and loose euchromatin. Our results emphasize the role of reader protein binding in chromatin compartmentalization. We show that reader proteins interact through an indirect mechanism facilitated by the shared chromatin “scaffold” to which they bind. Under a scenario where reader proteins compete for binding sites, we find that indirect interactions affect the program adopted by the chromatin fiber. By isolating indirect modes of epigenetic crosstalk, we demonstrate how the interplay between epigenetic patterning and environmental factors influences chromatin architecture.

59 BASIC BIOLOGICAL SCIENCES↗

Utilizing Machine Learning to Improve Neutralization Potency of an HIV-1 Antibody Targeting the gp41 N-Heptad Repeat

The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.

Biopolymers↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation

This dataset accompanies the publication "ProtNHF: Neural Hamiltonian Flows for Controllable Protein Sequence Generation". This paper introduces a new AI model for protein sequence generation. This dataset contains data related to experiments discussed in the publication. This includes generated sequences and evaluation metrics supporting all unconditional and bias-controlled experiments in the ProtNHF paper.

60 APPLIED LIFE SCIENCES↗

Can protein expression be ‘solved’?

Recombinant protein expression is central to biotechnology’s application in academic exploration as well as human health, climate applications and the bioeconomy in general. However, not all proteins can be expressed in all organisms, and the field lacks a predictive model of soluble protein overexpression that could replace laborious experimental trial-and-error. Here, we discuss the state of the field and identify the lack of large, high-fidelity datasets as the primary bottleneck to progress. We review possible assays that could be used for data collection to identify a path toward an extensible experimental platform for collecting soluble recombinant protein overexpression data across organisms. We suggest that the resulting dataset should be used to train increasingly generalizable predictive models of protein expression to answer the question: “How can predictive protein expression be solved?”.

59 BASIC BIOLOGICAL SCIENCES↗