Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein model quality assessment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

3D-equivariant graph neural networks for protein model quality assessment

Quality assessment (QA) of predicted protein tertiary structure models plays an important role in ranking and using them. With the recent development of deep learning end-to-end protein structure prediction techniques for generating highly confident tertiary structures for most proteins, it is important to explore corresponding QA strategies to evaluate and select the structural models predicted by them since these models have better quality and different properties than the models predicted by traditional tertiary structure prediction methods. We develop EnQA, a novel graph-based 3D-equivariant neural network method that is equivariant to rotation and translation of 3D objects to estimate the accuracy of protein structural models by leveraging the structural features acquired from the state-of-the-art tertiary structure prediction method—AlphaFold2. We train and test the method on both traditional model datasets (e.g. the datasets of the Critical Assessment of Techniques for Protein Structure Prediction) and a new dataset of high-quality structural models predicted only by AlphaFold2 for the proteins whose experimental structures were released recently. Our approach achieves state-of-the-art performance on protein structural models predicted by both traditional protein structure prediction methods and the latest end-to-end deep learning method—AlphaFold2. It performs even better than the model QA scores provided by AlphaFold2 itself. The results illustrate that the 3D-equivariant graph neural network is a promising approach to the evaluation of protein structural models. Integrating AlphaFold2 features with other complementary sequence and structural features is important for improving protein model QA.

59 BASIC BIOLOGICAL SCIENCES↗

Protein model quality assessment using rotation–equivariant transformations on point clouds

Machine learning research concerning protein structure has seen a surge in popularity over the last years with promising advances for basic science and drug discovery. Working with macromolecular structure in a machine learning context requires an adequate numerical representation, and researchers have extensively studied representations such as graphs, discretized 3D grids, and distance maps. As part of CASP14, we explored a new and conceptually simple representation in a blind experiment: atoms as points in 3D, each with associated features. These features—initially just the basic element type of each atom—are updated through a series of neural network layers featuring rotation-equivariant convolutions. Starting from all atoms, we further aggregate information at the level of alpha carbons before making a prediction at the level of the entire protein structure. We find that this approach yields competitive results in protein model quality assessment despite its simplicity and despite the fact that it incorporates minimal prior information and is trained on relatively little data. As a result, its performance and generality are particularly noteworthy in an era where highly complex, customized machine learning methods such as AlphaFold 2 have come to dominate protein structure prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15

Abstract Estimating the accuracy of quaternary structural models of protein complexes and assemblies (EMA) is important for predicting quaternary structures and applying them to studying protein function and interaction. The pairwise similarity between structural models is proven useful for estimating the quality of protein tertiary structural models, but it has been rarely applied to predicting the quality of quaternary structural models. Moreover, the pairwise similarity approach often fails when many structural models are of low quality and similar to each other. To address the gap, we developed a hybrid method (MULTICOM_qa) combining a pairwise similarity score (PSS) and an interface contact probability score (ICPS) based on the deep learning inter‐chain contact prediction for estimating protein complex model accuracy. It blindly participated in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 and performed very well in estimating the global structure accuracy of assembly models. The average per‐target correlation coefficient between the model quality scores predicted by MULTICOM_qa and the true quality scores of the models of CASP15 assembly targets is 0.66. The average per‐target ranking loss in using the predicted quality scores to rank the models is 0.14. It was able to select good models for most targets. Moreover, several key factors (i.e., target difficulty, model sampling difficulty, skewness of model quality, and similarity between good/bad models) for EMA are identified and analyzed. The results demonstrate that combining the multi‐model method (PSS) with the complementary single‐model method (ICPS) is a promising approach to EMA.

59 BASIC BIOLOGICAL SCIENCES↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Overall protein structure quality assessment using hydrogen-bonding parameters

Atomic model refinement at low resolution is often a challenging task. This is mostly because the experimental data are not sufficiently detailed to be described by atomic models. To make refinement practical and ensure that a refined atomic model is geometrically meaningful, additional information needs to be used such as restraints on Ramachandran plot distributions or residue side-chain rotameric states. However, using Ramachandran plots or rotameric states as refinement targets diminishes the validating power of these tools. Therefore, finding additional model-validation criteria that are not used or are difficult to use as refinement goals is desirable. Hydrogen bonds are one of the important noncovalent interactions that shape and maintain protein structure. These interactions can be characterized by a specific geometry of hydrogen donor and acceptor atoms. Systematic analysis of these geometries performed for quality-filtered high-resolution models of proteins from the Protein Data Bank shows that they have a distinct and a conserved distribution. Here, it is demonstrated how this information can be used for atomic model validation.

59 BASIC BIOLOGICAL SCIENCES↗

A gated graph transformer for protein complex structure quality assessment and its performance in CASP15

Abstract Motivation Proteins interact to form complexes to carry out essential biological functions. Computational methods such as AlphaFold-multimer have been developed to predict the quaternary structures of protein complexes. An important yet largely unsolved challenge in protein complex structure prediction is to accurately estimate the quality of predicted protein complex structures without any knowledge of the corresponding native structures. Such estimations can then be used to select high-quality predicted complex structures to facilitate biomedical research such as protein function analysis and drug discovery. Results In this work, we introduce a new gated neighborhood-modulating graph transformer to predict the quality of 3D protein complex structures. It incorporates node and edge gates within a graph transformer framework to control information flow during graph message passing. We trained, evaluated and tested the method (called DProQA) on newly-curated protein complex datasets before the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) and then blindly tested it in the 2022 CASP15 experiment. The method was ranked 3rd among the single-model quality assessment methods in CASP15 in terms of the ranking loss of TM-score on 36 complex targets. The rigorous internal and external experiments demonstrate that DProQA is effective in ranking protein complex structures. Availability and implementation The source code, data, and pre-trained models are available at https://github.com/jianlin-cheng/DProQA.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

Measurements of Protein Crystal Face Growth Rates

Protein crystal growth rates will be determined for several hyperthermophile proteins.; The growth rates will be assessed using available theoretical models, including kinetic roughening.; If/when kinetic roughening supersaturations are established, determinations of protein crystal quality over a range of supersaturations will also be assessed.; The results of our ground based effort may well address the existence of a correlation between fundamental growth mechanisms and protein crystal quality.

Gorti, S.↗

Accuracy-Based Annotation Quality Score (ABAQS) v1.0

Assessing genome annotation quality is crucial for downstream analyses, but current methods are inadequate for eukaryotes. We present Accuracy-Based Annotation Quality Score (ABAQS), a novel, minimal-data-driven method that comprehensively assesses annotation quality. ABAQS evaluates multiple factors, including genome completeness, gene model validity, and protein profile accuracy, outperforming other metrics like BUSCO and PSAURON. We applied ABAQS to over 2500 eukaryotic genomes and showed its robustness and effectiveness in evaluating genome annotation quality, making it a valuable tool for researchers working with genomic data. ABAQS reveals significant variation in annotation quality and highlights the importance of filtering in improving annotation quality and accuracy.

Haridas, Sajeet [Lawrence Berkeley National Labora↗

Systematic Evaluation of Counterpoise Correction in Density Functional Theory

A widespread belief persists that the Boys–Bernardi function counterpoise (CP) procedure “overcorrects” supramolecular interaction energies for the effects of basis-set superposition error. To the extent that this is true for correlated wave function methods, it is usually an artifact of low-quality basis sets. The question has not been considered systematically in the context of density functional theory, however, where basis-set convergence is generally less problematic. We present a systematic assessment of the CP procedure for a representative set of functionals and basis sets, considering both benchmark data sets of small dimers and larger supramolecular complexes. The latter include layered composite polymers with ~150 atoms and ligand–protein models with ~300 atoms. Provided that CP correction is used, we find that intermolecular interaction energies of nearly complete-basis quality can be obtained using only double-ζ basis sets. Furthermore, this is less expensive as compared to triple-ζ basis sets without CP correction. CP-corrected interaction energies are less sensitive to the presence of diffuse basis functions as compared to uncorrected energies, which is important because diffuse functions are expensive and often numerically problematic for large systems. Our results upend the conventional wisdom that CP “overcorrects” for basis-set incompleteness. In small basis sets, CP correction is mandatory in order to demonstrate that the results do not rest on error cancellation.

74 ATOMIC AND MOLECULAR PHYSICS↗

Smart culture medium optimization for recombinant protein production: Experimental, modeling, and AI/ML-driven strategies

Recombinant protein production (RPP) is central to biotechnology, where recombinant proteins are used as either end products or catalysts in the synthesis of chemicals, fuels, and materials. Among the major cost drivers, culture medium plays a pivotal role in determining protein yield and quality. This review presents a comprehensive perspective on the critical stages of “smart” culture medium optimization: planning, screening, modeling, optimization, and validation. In the planning stage, we examine the nutritional and energetic roles of medium components, including carbon, nitrogen, amino acids, salts, and trace metals, and their impacts on culture parameters such as pH, oxidative state, and osmolality. We highlight the variability in trace metal content due to water sources, culture vessels, and raw materials, which can substantially influence RPP. The screening stage covers Design of Experiments (DoE) approaches, assessing their theoretical basis, implementation, and limitations. For modeling, we describe methods that integrate experimental data to develop predictive models for smart medium formulation. Model-based optimization strategies can then be employed to select optimal media compositions for a given application. The validation stage aims to evaluate model predictions and provide feedback for model training and refinement. Finally, we survey mechanistic and artificial intelligence/machine learning (AI/ML)-driven models as integrated, transformational tools for predictive modeling of bioprocess conditions, nutrient availability, cellular metabolism, and protein quality, with the goal of optimizing culture media to enhance protein yields while reducing costs and environmental impact. We conclude by addressing the challenges of translating laboratory-scale medium optimization to industrial-scale settings and exploring future AI/ML-driven approaches that may overcome current bottlenecks and accelerate medium design for RPP. Overall, this review provides a unified framework for advancing smart medium design in RPP.

Artificial Intelligence/Machine Learning (AI/ML)↗

Response of thyroid follicular cells to gamma irradiation compared to proton irradiation: II. The role of connexin 32

The objective of this study was to determine whether connexin 32-type gap junctions contribute to the "contact effect" in follicular thyrocytes and whether the response is influenced by radiation quality. Our previous studies demonstrated that early-passage follicular cultures of Fischer rat thyroid cells express functional connexin 32 gap junctions, with later-passage cultures expressing a truncated nonfunctional form of the protein. This model allowed us to assess the role of connexin 32 in radiation responsiveness without relying solely on chemical manipulation of gap junctions. The survival curves generated after gamma irradiation revealed that early-passage follicular cultures had significantly lower values of alpha (0.04 Gy(-1)) than later-passage cultures (0.11 Gy(-1)) (P < 0.0001, n = 12). As an additional way to determine whether connexin 32 was contributing to the difference in survival, cultures were treated with heptanol, resulting in higher alpha values, with early-passage cultures (0.10 Gy(-1)) nearly equivalent to untreated late-passage cultures (0.11 Gy(-1)) (P > 0.1, n = 9). This strongly suggests that the presence of functional connexin 32-type gap junctions was contributing to radiation resistance in gamma-irradiated thyroid follicles. Survival curves from proton-irradiated cultures had alpha values that were not significantly different whether cells expressed functional connexin 32 (0.10 Gy(-1)), did not express connexin 32 (0.09 Gy(-1)), or were down-regulated (early-passage plus heptanol, 0.09 Gy(-1); late-passage plus heptanol, 0.12 Gy(-1)) (P > 0.1, n = 19). Thus, for proton irradiation, the presence of connexin 32-type gap junctional channels did not influence their radiosensitivity. Collectively, the data support the following conclusions. (1) The lower alpha values from the gamma-ray survival curves of the early-passage cultures suggest greater repair efficiency and/or enhanced resistance to radiation-induced damage, coincident with the expression of connexin 32-type gap junctions. (2) The increased sensitivity of FRTL-5 cells to proton irradiation was independent of their ability to communicate through connexin 32 gap junctions. (3) The fact that the beta components of the survival curves from both gamma rays and proton beams were similar (average 0.022 +/- 0.008 Gy(-2), P > 0.1, n = 39) suggests that at higher doses the loss of viability occurs at a relatively constant rate and is independent of radiation quality and the presence of functional gap junctions.

Non-NASA Center↗

Radiation Risk Assessment of the Individual Astronaut: A Complement to Radiation Interests at the NIH

Predicting human risks following exposure to space radiation is uncertain in part because of unpredictable distribution of high-LET and low-dose-derived damage amongst cells in tissues, unknown synergistic effects of microgravity upon gene- and protein-expression, and inadequately modeled processing of radiation-induced damage within cells to produce rare and late-appearing malignant cancers. Furthermore, estimation of risks of radiogenic outcome within small numbers of astronauts is not possible using classic epidemiologic study. It therefore seems useful to develop strategies of risk-assessment based upon large datasets acquired from correlated biological models useful for resolving radiogenic risk-assessment for irradiated individuals. In this regard, it is suggested that sensitive cellular biodosimeters that simultaneously report 1) the quantity of absorbed dose after exposure to ionizing radiation, 2) the quality of radiation delivering that dose, and 3) the biomolecular risk of malignant transformation be developed in order to resolve these NASA-specific challenges. Multiparametric cellular biodosimeters could be developed using analyses of gene-expression and protein-expression whereby large datasets of cellular response to radiation-induced damage are analyzed for markers predictive for acute response as well as cancer-risk. A new paradigm is accordingly addressed wherein genomic and proteomic datasets are registered and interrogated in order to provide statistically significant dose-dependent risk estimation in individual astronauts. This evaluation of the individual for assessment of radiogenic outcomes connects to NIH program in that such a paradigm also supports assignment of a given patient to a specific therapy, the diagnosis of response of that patient to therapy, and the prediction of risks accumulated by that patient during therapy - such as risks incurred by scatter and neutrons produced during high-energy Intensity-Modulated Radiation Therapy. Value of assessment of radiogenic outcome for individuals exposed to radiation is suggested to be common to both NASA and NIH.

Richmond, Robert C.↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Invasions eliminate the legacy effects of substrate history on microbial nitrogen cycling

Abstract Changes in substrate quality driven by climate, land use, or other forms of global change may represent a strong selective force on microbial communities. Invasion of new taxa into a community through dispersal, evolution, or recolonization could impact the outcome of this environmental selection. Here, we simulated substrate change with a trait‐based model of microbial litter decomposition (DEMENTpy) to assess the legacy effects of past substrate quality and the impact of selection by a new substrate on community decomposition activity. Simulations were run with different levels of invasion, including invasion from communities long‐adapted to the new substrate. Legacy effects were evident with substrate change for native communities differing in composition. Protein was the only substrate that exerted a strong enough selective force to affect community composition. Legacy effects disappeared when invaders came from substrates similar to the new substrate. Together, our simulations demonstrate that substrate quality changes associated with global change can lead to legacy effects on substrate degradation. In decomposing plant litter, such legacy effects can occur if substrate inputs shift to higher protein content and if invasion is low.

54 ENVIRONMENTAL SCIENCES↗

Microarray Data Analysis of Space Grown Arabidopsis Leaves for Genes Important in Vascular Patterning

Venation patterning in leaves is a major determinant of photosynthesis efficiency because of its dependency on vascular transport of photoassimilates, water, and minerals. Arabidopsis thaliana grown in microgravity show delayed growth and leaf maturation. Gene expression data from the roots, hypocotyl, and leaves of A. thaliana grown during spaceflight vs. ground control analyzed by Affymetrix microarray are available through NASAs GeneLab (GLDS-7). We analyzed the data for differential expression of genes in leaves resulting from the effects of spaceflight on vascular patterning. Two genes were found by preliminary analysis to be upregulated during spaceflight that may be related to vascular formation. The genes are responsible for coding an ARGOS like protein (potentially affecting cell elongation in the leaves), and an F-boxkelch-repeat protein (possibly contributing to protoxylem specification). Further analysis that will focus on raw data quality assessment and a moderated t-test may further confirm upregulation of the two genes and/or identify other gene candidates. Plants defective in these genes will then be assessed for phenotype by the mapping and quantification of leaf vascular patterning by NASAs VESsel GENeration (VESGEN) software to model specific vascular differences of plants grown in spaceflight.

Weitzeal, A. J.↗

Microarray Data Analysis of Space Grown Arabidopsis Leaves for Genes Important in Vascular Patterning

Venation patterning in leaves is a major determinant of photosynthesis efficiency because of its dependency on vascular transport of photoassimilates, water, and minerals. Arabidopsis thaliana grown in microgravity show delayed growth and leaf maturation. Gene expression data from the roots, hypocotyl, and leaves of A. thaliana grown during spaceflight vs. ground control analyzed by Affymetrix microarray are available through NASA's GeneLab (GLDS-7). We analyzed the data for differential expression of genes in leaves resulting from the effects of spaceflight on vascular patterning. Two genes were found by preliminary analysis to be upregulated during spaceflight that may be related to vascular formation. The genes are responsible for coding an ARGOS like protein (potentially affecting cell elongation in the leaves), and an F-box/kelch-repeat protein (possibly contributing to protoxylem specification). Further analysis that will focus on raw data quality assessment and a moderated t-test may further confirm upregulation of the two genes and/or identify other gene candidates. Plants defective in these genes will then be assessed for phenotype by the mapping and quantification of leaf vascular patterning by NASA's VESsel GENeration (VESGEN) software to model specific vascular differences of plants grown in spaceflight.

Weitzeal, A. J.↗

Microarray Data Analysis of Space Grown Arabidopsis Leaves for Genes Important in Vascular Patterning

Venation patterning in leaves is a major determinant of photosynthesis efficiency because of its dependency on vascular transport of photo-assimilates, water, and minerals. Arabidopsis thaliana grown in microgravity show delayed growth and leaf maturation. Gene expression data from the roots, hypocotyl, and leaves of A. thaliana grown during spaceflight vs. ground control analyzed by Affymetrix microarray are available through NASA's GeneLab (GLDS-7). We analyzed the data for differential expression of genes in leaves resulting from the effects of spaceflight on vascular patterning. Two genes were found by preliminary analysis to be up-regulated during spaceflight that may be related to vascular formation. The genes are responsible for coding an ARGOS (Auxin-Regulated Gene Involved in Organ Size)-like protein (potentially affecting cell elongation in the leaves), and an F-box/kelch-repeat protein (possibly contributing to protoxylem specification). Further analysis that will focus on raw data quality assessment and a moderated t-test may further confirm up-regulation of the two genes and/or identify other gene candidates. Plants defective in these genes will then be assessed for phenotype by the mapping and quantification of leaf vascular patterning by NASA's VESsel GENeration (VESGEN) software to model specific vascular differences of plants grown in spaceflight.

venation↗