Search NASA⌕ Search

SEARCH · Search NASA

Results for “estimation of protein model accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Combining pairwise structural similarity and deep learning interface contact prediction to estimate protein complex model accuracy in CASP15

Abstract Estimating the accuracy of quaternary structural models of protein complexes and assemblies (EMA) is important for predicting quaternary structures and applying them to studying protein function and interaction. The pairwise similarity between structural models is proven useful for estimating the quality of protein tertiary structural models, but it has been rarely applied to predicting the quality of quaternary structural models. Moreover, the pairwise similarity approach often fails when many structural models are of low quality and similar to each other. To address the gap, we developed a hybrid method (MULTICOM_qa) combining a pairwise similarity score (PSS) and an interface contact probability score (ICPS) based on the deep learning inter‐chain contact prediction for estimating protein complex model accuracy. It blindly participated in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 and performed very well in estimating the global structure accuracy of assembly models. The average per‐target correlation coefficient between the model quality scores predicted by MULTICOM_qa and the true quality scores of the models of CASP15 assembly targets is 0.66. The average per‐target ranking loss in using the predicted quality scores to rank the models is 0.14. It was able to select good models for most targets. Moreover, several key factors (i.e., target difficulty, model sampling difficulty, skewness of model quality, and similarity between good/bad models) for EMA are identified and analyzed. The results demonstrate that combining the multi‐model method (PSS) with the complementary single‐model method (ICPS) is a promising approach to EMA.

59 BASIC BIOLOGICAL SCIENCES↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

3D-equivariant graph neural networks for protein model quality assessment

Quality assessment (QA) of predicted protein tertiary structure models plays an important role in ranking and using them. With the recent development of deep learning end-to-end protein structure prediction techniques for generating highly confident tertiary structures for most proteins, it is important to explore corresponding QA strategies to evaluate and select the structural models predicted by them since these models have better quality and different properties than the models predicted by traditional tertiary structure prediction methods. We develop EnQA, a novel graph-based 3D-equivariant neural network method that is equivariant to rotation and translation of 3D objects to estimate the accuracy of protein structural models by leveraging the structural features acquired from the state-of-the-art tertiary structure prediction method—AlphaFold2. We train and test the method on both traditional model datasets (e.g. the datasets of the Critical Assessment of Techniques for Protein Structure Prediction) and a new dataset of high-quality structural models predicted only by AlphaFold2 for the proteins whose experimental structures were released recently. Our approach achieves state-of-the-art performance on protein structural models predicted by both traditional protein structure prediction methods and the latest end-to-end deep learning method—AlphaFold2. It performs even better than the model QA scores provided by AlphaFold2 itself. The results illustrate that the 3D-equivariant graph neural network is a promising approach to the evaluation of protein structural models. Integrating AlphaFold2 features with other complementary sequence and structural features is important for improving protein model QA.

59 BASIC BIOLOGICAL SCIENCES↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

Development of a Systematic and Extensible Force Field for Peptoids (STEPs)

Peptoids (N-substituted glycines) are a class of biomimetic polymers that have attracted significant attention due to their accessible synthesis and enzymatic and thermal stability relative to their naturally occurring counterparts (polypeptides). While these polymers provide the promise of more robust functional materials via hierarchical approaches, they present a new challenge for computational structure prediction for material design. The reliability of calculations hinges on the accuracy of interactions represented in the force field used to model peptoids. For proteins, structure prediction based on sequence and de novo design has made dramatic progress in recent years; however, these models are not readily transferable for peptoids. Current efforts to develop and implement peptoid-specific force fields are spread out, leading to replicated efforts and a fragmented collection of parameterized sidechains. Here, we developed a peptoid-specific force field containing 70 different side chains, using GAFF2 as starting point. The new model is validated based on the generation of Ramachandran-like plots from DFT optimization compared against force field reproduced potential energy and free energy surfaces as well as the reproduction of equilibrium cis/trans values for some residues experimentally known to form helical structures. In conclusion, equilibrium cis/trans distributions (Kct) are estimated for all parameterized residues to identify which residues have an intrinsic propensity for cis or trans states in the monomeric state.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Direct Estimation of Parameters in ODE Models Using WENDy: Weak-Form Estimation of Nonlinear Dynamics

Abstract We introduce the Weak-form Estimation of Nonlinear Dynamics (WENDy) method for estimating model parameters for non-linear systems of ODEs. Without relying on any numerical differential equation solvers, WENDy computes accurate estimates and is robust to large (biologically relevant) levels of measurement noise. For low dimensional systems with modest amounts of data, WENDy is competitive with conventional forward solver-based nonlinear least squares methods in terms of speed and accuracy. For both higher dimensional systems and stiff systems, WENDy is typically both faster (often by orders of magnitude) and more accurate than forward solver-based approaches. The core mathematical idea involves an efficient conversion of the strong form representation of a model to its weak form, and then solving a regression problem to perform parameter inference. The core statistical idea rests on the Errors-In-Variables framework, which necessitates the use of the iteratively reweighted least squares algorithm. Further improvements are obtained by using orthonormal test functions, created from a set of $$C^{\infty }$$ C ∞ bump functions of varying support sizes.We demonstrate the high robustness and computational efficiency by applying WENDy to estimate parameters in some common models from population biology, neuroscience, and biochemistry, including logistic growth, Lotka-Volterra, FitzHugh-Nagumo, Hindmarsh-Rose, and a Protein Transduction Benchmark model. Software and code for reproducing the examples is available at https://github.com/MathBioCU/WENDy .

97 MATHEMATICS AND COMPUTING↗

Airborne hyperspectral imaging of nitrogen deficiency on crop traits and yield of maize by machine learning and radiative transfer modeling

Nitrogen is an essential nutrient that directly affects plant photosynthesis, crop yield, and biomass production for bioenergy crops, but excessive application of nitrogen fertilizers can cause environmental degradation. To achieve sustainable nitrogen fertilizer management for precision agriculture, there is an urgent need for nondestructive and high spatial resolution monitoring of crop nitrogen and its allocation to photosynthetic proteins as that changes over time. Here, we used visible to shortwave infrared (400–2400 nm) airborne hyperspectral imaging with high spatial (0.5 m) and spectral (3–5 nm) resolutions to accurately estimate critical crop traits, i.e., nitrogen, chlorophyll, and photosynthetic capacity (CO 2 -saturated photosynthesis rate, V max,27 ), at leaf and canopy scales, and to assess nitrogen deficiency on crop yield. We conducted three airborne campaigns over a maize (Zea mays L.) field during the growing season of 2019. Physically based soil-canopy Radiative Transfer Modeling (RTM) and data-driven approaches i.e. Partial-Least Squares Regression (PLSR) were used to retrieve crop traits from hyperspectral reflectance, with ground truth of leaf nitrogen, chlorophyll, V max,27 , Leaf Area Index (LAI), and harvested grain yield. To improve computational efficiency of RTMs, Random Forest (RF) was used to mimic RTM simulations to generate machine learning surrogate models RTM-RF. The results show that prior knowledge of soil background and leaf angle distribution can significantly reduce the ill-posed RTM retrieval. RTM-RF achieved a high accuracy to predict leaf chlorophyll content (R 2 = 0.73) and LAI (R 2 = 0.75). Meanwhile, PLSR exhibited better accuracy to predict leaf chlorophyll content (R 2 = 0.79), nitrogen concentration (R 2 = 0.83), nitrogen content (R 2 = 0.77), and V max,27 (R 2 = 0.69) but required measured traits for model training. We also found that canopy structure signals can enhance the use of spectral data to predict nitrogen related photosynthetic traits, as combining RTM-RF LAI and PLSR leaf traits well predicted canopy-level traits (leaf traits × LAI) including canopy chlorophyll (R 2 = 0.80), nitrogen (R 2 = 0.85) and V max,27 (R 2 = 0.82). Compared to leaf traits, we further found that canopy-level photosynthetic traits, particularly canopy V max,27 , have higher correlation with maize grain yield. This study highlights the potential for synergistic use of process-based and data-driven approaches of hyperspectral imaging to quantify crop traits that facilitate precision agricultural management to secure food and bioenergy production.

54 ENVIRONMENTAL SCIENCES↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗