Search NASA⌕ Search

Engineering topics

St. John, Peter

Publications and source records attributed to St. John, Peter.

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

DOE STTR Phase I Final Report Report: Machine-learning Based Prediction of Thermal Limits for Conjugated Organic Materials

The SCANN-DT Phase I STTR project led by NLM Photonics and partnering with the National Renewable Energy Laboratory (NREL) sought to apply machine learning techniques based on graph neural networks (GNNs) towards the prediction of decomposition temperatures (Td) of organic semiconductors, based on prior work on bond dissociation energy (BDE) prediction as implemented in NREL’s ALFABET prediction tool. Using a curated set of experimental decomposition energies, the project examined GNN-based, classical quantitative structure-property relationship (QSPR) based on DFT calculations, and combinations of both methods to predict Td. While the best MAEs in Td achieved were near 40°C, below project targets, the project led to improvements in the ALFABET model, improvements in cloud-based implementations of NWChem software, and a substantial dataset of calculations on medium-sized conjugated organic molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep Green: Structural and Functional Genomic Characterization of Conserved Unannotated Green Lineage Proteins

Our overall objective is to improve and increase functional and structural predictions for a growing number of plant proteins of unknown structure and function (the Deep Green proteins), and to make our predictive data useful and accessible to the larger research community. The project is divided into five major objectives which include 1) Assembly and curation of Deep Green candidate protein sets; 2) in silico structural and functional predictions and network analyses; 3) assembly and validation of reverse genetic resources in Chlamydomonas reinhardtii (Chlamydomonas); 4) high throughput functional genomics characterization and prioritization in Chlamydomonas; and 5) structural validation of selected candidates and functional validation in two important reference plant species, Arabidopsis thaliana (Arabidopsis) and Setaria viridis (Setaria).

59 BASIC BIOLOGICAL SCIENCES↗

EvoProtGrad (Directed Evolution for Proteins with Gradients) [SWR-23-48]

A Python package for directed evolution on a protein sequence with gradient-based discrete Markov chain monte carlo (MCMC). Users are able to compose custom models that map sequence to function with pretrained models, including protein language models (PLMs), to guide and constrain search. Our package natively integrates with the HuggingFace platform and supports PLMs from transformers. Our MCMC sampler identifies promising amino acids to mutate via model gradients taken with respect to the input (i.e., sensitivity analysis). We allow users to compose their own custom target function for MCMC by leveraging the Product of Experts MCMC paradigm. Each model is an "expert" that contributes its own knowledge about the protein's fitness landscape to the overall target function. The sampler is designed to be more efficient and effective than brute force and random search while maintaining most of the generality and flexibility. Additional information can be found in the related publication: https://iopscience.iop.org/article/10.1088/2632-2153/accacd

Emami, Patrick↗

Deep Green Unannotated Protein Structures

The Deep Green list is based on the identification and curation of conserved unannotated proteins in three green lineage (Viridiplantae) model organisms; Arabidopsis thaliana, Chlamydomonas reinhardtii, and Setaria viridis. Preliminary characterization of Deep Green proteins and genes was done using various informatics tools and published data sets and is presented in Knoshaug, Sun, et al., 2023, submitted. The structures of these unannotated proteins were also predicted using AlphaFold (Jumper et al., 2021). The data deposited here are the AlphaFold structural predictions having the highest pLDDT score and thus identified as the best folded structure (ranked_0). These data enable others to do in-depth structural characterizations to aid in functional characterization leading to deeper understanding of plant biology. References: Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P. and Hassabis, D. (2021) Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. Knoshaug, E. P., Sun, P., Nag, A., Nguyen, H., Mattoon, E. M., Zhang, N., Liu, J., Chen, C., Cheng, J., Zhang, R., St. John, P., and Umen, J. (submitted) Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii, Arabidopsis thaliana, and Setaria viridis.

09 BIOMASS FUELS↗