Search NASA⌕ Search

Engineering topics

Nag, Ambarish

Publications and source records attributed to Nag, Ambarish.

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Green: Structural and Functional Genomic Characterization of Conserved Unannotated Green Lineage Proteins

Our overall objective is to improve and increase functional and structural predictions for a growing number of plant proteins of unknown structure and function (the Deep Green proteins), and to make our predictive data useful and accessible to the larger research community. The project is divided into five major objectives which include 1) Assembly and curation of Deep Green candidate protein sets; 2) in silico structural and functional predictions and network analyses; 3) assembly and validation of reverse genetic resources in Chlamydomonas reinhardtii (Chlamydomonas); 4) high throughput functional genomics characterization and prioritization in Chlamydomonas; and 5) structural validation of selected candidates and functional validation in two important reference plant species, Arabidopsis thaliana (Arabidopsis) and Setaria viridis (Setaria).

59 BASIC BIOLOGICAL SCIENCES↗

Deep Green Unannotated Protein Structures

The Deep Green list is based on the identification and curation of conserved unannotated proteins in three green lineage (Viridiplantae) model organisms; Arabidopsis thaliana, Chlamydomonas reinhardtii, and Setaria viridis. Preliminary characterization of Deep Green proteins and genes was done using various informatics tools and published data sets and is presented in Knoshaug, Sun, et al., 2023, submitted. The structures of these unannotated proteins were also predicted using AlphaFold (Jumper et al., 2021). The data deposited here are the AlphaFold structural predictions having the highest pLDDT score and thus identified as the best folded structure (ranked_0). These data enable others to do in-depth structural characterizations to aid in functional characterization leading to deeper understanding of plant biology. References: Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P. and Hassabis, D. (2021) Highly accurate protein structure prediction with AlphaFold. Nature, 596:583-589. Knoshaug, E. P., Sun, P., Nag, A., Nguyen, H., Mattoon, E. M., Zhang, N., Liu, J., Chen, C., Cheng, J., Zhang, R., St. John, P., and Umen, J. (submitted) Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii, Arabidopsis thaliana, and Setaria viridis.

09 BIOMASS FUELS↗

Towards deep computer vision for in-line defect detection in polymer electrolyte membrane fuel cell materials

Polymer Electrolyte Membrane (PEM) fuel cells are a promising source of alternative energy. However, their production is limited by a lack of well-established methods for quality control of their constituent materials like the membrane-electrode assembly during roll-to-roll manufacturing. One potential solution is the implementation of deep learning methods to detect unwanted defects through their detection in scanned images. Here we explore the detection of defects like scratches, pinholes, and scuffs in a sample dataset of PEM optical images using two deep learning algorithms: Patch Distribution Modeling (PaDiM) for unsupervised anomaly detection and Faster-RCNN for supervised object detection. Both methods achieve scores on performance metrics (ROC-AUC and PRO-AUC for PaDiM and AP for Faster-RCNN) that are comparable to their scores on benchmark datasets. These methods also have the potential to detect a wider range of defects compared to IR thermography and previous optical inspection methods. Overall, deep learning shows promise at detecting relevant defects of interest and has the potential to achieve real-time defect detection.

30 DIRECT ENERGY CONVERSION↗

Whole genome resequencing data from a collection of Clostridium Thermocellum strains

Clostridium thermocellum is an anaerobic thermophilic bacterium that natively ferments cellulose to ethanol and organic acids. This data set is a collection of whole genome resequencing data for several hundred strains of Clostridium thermocellum. It includes strains that have been engineered to increase ethanol production, strains that have been engineered to understand microbial physiology, and strains that have been adapted for desired phenotypes including increased ethanol tolerance. Resequencing data consists of paired Illumina reads, 100-150 bp on each end, with a ~500 bp insert size. One data file containing raw Illumina data (interleaved) is available for each strain. We also provide data describing the mutations identified in each strain, and distinguish between inherited and newly observed mutations. In addition to resequencing data, we also provide metadata describing the lineage of each strain, and any targeted genetic modifications.

resequencing bio energy fermentation↗