Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein domains”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ion-selective conformational stabilization of a disordered repeats-in-toxin protein domain

Ion-binding intrinsically disordered proteins (IDPs) recruit and bind to specific metal ions to perform critical biological functions. In proteins where ion binding and structural transitions are coupled, interactions with off-target toxic metals can dramatically disrupt protein structure and function, exemplified by lead and mercury poisoning. Understanding the complex mechanisms underlying how IDPs exclude or allow binding to different ionic species is crucial for addressing the origins of metal toxicity in biological systems. Here, we elucidate mechanisms of ion selectivity in an IDP that adopts a structure upon Ca 2+ binding. We probed ion-induced conformational changes of a repeats-in-toxin (RTX) protein domain in the presence of different ion ligands—Mg 2+ , Ca 2+ , Sr 2+ , and Ba 2+ —with chemical similarities but drastically different ionic radii. RTX adopts ion-selective conformations measured by x-ray crystallography, small-angle x-ray scattering, and circular dichroism. High-resolution x-ray structures reveal that Sr 2+ induces a nearly identical RTX structure as natively binding Ca 2+ , enabled by the intrinsic flexibility and disorder of the protein. Small-angle x-ray scattering and circular dichroism indicate that smaller Mg 2+ does not induce a significant conformational change in RTX, whereas larger Ba 2+ induces a partially folded structure. These results highlight the importance of geometric constraints imposed by protein structure in determining metal ion selectivity, yielding insights into how off-target ion binding may result in protein misfolding and malfunction.

Gudinas, Alana P. [Stanford Univ., CA (United Stat↗

Powdery mildew effectors AVR A1 and BEC1016 target the ER J‐domain protein Hv ERdj3B required for immunity in barley

Abstract The barley powdery mildew fungus, Blumeria hordei (Bh), secretes hundreds of candidate secreted effector proteins (CSEPs) to facilitate pathogen infection and colonization. One of these, CSEP0008, is directly recognized by the barley nucleotide‐binding leucine‐rich‐repeat (NLR) receptor MLA1 and therefore is designated AVR A1 . Here, we show that AVR A1 and the sequence‐unrelated Bh effector BEC1016 (CSEP0491) suppress immunity in barley. We used yeast two‐hybrid next‐generation interaction screens (Y2H‐NGIS), followed by binary Y2H and in planta protein–protein interactions studies, and identified a common barley target of AVR A1 and BEC1016, the endoplasmic reticulum (ER)‐localized J‐domain protein Hv ERdj3B. Silencing of this ER quality control (ERQC) protein increased Bh penetration. Hv ERdj3B is ER luminal, and we showed using split GFP that AVR A1 and BEC1016 translocate into the ER signal peptide‐independently. Overexpression of the two effectors impeded trafficking of a vacuolar marker through the ER; silencing of Hv ERdj3B also exhibited this same cellular phenotype, coinciding with the effectors targeting this ERQC component. Together, these results suggest that the barley innate immunity, preventing Bh entry into epidermal cells, requires ERQC. Here, the J‐domain protein Hv ERdj3B appears to be essential and can be regulated by AVR A1 and BEC1016. Plant disease resistance often occurs upon direct or indirect recognition of pathogen effectors by host NLR receptors. Previous work has shown that AVR A1 is directly recognized in the cytosol by the immune receptor MLA1. We speculate that the AVR A1 J‐domain target being inside the ER, where it is inapproachable by NLRs, has forced the plant to evolve this challenging direct recognition.

54 ENVIRONMENTAL SCIENCES↗

PLAT domain protein 1 (PLAT1/PLAFP) binds to the Arabidopsis thaliana plasma membrane and inserts a lipid

Robust agricultural yields depend on the plant's ability to fix carbon amid variable environmental conditions. Over seasonal and diurnal cycles, the plant must constantly adjust its metabolism according to available resources or external stressors. The metabolic changes that a plant undergoes in response to stress are well understood, but the long-distance signaling mechanisms that facilitate communication throughout the plant are less studied. The phloem is considered the predominant conduit for the bidirectional transport of these signals in the form of metabolites, nucleic acids, proteins, and lipids. Lipid trafficking through the phloem in particular attracted our attention due to its reliance on soluble lipid-binding proteins (LBP) that generate and solubilize otherwise membrane-associated lipids. The Phloem Lipid-Associated Family Protein (PLAFP) from Arabidopsis thaliana is generated in response to abiotic stress as is its lipid-ligand phosphatidic acid (PA). PLAFP is proposed to transport PA through the phloem in response to drought stress. To understand the interactions between PLAFP and PA, nearly 100 independent systems comprised of the protein and one PA, or a plasma membrane containing varying amounts of PA, were simulated using atomistic classical molecular dynamics methods. In these simulations, PLAFP is found to bind to plant plasma membrane models independent of the PA concentration. When bound to the membrane, PLAFP adopts a binding pose where W41 and R82 penetrate the membrane surface and anchor PLAFP. This triggers a separation of the two loop regions containing W41 and R82. Subsequent simulations indicate that PA insert into the β -sandwich of PLAFP, driven by interactions with multiple amino acids besides the W41 and R82 identified during the insertion process. Finally, fine-tuning the protein-membrane and protein-PA interface by mutating a selection of these amino acids may facilitate engineering plant signaling processes by modulating the binding response.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

Archaeal protein containing domain of unknown function 2193 undergoes oligomeric reconfiguration upon iron–sulfur cluster binding

Methanogenic archaea are particularly rich in iron–sulfur proteins, yet their roles remain largely enigmatic. Here, we characterized aMethanococcus voltae(Mvo) protein from the domain of unknown function (DUF) 2193 family, a group of proteins present primarily in archaea and characterized by a conserved cysteine‐rich C‐terminal motif.MvoDUF2193 was heterologously expressed and characterized by a range of spectroscopic and analytical methods. The results demonstrate thatMvoDUF2193 binds a single [4Fe–4S] cluster per subunit and that cluster occupancy regulates the transition from an apo tetramer to a [4Fe–4S] monomeric form. We hypothesize thatMvoDUF2193 serves a regulatory role in the cell, mediated by [Fe–S] cluster binding and changes in oligomeric state.

Biochemistry & Molecular Biology↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

High-Throughput Directed Evolution of Marine Microalgae and Phototrophic Consortia for Improved Biomass Yields (Final Report)

Primary project achievements include using selective pressures (O 2 , light, temperature) and developing culturing regimes for the diatom Nitzschia inconspicua str. hildebrandi to attain enrichments with an ~90% increase in areal biomass productivity relative to the parental strain under pond-mimicking conditions with high O 2 stress in laboratory bioreactors. The resulting strain (GAI-337) was tested further for dilution time, culture density, CO 2 supplementation, pH, temperature, and dissolved O 2 concentration under outdoor pond-mimicking conditions to improve areal productivities. These experiments yielded an optimum harvest and dilution time just after sunset, ~0.45 g AFDW L -1 initial culture density for maximal productivities, no requirement for CO 2 supplementation or pH control, maximal performance under a diel temperature curve going from 24 °C at night to 36 °C during the day, and benefits from some O 2 removal from the culture by bubbling with air. Using pond-mimicking laboratory bioreactors, N. inconspicua GAI-337 achieved ~42 g AFDW m -2 d -1 . Nutrient limitation experiments resulted in a biomass composition that equated to ~160 Gallons of Gasoline Equivalent energy per ton AFDW, highlighting the potential of GAI-337 as a promising renewable fuel feedstock strain. Genome resequencing has revealed genome alterations potentially contributing to the improved growth of GAI-337 in the laboratory. Based on the comparative analyses of the GAI-337 and GAI-229 (reference) strains, we identified 144 single nucleotide substitutions that resulted in amino acid change, 7 single nucleotide substitutions that resulted in protein truncation; 5 deletions; and 1 frameshift mutation. From the mutations that potentially affect expression of functionally annotated genes, particular interest was noted for an interferon-induced 6-16 family protein that may be involved in the host immune response against microbe invasion; the chaperone protein DnaK, which may function to protect the folding of proteins within the cell; and SPRY domain protein that is found in many eukaryotic proteins important in cell signaling pathways. Transcriptome analysis revealed over 1000 genes with increased transcript levels. Many of these and many of the genes with mutations are not yet functionally annotated and an increased bioinformatics effort is necessary to more completely analyze the Nitzschia inconspicua genome. Adaptive laboratory evolution (ALE) was performed for over 300 days using consecutive 0.5°C temperature increases in a constant temperature incubator to attain greater thermal tolerance in Nitzschia inconspicua. The adapted strain was able to grow at a constant temperature of 37.5°C; whereas this constant temperature was lethal to the parental control, which had an upper temperature boundary of 35.5°C prior to adaptive evolution. Several high-temperature clonal isolates were obtained from the evolved population following ALE, and increased temperature tolerance was observed in clonal adapted cultures. The final temperature adaptation was maintained through cryopreservation and was observed in multiple clonal isolates, including multiple clonal isolates with significantly increased cell size, indicating the potential occurrence of a sexual cycle during the clonal isolation process. A survey of Nannochloropsis strains was conducted for tolerances to high pH and high bicarbonate media. Nannochloropsis granulata showed promising growth in diel bioreactors and was successfully grown at the GAI Kauai farm site in long-term growth campaigns. Co-culturing using Nitzschia inconspicua, Nannochloropsis and a cyanobacterium were assembled in the laboratory to determine if productivity synergies could be attained. Although all strains grew well in the laboratory high-bicarbonate media individually, the cyanobacterium quickly outgrew the other strains in the laboratory consortium pushing the co-culture away from a diverse (and potentially synergistic assemblage) phototroph culture towards a monoculture dominated by the cyanobacterium. Several outdoor growth campaigns were conducted, with productivities ranging between 10-20 g/m 2 /d of biomass. The best performing strain in the laboratory (GAI-337) did not outperform reference strains at the Kauai farm under the conditions used. Addition growth campaigns are necessary under conditions that result in higher biomass (>20 g/m 2 /d) and that attain higher O 2 levels are likely necessary. Initial data indicate that the thermally adapted strain did slightly better than the control strain at higher temperatures; however, additional campaigns are necessary to establish statistical significance. In summary, Nitzschia inconspicua is able to attain exemplary biomass and lipid yields in the laboratory bioreactors. Strain evolution to both O 2 and temperature resulted in targeted strain improvements. Additional outdoor campaigns are necessary to determine if laboratory improvements translate to the field.

09 BIOMASS FUELS↗

Variation of Structure and Cellular Functions of Type IA Topoisomerases across the Tree of Life

Topoisomerases regulate the topological state of cellular genomes to prevent impediments to vital cellular processes, including replication and transcription from suboptimal supercoiling of double-stranded DNA, and to untangle topological barriers generated as replication or recombination intermediates. The subfamily of type IA topoisomerases are the only topoisomerases that can alter the interlinking of both DNA and RNA. In this article, we provide a review of the mechanisms by which four highly conserved N-terminal protein domains fold into a toroidal structure, enabling cleavage and religation of a single strand of DNA or RNA. We also explore how these conserved domains can be combined with numerous non-conserved protein sequences located in the C-terminal domains to form a diverse range of type IA topoisomerases in Archaea, Bacteria, and Eukarya. There is at least one type IA topoisomerase present in nearly every free-living organism. The variation in C-terminal domain sequences and interacting partners such as helicases enable type IA topoisomerases to conduct important cellular functions that require the passage of nucleic acids through the break of a single-strand DNA or RNA that is held by the conserved N-terminal toroidal domains. In addition, this review will exam a range of human genetic disorders that have been linked to the malfunction of type IA topoisomerase.

59 BASIC BIOLOGICAL SCIENCES↗

Bipartite chromatin recognition by Hop1 from two diverged Holozoa

In meiosis, ploidy reduction is driven by a complex series of DNA breakage and recombination events between homologous chromosomes, orchestrated by meiotic HORMA domain proteins (HORMADs). Meiotic HORMADs possess a central chromatin binding region (CBR) whose architecture varies across eukaryotic groups. Here, we determine high-resolution crystal structures of the meiotic HORMAD CBR from two diverged aquatic Holozoa,Schistosoma mansoniandPatiria miniata, which reveal tightly associated plant homeodomain (PHD) and winged helix-turn-helix (wHTH) domains. We show that PHD–wHTH CBRs bind duplex DNA through their wHTH domains, and identify key residues that disrupt this interaction. Combining experimental and predicted structures, we show that the CBRs’ PHDs likely interact with the tail of histone H3, and may discriminate between unmethylated and trimethylated H3 lysine 4. Finally, we show that Holozoa Hop1 CBRs bind nucleosomes in vitro in a bipartite manner involving both the PHD and wHTH domain. Our data reveal how meiotic HORMADs with PHD–wHTH CBRs can bind chromatin and potentially discriminate between chromatin states to drive meiotic recombination to specific chromosomal regions.

Life Sciences & Biomedicine - Other Topics↗

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Generating Protein Structures for Pathway Discovery Using Deep Learning

Resolving the intricate details of biological phenomena at the molecular level is fundamentally limited by both length- and time scales that can be probed experimentally. Molecular dynamics (MD) simulations at various scales are powerful tools frequently employed to offer valuable biological insights beyond experimental resolution. However, while it is relatively simple to observe long-lived, stable configurations of, for example, proteins, at the required spatial resolution, simulating the more interesting rare transitions between such states often takes orders of magnitude longer than what is feasible even on the largest supercomputers available today. One common aspect of this challenge is pathway discovery, where the start and end states of a scientific phenomenon are known or can be approximated, but the mechanistic details in between are unknown. Here, we propose a representation-learning-based solution that uses interpolation and extrapolation in an abstract representation space to synthesize potential transition states, which are automatically validated using MD simulations. The new simulations of the synthesized transition states are subsequently incorporated into the representation learning, leading to an iterative framework for targeted path sampling. Our approach is demonstrated by recovering the transition of a RAS-RAF protein domain (CRD) from membrane-free to interacting with the membrane using coarse-grain MD simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

Novel Z-DNA binding domains in giant viruses

Z-nucleic acid structures play vital roles in cellular processes and have implications in innate immunity due to their recognition by Zα domains containing proteins (Z-DNA/Z-RNA binding proteins, ZBPs). Although Zα domains have been identified in six proteins, including viral E3L, ORF112, and I73R, as well as, cellular ADAR1, ZBP1, and PKZ, their prevalence across living organisms remains largely unexplored. In this study, we introduce a computational approach to predict Zα domains, leading to the revelation of previously unidentified Zα domain-containing proteins in eukaryotic organisms, including non-metazoan species. Our findings encompass the discovery of new ZBPs in previously unexplored giant viruses, members of the Nucleocytoviricota phylum. Through experimental validation, we confirm the Zα functionality of select proteins, establishing their capability to induce the B-to-Z conversion. Additionally, we identify Zα-like domains within bacterial proteins. While these domains share certain features with Zα domains, they lack the ability to bind to Z-nucleic acids or facilitate the B-to-Z DNA conversion. Our findings significantly expand the ZBP family across a wide spectrum of organisms and raise intriguing questions about the evolutionary origins of Zα-containing proteins. Moreover, our study offers fresh perspectives on the functional significance of Zα domains in virus sensing and innate immunity and opens avenues for exploring hitherto undiscovered functions of ZBPs.

59 BASIC BIOLOGICAL SCIENCES↗

Structural and Biophysical Properties of a [4Fe-4S] Ferredoxin-Like Protein from Synechocystis sp. PCC 6803 with a Unique Two Domain Structure

Electron carrier proteins (ECPs), binding iron-sulfur clusters, are vital components within the intricate network of metabolic and photosynthetic reactions. They play a crucial role in the distribution of reducing equivalents. In Synechocystis sp. PCC 6803, the ECP network includes at least nine ferredoxins. Previous research, including global expression analyses and protein binding studies, has offered initial insights into the functional roles of individual ferredoxins within this network. This study primarily focuses on Ferredoxin 9 (slr2059). Through sequence analysis and computational modeling, Ferredoxin 9 emerges as a unique ECP with a distinctive two-domain architecture. It consists of a C-terminal iron-sulfur binding domain and an N-terminal domain with homology to Nil-domain proteins, connected by a structurally rigid 4-amino acid linker. Notably, in contrast to canonical [2Fe-2S] ferredoxins exemplified by PetF (ssl0020), which feature highly acidic surfaces facilitating electron transfer with photosystem I reaction centers, models of Ferredoxin 9 reveal a more neutral to basic protein surface. Using a combination of electron paramagnetic resonance spectroscopy and square-wave voltammetry on heterologously produced Ferredoxin 9, this study demonstrates that the protein coordinates 2x[4Fe-4S]2+/1+ redox-active and magnetically interacting clusters, with measured redox potentials of -420 +/- 9 mV and -516 +/- 10 mV vs SHE. A more in-depth analysis of Fdx9's unique structure and protein sequence suggests that this type of Nil-2[4Fe-4S] multi-domain ferredoxin is well conserved in cyanobacteria, bearing structural similarities to proteins involved in homocysteine synthesis in methanogens.

cyanobacteria↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Foldy: An open-source web application for interactive protein structure analysis

Foldy is a cloud-based application that allows non-computational biologists to easily utilize advanced AI-based structural biology tools, including AlphaFold and DiffDock. With many deployment options, it can be employed by individuals, labs, universities, and companies in the cloud without requiring hardware resources, but it can also be configured to utilize locally available computers. Foldy enables scientists to predict the structure of proteins and complexes up to 6000 amino acids with AlphaFold, visualize Pfam annotations, and dock ligands with AutoDock Vina and DiffDock. In our manuscript, we detail Foldy’s interface design, deployment strategies, and optimization for various user scenarios. We demonstrate its application through case studies including rational enzyme design and analyzing proteins with domains of unknown function. Furthermore, we compare Foldy’s interface and management capabilities with other open and closed source tools in the field, illustrating its practicality in managing complex data and computation tasks. Our manuscript underlines the benefits of Foldy as a day-to-day tool for life science researchers, and shows how Foldy can make modern tools more accessible and efficient.

59 BASIC BIOLOGICAL SCIENCES↗