Search NASASearch

SEARCH · Search NASA

Results for “Molecular Sequence Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES

Evidence for autotrophic growth of purple sulfur bacteria using pyrite as electron and sulfur source

ABSTRACT Purple sulfur bacteria (PSB) are capable of anoxygenic photosynthesis via oxidizing reduced sulfur compounds and are considered key drivers of the sulfur cycle in a range of anoxic environments. In this study, we show that Allochromatium vinosum (a PSB species) is capable of autotrophic growth using pyrite as the electron and sulfur source. Comparative growth profile, substrate characterization, and transcriptomic sequencing data provided valuable insight into the molecular mechanisms underlying the bacterial utilization of pyrite and autotrophic growth. Specifically, the pyrite-supported cell cultures (“py”’) demonstrated robust but much slower growth rates and distinct patterns from their sodium sulfide-amended positive controls. Up to ~200-fold upregulation of genes encoding various c - and b -type cytochromes was observed in “py,” pointing to the high relevance of these molecules in scavenging and relaying electrons from pyrite to cytoplasmic metabolisms. Conversely, extensive downregulation of genes related to LH and RC complex components indicates that the electron source may have direct control over the bacterial cells’ photosynthetic activity. In terms of sulfur metabolism, genes encoding periplasmic or membrane-bound proteins (e.g., FccAB and SoxYZ) were largely upregulated, whereas those encoding cytoplasmic proteins (e.g., Dsr and Apr groups) are extensively suppressed. Other notable differentially expressed genes are related to flagella/fimbriae/pilin(+), metal efflux(+), ferrienterochelin(−), and [NiFe] hydrogenases(+). Characterization of the biologically reacted pyrite indicates the presence of polymeric sulfur. These results have, for the first time, put the interplay of PSB and transition metal sulfide chemistry under the spotlight, with the potential to advance multiple fields, including metal and sulfur biogeochemistry, bacterial extracellular electron transfer, and artificial photosynthesis. IMPORTANCE Microbial utilization of solid-phase substrates constitutes a critical area of focus in environmental microbiology, offering valuable insights into microbial metabolic processes and adaptability. Recent advancements in this field have profoundly deepened our knowledge of microbial physiology pertinent to these scenarios and spurred innovations in biosynthesis and energy production. Furthermore, research into interactions between microbes and solid-phase substrates has directly linked microbial activities to the surrounding mineralogical environments, thereby enhancing our understanding of the relevant biogeochemical cycles. Our study represents a significant step forward in this field by demonstrating, for the first time, the autotrophic growth of purple sulfur bacteria using insoluble pyrite (FeS2) as both the electron and sulfur source. The presented comparative growth profiles, substrate characterizations, and transcriptomic sequencing data shed light on the relationships between electron donor types, photosynthetic reaction center activities, and potential extracellular electron transfer in these organisms capable of anoxygenic photosynthesis. Furthermore, the findings of our study may provide new insights into early-Earth biogeochemical evolutions, offering valuable constraints for understanding the environmental conditions and microbial processes that shaped our planet’s history.

59 BASIC BIOLOGICAL SCIENCES

Dataset_for_Molecular_Motion_Below_the_Glass_Transition_A_Solid-State_NMR_Study_of_Siloxane_Polymer_Dynamics Study

This dataset contains solid-state 1H and 13C NMR relaxometry data, differential scanning calorimetry (DSC) data, and size exclusion chromatography (SEC/GPC) data supporting the study of sub-glass-transition (sub-Tg) molecular dynamics in a composition- and sequence-controlled series of diphenyl-substituted polysiloxanes (PDMS, 14Ph, 33Ph, 50Ph, 67Ph, and 100Ph; 0–100% diphenylsiloxane content by mole).All solid-state NMR data were acquired on a 200 MHz Bruker Avance III HD spectrometer using a static 7 mm HX probe or a 4 mm HX probe under 4 kHz magic-angle spinning. Raw Bruker TopSpin experiment folders are included for: (1) variable-temperature 1H lineshape measurements used to determine linewidth (FWHM) as a function of temperature across the glass transition; (2) 1H T1 (saturation recovery with solid-echo detection), probing nanosecond-scale dynamics near the 1H Larmor frequency; (3) 1H T1rho (direct spin-lock, 62.5 kHz), probing microsecond-scale segmental dynamics; (4) 13C-detected Lee–Goldburg cross-polarization 1H T1rho (LGCPH T1rho) for 33Ph and 50Ph, resolving aromatic and aliphatic proton environments; and (5) 13C T1 relaxation for 33Ph and 50Ph. Differential scanning calorimetry data (TA Instruments DSC 25, −150 to +120 °C, up to +300 °C for 100Ph, 10 °C/min) are included for all six compositions and support the glass-transition temperatures in Table 1 and Figure 1. Size exclusion chromatography data (Agilent 1200 Series, PL-Gel 300 mixed-C column, THF mobile phase, polystyrene calibration standards) are included for the three synthesized copolymers (33Ph, 50Ph, 67Ph) and support the number-average molecular weights in Table 1. Processed data include per-composition relaxation-time summaries (Excel), curve-fitting and Bloembergen-Purcell-Pound (BPP) model analysis notebooks (Jupyter/Python), and Igor Pro (.pxp) master files used to generate the manuscript's figures.

Bloembergen-Purcell-Pound theory

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi

Soil microbiome resilience to short-term (30 days, 90 days) and long-term (1000 days) drought

This dataset contains data used for the paper "Drought duration does not impact soil microbiome resilience". The Related References will be updated with a full citation when available. Increasing global droughts exert large but poorly understood effects on the microbial communities and ecology of soil. Microbial communities generally show resilience and return to pre-drought conditions when short-term droughted soils are rewet; soils exposed to long-term drought, however, often show a lag upon rewetting, after which microbial communities may or may not return to their pre-stressed conditions. Though short-term droughts have been widely studied, long-term drought manipulation experiments remain rare, especially those that compare microbial response to short-term and long-term drought in tandem. We conducted a 1000-day drought simulation in controlled laboratory conditions with soil cores collected from a tidal freshwater ecosystem in Washington state, USA, and subsequently exposed them to rewetting for two weeks. We also included short-term (30-day and 90-day) drought and rewet treatments to directly compare microbial community and organic matter responses across drought durations. We found distinct microbial taxa belonging to Firmicutes and Actinobacteria enriched after the 1000-day drought, but not after the short-term droughts. While we hypothesized that the microbial community would recover from a short-term drought after rewetting to resemble pre-drought conditions, our results revealed community dissimilarities between rewet and pre-drought conditions across all drought durations. These findings suggest unique microbial life history strategies within certain microbial phyla that make them successful colonizers during an extended drought period, and the influence of environmental and physiological context on microbial responses to rewetting. The 16SrRNA gene amplicon dataset contains processed DNA sequences in the form of an ASV table with raw unrarefied read counts and representative sequences in .fasta format as described in the ESS-DIVE amplicon sequence reporting format (https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format/instructions). The Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) dataset consists of processed files containing presence absence data of molecular formulae and molecular characterization of FTICR resolved peaks. The Nuclear Magnetic Resonance (NMR) dataset contains files relevant to NMR spectra and peaks. A sample key file and a sample metadata file is included for the FTICR/NMR and 16S dataset respectively.

1000-day drought

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa

Tuning Copolymer Microstructure Using Ring-Opening Cross-Metathesis Polymerization

The capability of ring-opening cross-metathesis (RO/CM) polymerization to produce alternating copolymers was studied. By treating commercial polybutadiene (PB) with bulky oxanorbornene monomers and Ru-based olefin metathesis catalysts, alternating copolymers were produced under mild conditions with high sequence fidelities. Here, we found that alternating copolymers could be produced starting from a variety of butadiene sources including PB, cyclooctadiene (COD), or t,t,t-1,5,9-cyclododecatriene (CDT), highlighting for the first time the kinetic pathway independence of this process. Kinetic copolymerization analysis of an oxanorbonene monomer with CDT revealed that much higher monomer conversions were obtained compared with the analogous homopolymerizations and showed evidence of alternating monomer incorporation. Copolymerization of these monomers also enabled good control when targeting different molecular weights. Copolymer thermal analysis revealed a strong correlation between thermal behavior and alternating sequence fidelity, providing a second lever beyond composition to tune thermal behavior. These data demonstrate that a broad variety of polymer microstructures can be accessed via RO/CM polymerization and highlight the potential of CDT in alternating copolymer synthesis.

Foster, Jeffrey C. [Oak Ridge National Laboratory

A modular and extensible CHARMM-compatible model for all-atom simulation of polypeptoids

Peptoids (N-substituted glycines) are a class of sequence-defined synthetic peptidomimetic polymers with applications including drug delivery, catalysis, and biomimicry. Classical molecular simulations have been used to predict and understand the conformational dynamics of single chains and their self-assembly into morphologies including sheets, tubes, spheres, and fibrils. The CGenFF-NTOID model based on the CHARMM General Force Field has demonstrated success in accurate all-atom molecular modeling of peptoid structure and thermodynamics. Extension of this force field to new peptoid side chains has historically required reparameterization of side chain bonded interactions against ab initio data. This fitting protocol improves the accuracy of the force field but is also burdensome and precludes modular extensibility of the model to arbitrary peptoid sequences. In this work, we develop and demonstrate a Modular Side Chain CGenFF-NTOID (MoSiC-CGenFF-NTOID) as an extension of CGenFF-NTOID employing a modular decomposition of the peptoid backbone and side chain parameterizations, wherein arbitrary side chains within the large family of substituted methyl groups (i.e., –CH 3 , –CH 2 R, –CHRR', and –CRR'R") are directly ported from CGenFF. We validate this approach against ab initio calculations and experimental data to develop a MoSiC-CGenFF-NTOID model for all 20 natural amino acid side chains along with 13 commonly used synthetic side chains and present an extensible paradigm to efficiently determine whether a novel side chain can be directly incorporated into the model or whether refitting of the CGenFF parameters is warranted. We make the model freely available to the community along with a tool to perform automated initial structure generation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Towards verifiable cancer digital twins: tissue level modeling protocol for precision medicine

Cancer exhibits substantial heterogeneity, manifesting as distinct morphological and molecular variations across tumors, which frequently undermines the efficacy of conventional oncological treatments. Developments in multiomics and sequencing technologies have paved the way for unraveling this heterogeneity. Nevertheless, the complexity of the data gathered from these methods cannot be fully interpreted through multimodal data analysis alone. Mathematical modeling plays a crucial role in delineating the underlying mechanisms to explain sources of heterogeneity using patient-specific data. Intra-tumoral diversity necessitates the development of precision oncology therapies utilizing multiphysics, multiscale mathematical models for cancer. This review discusses recent advancements in computational methodologies for precision oncology, highlighting the potential of cancer digital twins to enhance patient-specific decision-making in clinical settings. We review computational efforts in building patient-informed cellular and tissue-level models for cancer and propose a computational framework that utilizes agent-based modeling as an effective conduit to integrate cancer systems models that encode signaling at the cellular scale with digital twin models that predict tissue-level response in a tumor microenvironment customized to patient information. Furthermore, we discuss machine learning approaches to building surrogates for these complex mathematical models. These surrogates can potentially be used to conduct sensitivity analysis, verification, validation, and uncertainty quantification, which is especially important for tumor studies due to their dynamic nature.

60 APPLIED LIFE SCIENCES

Identifying impacts of contact tracing on HIV epidemiological inference from phylogenetic data

Abstract Robust sampling methods are foundational to inferences using phylogenies. Yet the impact of using contact tracing, a type of non-uniform sampling used in public health applications such as infectious disease outbreak investigations, has not been investigated in the molecular epidemiology field. To understand how contact tracing influences a recovered phylogeny, we developed a new simulation tool called SEEPS (Sequence Evolution and Epidemiological Process Simulator) that allows for the simulation of contact tracing and the resulting transmission tree, pathogen phylogeny, and corresponding virus genetic sequences. Importantly, SEEPS takes within-host evolution into account when generating pathogen phylogenies and sequences from transmission histories. Using SEEPS, we demonstrate that contact tracing can significantly impact the structure of the resulting tree, as described by popular tree statistics. Contact tracing generates phylogenies that are less balanced than the underlying transmission process, less representative of the larger epidemiological process, and affects the internal/external branch length ratios that characterize specific epidemiological scenarios. We also examined real data from a 2007–2008 Swedish HIV-1 outbreak and the broader 1998–2010 European HIV-1 epidemic to highlight the differences in contact tracing and expected phylogenies. Aided by SEEPS, we show that the data collection of the Swedish outbreak was strongly influenced by contact tracing even after downsampling, while the broader European Union epidemic showed little evidence of universal contact tracing, agreeing with the known epidemiological information about sampling and spread. Overall, our results highlight the importance of including possible non-uniform sampling schemes when examining phylogenetic trees. For that, SEEPS serves as a useful tool to evaluate such impacts, thereby facilitating better phylogenetic inferences of the characteristics of a disease outbreak. SEEPS is available at https://github.com/MolEvolEpid/SEEPS.

Virology

RhizoGrid Indexed Sorghum Rhizosphere Multi-Omics

PerCon SFA project data dentification of spatially resolved biomarkers of drought in Sorghum bicolor rhizosphere molecular-microbe interactions using a novel root cartography "RhizoGrid" system for sampling plants under drought and control conditions across 10 equally sized root zone environments (4 quadrants each). Each quadrant was sampled and processed for 16S amplicon, metabolomics, and X-ray computed tomography (XCT). Data download includes experimental metadata and results files for 16S rRNA sequence analysis of microbial community assembly (processed data files), liquid chromatography mass spectrometry (LC-MS) metabolomics analysis of microbial community root exudates (processed data files), X-ray computed tomography (XCT) spatial gradient analysis (raw and processed data files) of microbial community composition, and related computational modeling outputs.

59 BASIC BIOLOGICAL SCIENCES

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N

The 1.3 Å resolution structure of the truncated group Ia type IV pilin from Pseudomonas aeruginosa strain P1

The type IV pilus is a diverse molecular machine capable of conferring a variety of functions and is produced by a wide range of bacterial species. The ability of the pilus to perform host-cell adherence makes it a viable target for the development of vaccines against infection by human pathogens such as Pseudomonas aeruginosa . Here, the 1.3 Å resolution crystal structure of the N-terminally truncated type IV pilin from P. aeruginosa strain P1 (ΔP1) is reported, the first structure of its phylogenetically linked group (group I) to be discussed in the literature. The structure was solved from X-ray diffraction data that were collected 20 years ago with a molecular-replacement search model generated using AlphaFold ; the effectiveness of other search models was analyzed. Examination of the high-resolution ΔP1 structure revealed a solvent network that aids in maintaining the fold of the protein. On comparing the sequence and structure of P1 with a variety of type IV pilins, it was observed that there are cases of higher structural similarities between the phylogenetic groups of P. aeruginosa than there are between the same phylogenetic group, indicating that a structural grouping of pilins may be necessary in developing antivirulence drugs and vaccines. These analyses also identified the α–β loop as the most structurally diverse domain of the pilins, which could allow it to serve a role in pilus recognition. Studies of ΔP1 in vitro polymerization demonstrate that the optimal hydrophobic catalyst for the oligomerization of the pilus from strain K122 is not conducive for pilus formation of ΔP1; a model of a three-start helical assembly using the ΔP1 structure indicates that the α–β loop and the D-loop prevent in vitro polymerization.

Bragagnolo, Nicholas

Machine learning approaches for influenza A virus risk assessment identifies predictive correlates using ferret model in vivo data

In vivo assessments of influenza A virus (IAV) pathogenicity and transmissibility in ferrets represent a crucial component of many pandemic risk assessment rubrics, but few systematic efforts to identify which data from in vivo experimentation are most useful for predicting pathogenesis and transmission outcomes have been conducted. To this aim, we aggregated viral and molecular data from 125 contemporary IAV (H1, H2, H3, H5, H7, and H9 subtypes) evaluated in ferrets under a consistent protocol. Three overarching predictive classification outcomes (lethality, morbidity, transmissibility) were constructed using machine learning (ML) techniques, employing datasets emphasizing virological and clinical parameters from inoculated ferrets, limited to viral sequence-based information, or combining both data types. Among 11 different ML algorithms tested and assessed, gradient boosting machines and random forest algorithms yielded the highest performance, with models for lethality and transmission consistently better performing than models predicting morbidity. Comparisons of feature selection among models was performed, and highest performing models were validated with results from external risk assessment studies. Our findings show that ML algorithms can be used to summarize complex in vivo experimental work into succinct summaries that inform and enhance risk assessment criteria for pandemic preparedness that take in vivo data into account.

59 BASIC BIOLOGICAL SCIENCES

Parallel measurement of transcriptomes and proteomes from same single cells using nanodroplet splitting

Single-cell multiomics provides comprehensive insights into gene regulatory networks, cellular diversity, and temporal dynamics. Here, we introduce nanoSPLITS (nanodroplet SPlitting for Linked-multimodal Investigations of Trace Samples), an integrated platform that enables global profiling of the transcriptome and proteome from same single cells via RNA sequencing and mass spectrometry-based proteomics, respectively. Benchmarking of nanoSPLITS demonstrates high measurement precision with deep proteomic and transcriptomic profiling of single-cells. We apply nanoSPLITS to cyclin-dependent kinase 1 inhibited cells and found phospho-signaling events could be quantified alongside global protein and mRNA measurements, providing insights into cell cycle regulation. We extend nanoSPLITS to primary cells isolated from human pancreatic islets, introducing an efficient approach for facile identification of unknown cell types and their protein markers by mapping transcriptomic data to existing large-scale single-cell RNA sequencing reference databases. Accordingly, we establish nanoSPLITS as a multiomic technology incorporating global proteomics and anticipate the approach will be critical to furthering our understanding of biological systems.

59 BASIC BIOLOGICAL SCIENCES

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State

Cholesterol-dependent enzyme activity of human TSPO1

The amino acid sequence of the tryptophan-rich sensory proteins (TSPO) is substantially conserved throughout all kingdoms of life. Human mitochondrial TSPO1 (HsTSPO1) binds to porphyrins and steroids, although its interactions with these molecules remains unknown.HsTSPO1 is associated with numerous physiological and pathological disorders, but the underlying molecular mechanisms are unknown. Here, we disclose the finding of human mitochondrial TSPO as a cholesterol-dependent protoporphyrin IX oxygenase. The results of our biochemical characterization are consistent with structural data and evolutionary analysis. The dependence ofHsTSPO1 activity on cholesterol may be the result of the coevolution of this membrane protein with the membrane system. Our study provides a molecular foundation for comprehending the various roles played by mitochondrial TSPO in normal physiological and pathological situations.

Science & Technology - Other Topics